Application Load Balancer — Advanced Routing & Target Groups
Recap: Where the Fleet Left Off
Day 17 ended on a specific operational problem: an instance that has just been launched is not yet useful, and an instance that is about to be terminated is still holding live connections. Lifecycle hooks solve the first half of that problem by pausing an instance in Pending:Wait so bootstrap scripts can finish before it is considered healthy, and the mirror-image hook pauses it in Terminating:Wait so drain logic can run. The canonical drain action is to deregister the instance from the load balancer and wait for in-flight requests to complete, which is the point where yesterday's material and today's meet. Warm pools attack the same latency problem from the other direction: instead of making boot faster, they keep pre-initialized stopped instances on standby so scale-out promotes an already-bootstrapped instance rather than paying full boot and bootstrap cost.
Today extends that thread rather than replacing it. Lifecycle hooks and warm pools govern when a target is allowed to receive traffic; the Application Load Balancer governs which target receives it and in what proportion. The deregistration step inside a Terminating:Wait hook is only meaningful because the ALB maintains a target group health model that can be told to stop sending new requests to a specific target. Once you can control membership and health, the next question is whether you can control weight — and that is where weighted target groups turn a load balancer into a deployment mechanism.
Foundations You'll Need Today
Today's material is about a load balancer, and the whole day assumes you already know what one is and what its moving parts are called. If you have never configured one, the vocabulary can make an otherwise simple idea sound complicated, so let's build it up from the problem it solves.
What a load balancer is, and why you would want one
Imagine you run a web application on a single server. Users type your domain name, that name resolves to the server's IP address, and the server answers. This works until it doesn't: if the server is busy, users wait; if the server crashes, the site is down; if you want to deploy a new version, you have to take the site offline to do it. The obvious fix is to run several identical servers instead of one, so that no single machine is a bottleneck or a single point of failure. But now you have a new problem — which of those servers should a given user's request go to? You cannot hand out five IP addresses and hope users pick a working one.
A load balancer is the answer to that second problem. It is a service that sits in front of your servers, holds a single stable address that users connect to, and decides which of the servers behind it should handle each incoming request. Your users only ever know about the load balancer's address; the servers behind it can come and go, be replaced, or be scaled up and down without anyone outside noticing. That is the core value: a stable front door in front of a changing set of workers.
Listeners, rules, and target groups
Three pieces of vocabulary describe how a load balancer makes its decisions, and today's material uses all three constantly. A listener is the part that waits for incoming traffic on a particular port and protocol — for example, "listen for HTTP requests on port 80." A listener is the front door of the load balancer.
Behind that door, the load balancer needs to know where to send each request. A rule is a condition-and-action pair: "if the request path starts with /api, send it to the API servers; otherwise send it to the website servers." Rules are what let one load balancer serve several different applications or route different kinds of traffic to different places. When several rules could match, they are checked in a defined order and the first match wins — a detail that matters a lot later today.
A target group is simply a named collection of the servers (or other compute units) that a rule can send traffic to. Instead of listing individual servers in every rule, you point the rule at a target group and manage the membership of that group separately. This separation is the reason you can add or remove servers, or shift traffic between two versions of an application, without editing the routing rules at all.
Health checks
If a load balancer blindly forwards requests to every server in a target group, it will keep sending traffic to a server that has crashed. A health check is a periodic test the load balancer runs against each target — typically an HTTP request to a specific path, like /health, expecting a successful response. If a target fails the check enough times in a row, the load balancer stops sending it traffic until it starts passing again. This is what makes a fleet of servers self-healing from the user's perspective: a broken server quietly drops out of rotation instead of serving errors. Today's material spends real time on health checks because a badly configured one can take a perfectly healthy application completely offline.
Layer 4 versus layer 7
Network traffic is often described in terms of "layers," and the distinction between layer 4 and layer 7 is the single most important one for choosing a load balancer. Layer 4 is the level of connections and ports: a layer 4 load balancer can see that a connection arrived on port 443 from a particular IP address, but it cannot see what is inside that connection. Layer 7 is the level of the application protocol — for HTTP, that means it can read the URL path, the hostname, and the request headers. A layer 7 load balancer can therefore make routing decisions based on what the request is asking for, while a layer 4 load balancer can only decide based on where the connection came from and what port it targeted. Today's service, the Application Load Balancer, is a layer 7 load balancer, and nearly every capability discussed below flows from that fact.
The network context: subnets, Availability Zones, and security groups
Two more pieces of background show up in today's lab and gotchas. First, AWS runs its infrastructure in geographic regions, and each region is divided into isolated data-center groups called Availability Zones (AZs). A subnet is a slice of a region's network that lives in exactly one AZ. When a service is described as "highly available," it usually means its components are spread across at least two AZs so that the failure of one data-center group does not take the whole thing down. A load balancer is no exception: you choose which AZs it spans, and choosing only one makes the load balancer itself a single point of failure.
Second, a security group is a firewall attached to a resource that controls which network traffic is allowed in and out. Crucially, the load balancer has its own security group and the servers behind it have theirs, and both must permit the traffic in question. A very common real-world mistake — and a recurring exam trap — is a server whose security group allows traffic from users but not from the load balancer, which produces errors even though the application itself is fine.
With that grounding, here's why the Application Load Balancer exists, what problem it actually solves, and how its routing decisions are made.
1. Why This Is on the Exam
The Application Load Balancer sits at the intersection of three SAP-C02 concerns that are otherwise handled by separate services: traffic distribution, deployment safety, and TLS termination. That overlap is exactly why it appears so often. A question about blue/green deployment could be answered with CodeDeploy, with Route 53 weighted records, or with ALB weighted target groups, and the exam expects you to know which of those three is appropriate given the constraint stated in the scenario. A question about routing /api to one service and /static to another could be answered with a reverse proxy on EC2, with CloudFront behaviors, or with ALB listener rules — and again the discriminator is usually operational overhead rather than raw capability.
The architectural problem ALB solves is that a fleet of interchangeable compute units needs a stable, highly available entry point that can make routing decisions at layer 7. A Network Load Balancer can distribute connections but cannot inspect an HTTP path or a header, so it cannot split traffic by URL. A Classic Load Balancer can do rudimentary layer-7 work but lacks target groups, host-based rules, and native Lambda targets. ALB exists in the gap: it terminates HTTP and HTTPS, evaluates a prioritized list of rules against each request, and forwards to a named target group whose membership and health are managed independently of the listener configuration. That separation — listener rules on one side, target group membership on the other — is the design decision that makes canary releases and blue/green cutovers possible without touching DNS.
On the exam blueprint this maps most directly to the resilient architectures domain, because the failure-handling story (health checks, cross-zone balancing, connection draining) is what keeps a multi-AZ application available. It also shows up in the migration and modernization domain whenever a scenario describes decomposing a monolith into path-routed services, and in the cost domain when the correct answer is to consolidate several small load balancers behind one ALB with multiple listener rules. Expect at least one question where the deciding factor is that ALB is the only load balancer type that supports a Lambda target or a weighted split within a single rule.
2. How a Request Actually Flows Through an ALB
An ALB is not a single appliance; it is a set of load balancer nodes that AWS provisions across the Availability Zones you enable, with at least one node per enabled zone. When you create an internet-facing ALB you get a DNS name, and that name resolves to the IP addresses of those nodes. Because the node set changes as AWS scales the load balancer, AWS recommends that you never resolve the ALB DNS name to IPs and pin them — always reference the DNS name, or put a Route 53 alias record in front of it. This is the first mechanism detail worth internalizing: the ALB's own availability is a function of how many zones you enabled, and enabling only one zone makes the load balancer itself a single point of failure regardless of how many targets you registered.
When a request arrives, the listener evaluates its rules in priority order, lowest number first, and stops at the first match. Each rule has a condition (host header, path pattern, HTTP header, query string, source IP, or HTTP method) and an action. The action is either a forward to one or more target groups with optional weights, a redirect, a fixed response, or an authenticate action against a Cognito or OIDC identity provider. If no rule matches, the listener's default action applies. This ordering matters operationally: a broad rule with a low priority number will shadow every more specific rule behind it, and that is one of the most common misconfigurations in production ALBs.
Once a target group is selected, the ALB picks a target using the group's routing algorithm. The default is round robin, and the alternative is least outstanding requests, which biases toward targets with fewer in-flight requests and tends to produce better tail latency when request durations vary widely. The ALB then opens or reuses a connection to the target on the port and protocol the target group specifies, and the target group's health check determines whether that target was eligible in the first place. Health checks run on their own interval and threshold settings, independent of request traffic, which means a target can be receiving requests and still be marked unhealthy by the next check — the two systems are related but not synchronized.
Two protocol details are worth holding onto. First, ALB supports HTTP/1.1, HTTP/2, and gRPC, and it can translate between them: a client speaking HTTP/2 to the ALB can be forwarded to a target speaking HTTP/1.1. Second, when you enable a target group for a Lambda function, the ALB invokes the function with an event payload describing the request and expects a structured response, which means the Lambda target path has no persistent connection and no health check in the traditional sense. That difference is why Lambda targets behave differently under load and why they are usually the wrong answer when a scenario emphasizes steady high-throughput connection reuse.
3. The Core Decision Boundary: Rule Conditions vs. Target Group Weights
Almost every ALB scenario question reduces to a single fork: is the routing decision based on a property of the request, or on a proportion you want to control? If the answer is a property of the request — the hostname, the path, a header, a query parameter — the mechanism is a listener rule condition, and the rule forwards to exactly one target group. If the answer is a proportion — send 10% of traffic to the new version, or split evenly between two identical services — the mechanism is a single rule that forwards to two target groups with weights. Confusing these two produces designs that technically work but are operationally wrong, such as creating two listener rules with identical conditions and hoping the ALB splits between them, which it will not do; it will match the first rule and ignore the second.
The second-order decision is where the split should live at all. Weighted target groups keep the split inside the load balancer, which means it is instant, requires no DNS propagation, and can be changed by an API call. Route 53 weighted records put the split in DNS, which means clients cache the answer for the record's TTL and the effective split drifts from the configured one until caches expire. For a canary that you intend to ramp from 1% to 100% over an hour, the ALB approach is strictly better. For a split between two entirely separate stacks in different regions, DNS is the only option because a single ALB cannot span regions.
| Requirement in the scenario | Mechanism | Why not the alternative |
|---|---|---|
Route /api/* to service A, everything else to service B | Two listener rules with path conditions | Weights cannot express a path predicate |
| Send 5% of all traffic to a new version | One rule, two target groups, weighted 95/5 | Two identical rules never both match |
Route by customer hostname (a.example.com vs b.example.com) | Host-header conditions on separate rules | Weights are blind to the request |
| Shift traffic between two regions | Route 53 weighted or latency records | An ALB is regional and cannot span regions |
| Serve a maintenance page without a backend | Fixed-response action on a rule | Forwarding requires a healthy target |
| Require login before reaching the app | Authenticate action (Cognito or OIDC) | Application-level auth adds code to every service |
The exam phrasing that signals the weighted-target-group answer is usually a stated desire to test a new version with a small, controllable, quickly reversible share of production traffic. The phrasing that signals a listener rule is usually a stated difference in the request itself — a path prefix, a subdomain, a mobile client sending a distinguishing header. When a scenario mentions both, the correct design is normally a listener rule that selects the target group pair, with weights applied inside that rule, so that only the canary's slice of the path-routed traffic is split.
4. Configuration Modes and Their Tradeoffs
The knobs on an ALB fall into four groups: scheme and zones, listener and rule configuration, target group settings, and attributes. The scheme decision is made once at creation and cannot be changed — an internet-facing ALB has public IPs on its nodes and an internal ALB does not. This is a genuine constraint rather than a preference, and it is why scenarios that describe a service that must be reachable only from inside a VPC and also from an on-premises network over Direct Connect point at an internal ALB with a private IP, not an internet-facing one with security group restrictions. The zone decision is also effectively permanent in the sense that enabling more zones later is easy but running with one zone is a design error you cannot patch around.
Listener configuration is where most of the operational flexibility lives. Each listener is bound to a port and protocol, and HTTPS listeners require an ACM certificate or an imported one. Rules within a listener are evaluated by priority, and the default action is the fallback. The tradeoff here is between rule count and rule clarity: you can express a great deal with a handful of well-ordered rules, but every additional rule is another chance to shadow a later one. A useful discipline is to reserve low priority numbers for the most specific conditions and let the default action handle the long tail, rather than writing an exhaustive set of rules that must all be kept in the right order.
Target group settings carry the health check configuration, the routing algorithm, the deregistration delay, and the stickiness settings. Deregistration delay is the one that most often surprises people: when a target is deregistered or fails a health check, the ALB stops sending new requests but keeps existing connections open for the configured delay, defaulting to 300 seconds, so that in-flight requests can finish. Setting it too high makes deployments slow because every removed target lingers; setting it too low causes 5xx errors during scale-in because requests are cut off mid-flight. This is the same drain concept that yesterday's Terminating:Wait lifecycle hook implements, and the two must be tuned together or one will undo the other.
Attributes are the least visible group and the most likely to be the hidden answer in a scenario. Cross-zone load balancing is on by default for ALBs and distributes requests evenly across all targets in all enabled zones rather than evenly across zones. Slow start gives a newly registered target a ramp-up period during which it receives a linearly increasing share of requests, which matters for JVM-based services that need to warm caches. Sticky sessions bind a client to a target using a cookie, which is sometimes necessary for stateful applications but undermines even distribution and complicates scale-in. Idle timeout controls how long an idle connection is held, and it must be shorter than the keep-alive timeout of the backend or the backend will close connections the ALB still believes are open.
5. Sizing, Limits and Quotas
ALB capacity is not something you provision. There is no instance size and no capacity unit to buy; AWS scales the load balancer's nodes in response to traffic, and the practical limits are expressed as quotas on configuration objects and as soft limits on new connections and requests per second that can be raised through a support request. This is a meaningful difference from an NLB, where you can attach Elastic IPs and reason about static addresses, and from a Gateway Load Balancer, which is built for a different purpose entirely. For exam purposes the important consequence is that "the load balancer is too small" is almost never the correct diagnosis; the bottleneck is nearly always the targets, the health check configuration, or a quota on a configuration object.
The quotas that actually constrain designs are the ones on rules and target groups. A listener supports a bounded number of rules, and each rule supports a bounded number of conditions and a bounded number of target groups in a forward action. Target groups have their own limits on registered targets, and each target group can be associated with a limited number of load balancers. Certificates are limited per load balancer, which matters when a scenario describes hosting hundreds of customer domains — the answer there is usually a wildcard certificate or SNI with a smaller set of certificates rather than one certificate per domain. These numbers change over time, so the exam-relevant skill is knowing which quota to check, not memorizing the current value.
| Dimension | What to know | Design consequence |
|---|---|---|
| Capacity model | Managed and auto-scaled; no instance sizing | Do not answer "resize the ALB" |
| Enabled zones | At least one node per enabled zone | One zone makes the ALB itself a SPOF |
| Rules per listener | Bounded, raisable by support request | Very large routing tables may need CloudFront or a different pattern |
| Targets per target group | Bounded, raisable | Huge fleets may need multiple target groups |
| Certificates per ALB | Bounded; SNI selects by hostname | Many custom domains need wildcards or a certificate strategy |
| Deregistration delay | Configurable, default 300 seconds | Directly sets deployment and scale-in duration |
| Idle timeout | Configurable, default 60 seconds | Must be shorter than backend keep-alive |
Two sizing-adjacent behaviors are worth separating from quotas. First, the ALB's health check interval and thresholds determine how quickly a failing target is removed, and that detection latency is a real availability cost during an incident — a target that fails immediately after a check will keep receiving traffic until the next check plus the unhealthy threshold. Second, connection reuse means the ALB is not opening a new backend connection per request, so backend connection counts are governed by concurrency and keep-alive settings rather than by request rate. A scenario that describes exhausting a database's connection limit through an ALB is usually describing a target-side pooling problem, not a load balancer limit.
6. Failure Modes and What They Look Like in Production
The most common ALB failure signature is a burst of HTTP 502 responses. A 502 from an ALB means the load balancer could not get a valid response from the target, and the usual causes are a target that crashed mid-request, a target whose security group no longer permits traffic from the ALB, a target listening on a different port than the target group expects, or a target that is being deregistered while still receiving requests because the deregistration delay is too short. The first diagnostic move is to check target health in the target group and then check the target's own logs for the request that failed — the ALB access logs will show the request reaching the load balancer, and the absence of a corresponding application log entry localizes the problem to the network path or the process.
A 503 from an ALB almost always means there are no healthy targets. That is a different failure than 502 and points at the health check rather than the application. The classic version of this is a health check path that returns a redirect or requires authentication, so every target is marked unhealthy even though the application is serving real traffic fine. The second classic version is a health check that is too aggressive relative to the application's startup time, so targets are marked unhealthy during boot and never get the chance to become healthy. Both are configuration errors that produce a total outage from a working application, which is why the health check path deserves the same review as the application code.
Timeouts present as 504 responses and are the hardest of the three to diagnose because they can originate on either side. The ALB's idle timeout governs how long it will wait on an idle connection; if the backend takes longer than that to produce a response, the ALB closes the connection and the client sees a 504. The mirror-image problem is a backend keep-alive timeout shorter than the ALB's idle timeout, which causes the backend to close connections the ALB still considers usable, producing intermittent 502s that are maddening to reproduce. The standard fix is to set the ALB idle timeout slightly longer than the backend's keep-alive timeout so the backend is always the side that closes first.
Two subtler failure modes are worth naming. Uneven distribution across targets usually means sticky sessions are enabled, or that one target is in a different zone and cross-zone balancing is disabled, or that the routing algorithm is least-outstanding-requests and one target is genuinely slower. And a canary that appears to receive no traffic at all is usually a weight configuration error rather than a routing error — a weight of zero on the new target group, or a second listener rule with a lower priority number that matches first and forwards everything to the old target group. Both are configuration mistakes that look like application bugs from the outside.
7. The Operational and SRE Angle
From an SRE perspective the ALB is the single best place to measure the user-visible health of a service, because it sees every request before any application code does. The metrics that matter most are the target response time, the request count split by target group, the HTTP code counts broken out by 2xx, 3xx, 4xx, and 5xx, and the unhealthy host count. The 5xx count is the closest thing to a direct SLO signal available without instrumenting the application, and an alarm on the 5xx rate as a proportion of total requests is a reasonable default for any service behind an ALB. The unhealthy host count is the leading indicator: it moves before user-visible errors do, because the ALB removes targets from rotation before the error rate climbs.
Access logs are the other half of the observability story. ALB access logs record the request time, the target that handled it, the target and load balancer response codes separately, and the time spent in each phase of the request. That separation is what makes them more useful than application logs for latency work: you can tell whether time was spent waiting for the target to respond or in the connection setup, which distinguishes an application problem from a networking problem. Enabling access logs to S3 and querying them with Athena is a common pattern, and it is the mechanism behind most "which endpoint is slow" investigations.
The runbook shape for an ALB-fronted service has a predictable structure. First, check target health and the unhealthy host count to determine whether the problem is membership or performance. Second, check the 5xx breakdown to distinguish 502 (target unreachable or crashed) from 503 (no healthy targets) from 504 (timeout). Third, check whether a recent deployment changed the target group membership or the listener rules, because a rule ordering mistake produces symptoms that look exactly like an application routing bug. Fourth, check the deregistration delay and idle timeout against the backend's keep-alive settings, since a mismatch there produces intermittent errors that correlate with deployments and scale events rather than with load.
For deployment safety, the SRE-relevant property of weighted target groups is that the weight is an API-controlled number, which means a canary can be automated: shift 5%, watch the 5xx rate and target response time for a defined bake period, and either continue the ramp or set the weight back to zero. That rollback is a single API call with no DNS propagation delay and no client-side caching, which is why it is a better rollback mechanism than a DNS-based split. The corresponding discipline is to alarm on the canary target group's error rate separately from the stable group's, because a 5% canary with a 20% error rate only moves the aggregate error rate by one percentage point — invisible in the aggregate, obvious when measured per target group.
8. Edge Cases and Exam Gotchas
The single most-tested gotcha is that two listener rules with identical conditions do not split traffic. The ALB evaluates rules in priority order and stops at the first match, so the second rule is dead configuration. Any scenario that describes a percentage split must be answered with weights inside one rule, and any scenario that describes two rules with the same condition is describing a bug. A closely related trap is rule ordering: a rule matching /* with a low priority number will shadow a rule matching /api/* with a higher number, and the symptom is that the specific route appears to be ignored.
Security groups are the second recurring trap. An ALB has its own security group, and the targets have theirs; the target's security group must allow traffic from the ALB's security group on the target port. A scenario that describes targets that are healthy in the console but return 502s is often describing a security group that permits the client but not the load balancer. The related detail is that an internal ALB's nodes have private IPs from the subnet CIDR, so a target security group rule referencing the ALB's security group is more robust than one referencing a CIDR range that may change.
Health check configuration produces a cluster of gotchas. The health check port can differ from the traffic port, which is useful for a dedicated health endpoint but is also a common source of confusion. The health check protocol can be HTTP even when the traffic protocol is HTTPS, which is fine internally but must be deliberate. A health check that returns a 3xx is treated as unhealthy unless you configure the matcher to accept it, so an application that redirects unauthenticated requests to a login page will fail its health check by default. And a health check path that hits a database-backed endpoint turns a database slowdown into a total outage, because every target fails its check at once — the standard mitigation is a shallow health check that verifies the process is alive and a separate deep check that is not used for load balancer membership.
The remaining gotchas are about scope. An ALB is regional, so any scenario requiring a single endpoint across regions is describing Global Accelerator or Route 53, not an ALB. An ALB cannot do TCP or UDP passthrough, so a scenario about a non-HTTP protocol is describing an NLB. Lambda targets have no health check and no connection reuse, so a scenario emphasizing sustained throughput or long-lived connections is not a Lambda target scenario. And an ALB does not terminate traffic for a private service consumed across accounts — that is PrivateLink, which presents an NLB-backed endpoint service rather than an ALB.
9. ALB vs. NLB vs. CLB vs. CloudFront
The load balancer family splits along the layer at which the decision is made. An ALB makes decisions at layer 7 and therefore understands HTTP semantics; an NLB makes decisions at layer 4 and therefore understands connections and ports but not paths. A Classic Load Balancer is the legacy option that predates target groups and should not appear as the correct answer in a new design. CloudFront is not a load balancer at all but a CDN whose behaviors can perform some of the same routing work at the edge, which is why it sometimes appears as a distractor in routing questions.
| Service | Decision layer | Pick it when… |
|---|---|---|
| Application Load Balancer | Layer 7 (HTTP/HTTPS, gRPC) | You need path, host, header, or query routing; weighted splits; redirects; or Lambda targets |
| Network Load Balancer | Layer 4 (TCP/UDP/TLS) | You need extreme throughput, static IPs, or non-HTTP protocols |
| Classic Load Balancer | Layer 4 and limited layer 7 | Legacy only; not the answer for a new design |
| CloudFront | Edge, layer 7 | You need global caching, edge TLS, or WAF at the edge in front of an ALB |
| API Gateway | Layer 7, API semantics | You need API keys, usage plans, request validation, or native AWS service integrations |
| Global Accelerator | Network, anycast | You need a static global IP with fast regional failover |
The practical rule is to start from the protocol and the predicate. If the protocol is HTTP and the routing predicate is something in the request, the answer is an ALB. If the protocol is not HTTP, or the requirement is a static IP address, the answer is an NLB. If the requirement is a global entry point with caching, the answer is CloudFront in front of an ALB, and the two are complementary rather than alternatives — CloudFront terminates the client connection at the edge and forwards to the ALB, which then routes to targets. If the requirement is API management rather than traffic distribution, the answer is API Gateway, and the fact that API Gateway can also front a Lambda does not make it a substitute for an ALB in front of a container fleet.
One comparison that comes up specifically in migration scenarios is ALB versus a self-managed reverse proxy such as NGINX on EC2. The functional overlap is real, but the operational difference is that the ALB is a managed, multi-AZ, auto-scaled service with no patching, no capacity planning, and native integration with ACM, WAF, and Auto Scaling group lifecycle. A scenario that emphasizes reducing operational overhead while keeping path-based routing is pointing at the ALB, and a scenario that requires a custom module or a protocol the ALB does not support is the rare case where the self-managed proxy is correct.
Hands-on Lab: A Weighted Canary Behind One Listener Rule
The goal is to stand up a single ALB that serves two versions of the same service and lets you shift traffic between them by changing one number, then verify the shift with per-target-group metrics. Work in a sandbox account and tear everything down at the end.
1. Create the network and the two target groups. In a VPC with at least two public subnets in different Availability Zones, create two target groups of type ip or instance — call them tg-blue and tg-green. Give both the same health check path, for example /health, and the same port. Register at least two targets in each group so that a single target failure does not take a whole version out of rotation. Confirm both groups report all targets healthy before continuing; a canary built on an unhealthy target group produces 502s that look like a code problem.
2. Create the ALB and one listener. Create an internet-facing ALB across the same two subnets and attach a security group that permits inbound HTTP from your IP. Create an HTTP listener on port 80 with a default action that forwards to tg-blue with weight 100. Verify that requests to the ALB DNS name reach the blue version and that the target group shows healthy targets receiving traffic.
3. Add the weighted rule. Add a listener rule with a condition that matches the traffic you want to canary — for a full canary, a path pattern of /*; for a scoped canary, a path prefix such as /checkout/*. Set the action to forward to two target groups: tg-blue with weight 90 and tg-green with weight 10. Give the rule a priority number lower than any broader rule so it is evaluated first. This is the step where the exam gotcha lives: the split is expressed as weights inside one rule, not as two rules with the same condition.
4. Verify the split empirically. Send a few hundred requests to the ALB in a loop and count how many responses identify as the green version. Expect roughly the configured proportion, with variance that shrinks as the sample grows. Then check the ALB's per-target-group request count metric in CloudWatch and confirm both groups are receiving traffic in approximately the configured ratio. If green receives nothing, check the weight value and the rule priority before suspecting the application.
5. Exercise the rollback path. Change the weights to 100/0 and confirm that green stops receiving traffic within seconds, with no DNS change and no client-side caching involved. Then change them to 50/50 and confirm the shift is equally immediate. This is the property that makes weighted target groups a better canary mechanism than a DNS split, and it is worth feeling directly rather than taking on faith.
6. Break it deliberately. Stop the application on one green target and watch the target group mark it unhealthy while the blue group continues serving. Then set the green target group's health check path to something that returns a 404 and observe the whole green group go unhealthy — the ALB will still route the configured weight to it, and the result is 502s for that share of traffic. This is the failure mode to recognize in production: a canary that is configured correctly but unhealthy produces errors proportional to its weight.
7. Tune the drain and timeout settings. Set the target group's deregistration delay to a short value such as 30 seconds, deregister one target, and observe how quickly it leaves rotation. Then compare that against the ALB idle timeout and the backend's keep-alive timeout, and set the idle timeout slightly higher than the backend's keep-alive so the backend is always the side that closes first. Record the three values in your notes; the relationship between them is a recurring exam theme.
8. Clean up. Delete the listener rule, the listener, the ALB, both target groups, and any instances or tasks you launched. Confirm in the console that no load balancer remains, since an idle ALB still bills hourly.
Scenario Question Drills
Q1. A team wants to release a new version of a service to 5% of production traffic and be able to revert within seconds if error rates rise. The service runs behind an Application Load Balancer. What should they configure?
Q2. An ALB listener has a rule with priority 10 matching /* and forwarding to the default target group, and a rule with priority 20 matching /api/* forwarding to the API target group. Requests to /api/orders are reaching the default target group. Why?
Q3. Users report intermittent 502 errors from an ALB-fronted service. The application logs show no corresponding errors, and the targets are healthy in the console. What is the most likely cause?
Q4. A service must be reachable from an on-premises network over Direct Connect and from other VPCs, but must never be reachable from the public internet. Which ALB configuration is correct?
Q5. Every target in a target group is marked unhealthy, but the application responds correctly when tested directly from an instance in the same subnet. The health check path is /dashboard, which redirects unauthenticated requests to a login page. What is the fix?
Q6. A team needs to serve a static maintenance page during a database migration without running any backend. What ALB feature should they use?
Q7. A canary target group is configured with a weight of 10, but CloudWatch shows it receiving zero requests. The targets are healthy. What should be checked first?
Q8. A service behind an ALB must accept gRPC traffic from clients while the backend containers speak HTTP/1.1. Is this supported?
Q9. During a scale-in event, users see a brief spike of 5xx errors as instances are removed. The application handles requests in under two seconds. What is the most likely cause?
Q10. A company hosts 200 customer-specific domains on a single ALB and needs TLS for each. What is the recommended approach?
Q11. A team wants a newly launched instance to receive a gradually increasing share of traffic rather than the full round-robin share immediately, because the application needs to warm its cache. What ALB feature addresses this?
Q12. A scenario requires a single global entry point with a static IP address that fails over between two regions in seconds, in front of ALBs in each region. What should be used?
Q13. An ALB forwards to a Lambda target. Which statement about this configuration is accurate?
Q14. A team wants to alarm on the user-visible error rate of a service behind an ALB before application-level instrumentation is in place. Which metric is the most direct signal?
Q15. A canary is running at 5% weight. The canary version has a 20% error rate while the stable version is healthy. The aggregate ALB error rate has barely moved. What should the team do to detect this reliably?
Peek into Tomorrow
Everything in today's design assumes the backend is a long-running process that the load balancer can hold a connection to and health-check on an interval. That assumption is what makes weighted target groups, deregistration delay, and slow start meaningful — they all describe the lifecycle of a persistent target. Tomorrow's material breaks that assumption. When the compute unit is a Lambda function, there is no persistent connection to drain, no instance to warm with a slow start, and no health check that can mark a function unhealthy before it is invoked. The load balancer's job of deciding which target receives a request is replaced by a different question: how many concurrent invocations should be allowed to exist at once, and what happens to the requests that arrive when that ceiling is reached.
That question has two answers that sound similar and behave very differently. One caps concurrency and makes excess invocations wait or fail; the other pre-warms execution environments so that a burst of traffic does not pay the initialization cost that a cold start imposes. Choosing between them is the same kind of decision as choosing between a weighted canary and a full cutover — it depends on whether the constraint is a downstream system that must be protected or a latency budget that must be met. The other unresolved thread is what happens when the event source is a stream rather than an HTTP request, because a stream delivers batches and a batch can fail partially, which is a failure mode with no analogue in the ALB world we just finished.
Sources
- What is an Application Load Balancer? — AWS Elastic Load Balancing User Guide
- Listeners for your Application Load Balancer — AWS Elastic Load Balancing User Guide
- Target groups for your Application Load Balancers — AWS Elastic Load Balancing User Guide
- Target group attributes — AWS Elastic Load Balancing User Guide
- Quotas for your Application Load Balancers — AWS Elastic Load Balancing User Guide
- Access logs for your Application Load Balancer — AWS Elastic Load Balancing User Guide
- Lambda targets for your Application Load Balancer — AWS Elastic Load Balancing User Guide
- CloudWatch metrics for your Application Load Balancer — AWS Elastic Load Balancing User Guide
- Blue/Green Deployments on AWS — AWS Whitepaper
- Reliability Pillar — AWS Well-Architected Framework