Day 38 of 70 · Week 6
Day 38 / 70 Week 6 of 14 Phase 3: SRE Observability, Resilience & DR

Route 53 Routing Policies & Health Checks

🕑 ~58 min read · 2 services covered
Route 53 Latency/Weighted/Failover/Geoproximity Routing

Recap: From Deterministic Failover to Everyday Traffic Steering

Route 53 Application Recovery Controller answered a specific question: when a region is genuinely lost, how do you flip traffic in a way you can prove was correct? Its readiness checks continuously validate that a standby region actually has the capacity and configuration to absorb production load, and its routing controls live on a control plane independent of any single region, so the failover mechanism itself does not fail with the thing it is failing over from. That combination is what makes ARC failover deterministic and auditable rather than a hopeful side effect of health checks.

Today sits on the same service family and inverts the failure mode. ARC is about the rare, catastrophic, human-authorized event: a region is gone, an operator or automation makes a deliberate call, and the audit trail matters as much as the outcome. Route 53 routing policies and health checks are about the continuous, unglamorous, automatic case — a single endpoint degrades, a canary release starts misbehaving, one Availability Zone's targets stop answering, and DNS quietly stops handing out that address without anyone being paged. Same service, opposite failure mode: ARC handles the failure you plan for, health checks handle the failure you hope never needs a human.

Foundations You'll Need Today

Today's material is about the system that turns a human-readable name like www.example.com into the numeric address a computer actually connects to, and about how that translation can be made to change automatically when something breaks. Before the routing policies make sense, five pieces of background are worth having straight.

DNS: the internet's phone book, and who is allowed to answer

Computers find each other by numeric address, but people remember names. DNS is the global system that translates one into the other. It is not a single database; it is a hierarchy of servers, and the important distinction for today is between a resolver and an authoritative server. A resolver is the server your laptop or your company network asks first — it does the legwork of chasing down an answer and then remembers (caches) it for a while. An authoritative server is the one that actually owns the answer for a given domain and is the final word on it. Route 53 is an authoritative DNS service: you tell it what the correct answers are for your domain, and it hands those answers to resolvers when they ask. That is why a routing policy is described as a rule evaluated at query time — every time a resolver asks, Route 53 decides what to say, rather than having a fixed answer written down once.

Records, record sets, and the alias-versus-CNAME distinction

A DNS record is a single entry that maps a name to something — most commonly a name to an IP address. A record set is the group of records that share the same name and the same type, which matters because routing policies operate on the whole set: a weighted configuration is several records with the same name, each carrying a different weight. Two record types come up constantly today. A CNAME is a record that says "this name is really just another name, go look that one up instead" — useful for pointing one hostname at another, but the DNS standard forbids using one at the very top of a domain (the zone apex, the bare example.com with nothing in front of it). An alias record is Route 53's own answer to that limitation: it points at an AWS resource such as a load balancer or a CloudFront distribution, works at the zone apex where a CNAME cannot, and is free to query. Most production designs use alias records for the final answer and CNAMEs only when they must reference a hostname outside AWS.

TTL and caching: why DNS changes are never instant

Every DNS answer carries a time to live, or TTL — a number of seconds that resolvers and clients are allowed to keep using that answer before they must ask again. This is the single most important number in today's material, because it is the reason a failover is never instantaneous. If Route 53 stops handing out an address but the answer is still cached with 300 seconds left on its TTL, clients keep connecting to the old address for up to five minutes. The practical consequence is that the real failover window is the time it takes to detect the problem plus the TTL, and lowering the TTL is the standard lever for shortening it, at the cost of more DNS queries.

Health checks: automated "is this thing actually working?" probes

A health check is a small, separate piece of configuration that asks a specific question on a schedule — typically an HTTP request to a URL, or a raw TCP connection — and records whether the answer looked healthy. The two settings that matter are the interval (how often it asks) and the failure threshold (how many consecutive bad answers before it declares the endpoint unhealthy). Requiring several consecutive failures is deliberate: it stops a single dropped packet or a momentary hiccup from flapping your DNS answers back and forth. The reason health checks matter so much today is that they are what turns a static routing rule into a self-healing one — a record whose health check is failing is simply left out of the answer, with no deployment and no human involved.

Regions, Availability Zones, and load balancers

AWS runs its services in geographic regions, and each region contains several isolated data centers called Availability Zones. Designing for resilience usually means spreading across AZs within a region, and sometimes across regions entirely, which is why today's policies are described as steering traffic "across regions or endpoints." A load balancer is the component that sits in front of a group of servers and distributes incoming requests among them, checking each server's health and quietly removing any that stop responding. The Application Load Balancer (ALB) works at the level of individual HTTP requests and can route by URL path or hostname; the Network Load Balancer (NLB) works at the level of raw connections and is chosen for extreme throughput or for a fixed IP address. Today's material sits one layer above all of this: Route 53 decides which region or which load balancer a user should reach in the first place, and the load balancer then decides which server behind it handles the request.

With that grounding — DNS as a query-time decision, records and record sets, TTL as the reason failover is never instant, health checks as the automated signal, and load balancers as the layer just below — here is what Route 53's routing policies actually do and which one a given scenario forces you into.

1. Why This Is on the Exam

Route 53 is the only service in the AWS portfolio that sits in front of essentially every other service, which means it appears in scenario questions that are nominally about something else. A question about a multi-region active-active application, a blue/green deployment, a canary release, a hybrid DNS design, or a disaster recovery runbook will almost always have a routing policy decision buried in it, and that decision is frequently the difference between the two most plausible answer choices. The exam does not ask you to recite the list of policies; it asks you to recognize which policy the stated constraints force you into, and then to notice whether the answer choice also handles the health-check behavior the scenario implies.

This maps most directly to Domain 2, Design Resilient Architectures, because routing policy plus health checks is the mechanism by which a multi-AZ or multi-region design actually fails over. It also reaches into Domain 1 through hybrid DNS and private hosted zones, and into Domain 4 whenever a scenario asks you to shift a small percentage of traffic to a new version and roll back cheaply if it goes wrong. The recurring trap is that candidates learn the policy names as a vocabulary list and then pick the one whose name sounds closest to the scenario's adjective — "lowest latency" maps to latency-based, "closest" maps to geoproximity — without checking whether the scenario also requires failover, which is a separate configuration on top of the policy, not a property of it.

There is a second reason this day carries weight. Health checks are the part of Route 53 that most candidates under-specify. A routing policy without a health check is a static answer; a routing policy with a health check is a self-healing system. Exam scenarios that describe an outage and ask what would have prevented it are usually testing whether you attached health checks to the records, whether you pointed them at the right thing (an endpoint, a calculated set of other checks, or a CloudWatch alarm), and whether you understood that a health check failing does not remove a record from a simple routing policy at all.

2. How Routing and Health Checking Actually Work

Route 53 is an authoritative DNS service, and the mental model that makes everything else fall into place is that a routing policy is a rule evaluated at query time, not a static zone file. When a resolver asks for a name, Route 53 evaluates the record set for that name, applies the policy, filters out any records whose associated health check is failing, and returns the surviving answer or answers. That evaluation happens per query, which is why weighted routing can shift percentages gradually and why a failing health check can take an endpoint out of rotation without any deployment. The client never learns that a policy exists; it just receives an answer, and the answer changes as conditions change.

Health checks are separate resources that Route 53 evaluates on its own schedule, independent of incoming queries. A health check can monitor an endpoint directly over HTTP, HTTPS, or TCP, or it can be a calculated health check that combines the status of several child checks with boolean logic, or it can be tied to a CloudWatch alarm so that any metric you can alarm on becomes a routing signal. The endpoint health checker originates requests from a set of Route 53 health-checker IP ranges distributed globally, and it considers an endpoint healthy only when a threshold number of consecutive checks succeed. That threshold behavior is what prevents a single dropped packet from flapping your DNS answers.

The interaction between the two is where most of the exam's nuance lives. Health checks attach to individual records within a record set, so a weighted record set with three records can have three different health checks, and a failing one is simply excluded from the weighted distribution — the remaining weights are renormalized across the healthy records. For failover routing, the primary and secondary records each carry a health check, and the secondary is only returned when the primary's check is failing. For simple routing, there is exactly one record and no health-check-driven removal, which is the single most commonly missed fact in this topic: you cannot make a simple routing record fail over, because there is nothing to fail over to.

Alias records deserve their own note because they change what you are pointing at. An alias record points at an AWS resource — an ALB, a CloudFront distribution, an S3 website endpoint, another Route 53 record in the same zone — and Route 53 resolves it to the resource's current addresses without you managing them. Alias records are free to query, they can be used at the zone apex where a CNAME cannot, and they inherit the health of the target when the target is an AWS resource that Route 53 can evaluate. Most production designs use alias records for the leaf answers and CNAMEs only where a non-AWS hostname must be referenced.

3. The Core Decision Boundary: Which Policy Does the Constraint Force?

Every routing question reduces to one fork: is the scenario asking you to distribute traffic by a measurable property of the request, or to select a single answer based on a condition? Distribution policies — weighted, latency, geolocation, geoproximity, multi-value — return one or more answers chosen by a rule and are used when you want traffic spread across endpoints. Selection policies — failover, and simple in its degenerate case — return a specific answer and are used when you want a deterministic choice between a primary and a standby. Getting this fork right eliminates most wrong answers immediately, because a scenario that says "route users to the region closest to them" cannot be satisfied by failover, and a scenario that says "serve from the standby only when the primary is down" cannot be satisfied by latency routing.

The second-order question is what the policy measures. Latency-based routing measures the network latency AWS observes between the requester and each region, which is a property of the network path and not of geography — a user in one city may be routed to a region that is not the geographically nearest one because the measured path is faster. Geolocation routing measures where the DNS query originated, at the level of continent, country, or US state, and returns the answer you mapped to that location regardless of how fast the path is. Geoproximity routing measures distance from the resource to the requester and lets you bias the result, which is the only policy that lets you deliberately pull traffic toward or away from a region by a tunable amount. These three are the ones candidates conflate most, and the distinguishing question is always "what is being measured, and can I tune it?"

PolicyWhat it measuresAnswers returnedHealth-check failoverTypical use
SimpleNothingOneNoSingle endpoint, no redundancy
WeightedAssigned weight per recordOne (probabilistically)Yes, per recordCanary, A/B, gradual migration
LatencyMeasured network latency to regionOne (lowest latency)Yes, per recordMulti-region active-active
FailoverPrimary health stateOne (primary or secondary)Yes, that is the pointActive-passive DR
GeolocationQuery origin (continent/country/state)One (mapped)Yes, per recordData residency, localization
GeoproximityDistance, with tunable biasOne (nearest after bias)Yes, per recordShifting traffic share between regions
Multi-value answerNothing (returns up to eight healthy records)Up to eightYes, per recordClient-side load spreading without an ELB

Multi-value answer routing is the odd one out and is worth calling out because its name misleads. It is not a load balancer and it does not weight anything; it returns up to eight healthy records in a single response and lets the client pick, which is useful when you want crude redundancy without paying for an ELB but is not a substitute for one. If a scenario needs session affinity, connection draining, or per-target health semantics richer than "is this IP answering," multi-value is the wrong answer and an ALB behind a single alias record is the right one.

4. Configuration Modes and Their Tradeoffs

Weighted routing is the policy most often configured incorrectly, because the weights are relative rather than absolute and because the behavior when a record fails is not what people expect. Each record in a weighted set carries a number, and Route 53 returns a record with probability proportional to its weight divided by the sum of all weights in the set. Setting one record to 95 and another to 5 gives you a five percent canary, but setting them to 19 and 1 gives you exactly the same split — the numbers are ratios, not percentages. When a record's health check fails, that record is removed from the set and the remaining weights are renormalized, so a 95/5 split with the 5 failing becomes 100/0 rather than 95/5. That renormalization is usually what you want, but it means a canary that fails does not leave you serving 95 percent of traffic to a healthy primary and five percent to nothing; it leaves you serving everything to the primary, which is the correct outcome and worth stating explicitly in a runbook.

Failover routing is the simplest policy to reason about and the easiest to get subtly wrong. You designate one record as primary and one as secondary, attach a health check to the primary, and Route 53 returns the secondary only while the primary's check is failing. The subtlety is that the secondary is not health-checked in the same way — if the secondary is also unhealthy, Route 53 will still return it, because failover routing has no third option. A robust active-passive design therefore either health-checks the secondary too and accepts that a double failure returns nothing, or pairs failover routing with a separate mechanism that alerts on secondary unhealthiness. The other subtlety is TTL: failover is only as fast as your DNS TTL allows, because resolvers and clients cache the previous answer until it expires.

Latency and geoproximity routing both spread traffic across regions but differ in tunability. Latency routing is entirely automatic — you cannot tell Route 53 to prefer one region over another for a given user, only to prefer the lowest-latency healthy one. Geoproximity routing adds a bias parameter per record, expressed as a positive or negative number, which expands or shrinks the geographic area from which a region attracts traffic. That bias is the mechanism for deliberately shifting load between two healthy regions without taking either out of service, which is exactly what a gradual regional migration or a capacity rebalance looks like. If a scenario says "shift traffic gradually from us-east-1 to us-west-2 while both are healthy," geoproximity with bias is the answer; latency routing cannot express that intent.

Health-check configuration has its own tradeoff surface. The interval determines how quickly you detect failure and how much you pay, with a fast interval costing more per check. The failure threshold determines how many consecutive failures are required before the endpoint is considered unhealthy, and the success threshold determines how many consecutive successes are required before it is considered healthy again. Raising the failure threshold reduces false positives from transient blips at the cost of slower detection; raising the success threshold prevents flapping back into rotation too eagerly. For a canary deployment, a low failure threshold with a fast interval is usually right, because you want to pull a bad canary out quickly and you are not worried about a brief false positive on a five percent slice. For a primary region in an active-passive design, a slightly higher failure threshold is defensible because a spurious failover is expensive.

5. Sizing, Limits and Quotas

The numbers that matter here are mostly about how many records and health checks you can have and how fast health checking can react, and they are worth knowing because exam scenarios occasionally hinge on whether a design is even expressible. Route 53 allows a large number of records per hosted zone and a large number of hosted zones per account, with the default hosted zone quota being 500 per account and adjustable upward by request. Record sets within a single name are limited to 100 records for most routing policies, which is the practical ceiling on how finely you can slice a weighted distribution — you cannot express a thousand-way split within one record set.

Health checks have their own quotas and their own cost model. The default quota is 200 health checks per account, adjustable by request, and each health check is billed monthly with a higher rate for the fast interval and for checks that use HTTPS with SNI or that are configured with string matching. Calculated health checks and CloudWatch-alarm-based health checks count against the same quota. The fast interval is 10 seconds with a 30-second default, and the standard interval is 30 seconds; the failure threshold defaults to three consecutive failures, which means a standard-interval check takes roughly 90 seconds to declare an endpoint unhealthy before DNS even begins to change.

TTL is the number that most often determines real-world failover time and is the one candidates forget to include in their arithmetic. A record with a 300-second TTL can be cached by resolvers and clients for five minutes after Route 53 stops returning it, so the effective failover window is the health-check detection time plus the TTL, not just the detection time. Lowering TTL to 60 seconds shortens that window at the cost of more query volume and slightly higher resolution latency, and it is the standard tradeoff for records that participate in failover. Alias records to AWS resources do not have a user-settable TTL in the same way, because Route 53 resolves them to the resource's addresses, which is one more reason alias records are preferred for the leaf answers in a failover design.

SettingDefault / typical valueWhy it matters
Hosted zones per account500 (adjustable)Ceiling on how many domains you can host
Records per record set100Limits how finely a weighted split can be sliced
Health checks per account200 (adjustable)Shared across endpoint, calculated, and alarm checks
Health check interval30s standard, 10s fastSets detection latency; fast interval costs more
Failure threshold3 consecutive failures~90s to declare unhealthy at standard interval
Record TTL300s common, 60s for failoverAdds directly to effective failover time
Multi-value answersUp to 8 healthy recordsCeiling on client-side spreading without an ELB

6. Failure Modes and What They Look Like in Production

The most common production failure is a health check that is technically working but monitoring the wrong thing. A health check pointed at an ALB's DNS name over HTTPS will report healthy as long as the load balancer answers, even if every target behind it is failing, because the ALB itself is still responding — possibly with a 503 that the health check is not configured to treat as unhealthy. The symptom is a routing policy that never fails over during a real application outage, and the first diagnostic move is to look at what the health check is actually requesting and what status codes it accepts. Pointing the check at a dedicated deep-health endpoint that exercises the application's dependencies, and configuring it to accept only 200, is the standard fix.

The second common failure is a health check that is too aggressive and causes flapping. If the failure threshold is one and the interval is fast, a brief network hiccup or a garbage-collection pause can mark a healthy endpoint unhealthy, Route 53 removes it, traffic shifts, the endpoint recovers, and the check marks it healthy again — and the cycle repeats. The symptom is oscillating DNS answers and traffic that never settles, visible in the health-check status history and in uneven request distribution across targets. The fix is to raise the failure threshold, raise the success threshold so recovery is also confirmed rather than assumed, and make sure the health check's timeout is comfortably below its interval so slow responses are not counted as failures by accident.

The third failure is the one that produces the most confusing incident reports: a failover that does not happen because the record was never eligible to fail over. Simple routing records have no failover behavior, so a scenario where the only record points at a dead endpoint will keep returning that endpoint until someone changes the record. Similarly, a weighted record set where every record has failed returns nothing at all, which surfaces to users as NXDOMAIN or SERVFAIL rather than as a slow response. And a failover configuration where the secondary record has no health check will happily return a secondary that is also down. The diagnostic move in all three cases is the same: enumerate the record set, confirm each record's health-check association, and confirm that at least one record can be healthy under the current conditions.

Finally, there is the failure mode that is not a failure of Route 53 at all but of the assumption that DNS is instantaneous. A team declares a five-minute RTO, configures failover routing correctly, and then discovers that clients with long-lived connections never re-resolve and keep talking to the failed endpoint for hours. The symptom is a failover that looks successful in Route 53's console while a subset of users remains broken. The fix is architectural rather than DNS-level: keep TTLs low, ensure clients re-resolve on connection failure, and for the cases where DNS caching cannot be controlled, use a network-layer mechanism such as Global Accelerator that does not depend on the client re-resolving at all.

7. The Operational and SRE Angle

From an SRE perspective, routing policy and health checks are the outermost layer of your availability story, and they deserve the same monitoring discipline as the application itself. The signals worth watching are health-check status transitions, which Route 53 publishes as CloudWatch metrics per health check, and the distribution of answers actually being returned. A health check that has been flapping — alternating between healthy and unhealthy — is a leading indicator of an unstable endpoint, and it is visible in the status metric before it shows up as user-visible latency. Alarming on health-check status changes rather than only on the resulting error rate gives you a head start on the incident.

The SLO implication is that your effective availability is bounded by the routing layer's ability to detect and route around failure, not just by the health of individual endpoints. If your health check takes 90 seconds to detect a failure and your TTL is 300 seconds, then a single-endpoint failure costs you roughly six and a half minutes of degraded service for some fraction of users, and that number belongs in your error budget arithmetic. Writing it down explicitly — detection time plus TTL equals worst-case failover window — turns a vague "we have failover" claim into a number you can defend, and it is exactly the kind of reasoning the exam rewards when it asks you to choose between two failover designs.

The runbook shape follows from that. A routing-layer runbook should answer four questions in order: which record set is involved and what policy does it use, what is the current health-check status for each record in that set, what is the TTL and therefore how long until clients see a change, and is the standby actually healthy right now. That last question is the one that separates a working runbook from a hopeful one, and it is the same question Route 53 ARC's readiness checks answer continuously for the region-level case. For day-to-day operations, the practical discipline is to treat health checks as production configuration under change control, review them whenever the application's dependency graph changes, and re-test failover deliberately rather than assuming it still works because it worked when it was built.

8. Edge Cases and Exam Gotchas

The single most reliable gotcha is that simple routing has no health-check-driven failover. If a scenario describes a single record and asks how to make it resilient, the answer is never "add a health check to the simple record" — it is to move to a policy that supports multiple records, or to put a load balancer behind the record and let the load balancer handle target health. Candidates lose points here because "attach a health check" is the correct instinct for almost every other policy and is simply unavailable for this one.

The second gotcha is the distinction between latency and geoproximity when the scenario mentions compliance or data residency. Latency routing optimizes for speed and will happily send a European user to a US region if that path measures faster, which violates a residency requirement. Geolocation routing pins users to a mapped location regardless of speed, which satisfies residency but can be slower. Geoproximity is a distance-based compromise with a tunable bias and is not a residency control. If the scenario says "data must remain in the EU," the answer is geolocation, not latency, and not geoproximity.

The third gotcha is that health checks can be based on CloudWatch alarms, which means any metric you can alarm on becomes a routing input. This is the escape hatch for conditions Route 53 cannot probe directly — a queue depth, a custom business metric, a database replication lag — and scenarios that describe routing based on something other than endpoint reachability are usually pointing at this. The related gotcha is calculated health checks, which combine child checks with boolean logic and are the mechanism for expressing "healthy if at least two of these three regions are healthy," a pattern that appears in multi-region designs.

The remaining gotchas are smaller but recur. Alias records are required at the zone apex because CNAMEs are not permitted there, and alias records to AWS resources are free to query. Weighted records with all weights set to zero behave as if all records are equally weighted, which is a common accidental configuration. A health check that monitors an endpoint by IP address rather than by hostname will not send the Host header the application expects, producing false unhealthy results. And multi-value answer routing returns up to eight records but does not perform any health-based load balancing beyond excluding unhealthy records, so it is not a substitute for an ELB when the scenario needs connection-level behavior.

9. This vs. the Services It Gets Confused With

Route 53 routing policies are frequently confused with Global Accelerator and with load balancers, and the confusion is understandable because all three distribute traffic. The distinction is the layer at which the decision is made and what the client experiences. Route 53 makes the decision in DNS, which means the client receives an address and connects to it directly, which in turn means the decision is cached and the failover speed is bounded by TTL. Global Accelerator makes the decision at the network edge using static anycast IPs, so the client connects to an address that never changes and the rerouting happens inside AWS's network without any DNS caching delay. A load balancer makes the decision per connection or per request within a single region, with rich health semantics for individual targets.

MechanismDecision layerFailover speedHealth granularityPick it when…
Route 53 routing policyDNS resolutionDetection + TTLPer recordYou need geographic, weighted, or latency-based steering across regions or endpoints
Route 53 health checkDNS resolutionDetection + TTLEndpoint, calculated, or alarm-basedYou need automatic removal of unhealthy answers
Global AcceleratorAWS network edge (anycast)Seconds, no DNS cachingEndpoint group per regionClients cache DNS aggressively or you need instant regional failover
Application Load BalancerPer request, within a regionSeconds, target-levelPer target, with path-based checksYou need path/host routing, sticky sessions, or per-target draining
Network Load BalancerPer connection, within a regionSeconds, target-levelPer target, TCP/HTTPYou need extreme throughput or static IPs per AZ

The practical rules that fall out of this table are worth stating as decision rules. Pick a Route 53 routing policy when the steering decision is about which region or endpoint a user should reach and you can tolerate a failover window measured in detection time plus TTL. Pick Global Accelerator when the client population caches DNS beyond your control, when you need a fixed IP allowlist, or when the failover window must be seconds rather than minutes. Pick an ALB when the decision is about which target within a region should serve a request and you need request-level routing or session affinity. And pick Route 53 ARC, from yesterday, when the decision is a deliberate regional cutover that must be auditable and must not depend on the health of the region you are cutting over from.

Hands-on Lab: Weighted Canary with Automatic Health-Check Rollback (45 min)

The goal is to build a weighted routing configuration that sends five percent of traffic to a canary endpoint and automatically stops sending traffic to it if the canary becomes unhealthy, then to prove the rollback works by deliberately breaking the canary. Work in a test hosted zone rather than a production one, and use two distinct endpoints so the health-check behavior is observable.

1. Establish the two endpoints. Create two targets that are clearly distinguishable in a response body — for example two ALBs, or two EC2 instances behind a single ALB with different target groups, or two S3 static website endpoints. Each should return something that identifies which one answered, such as a version string in the response body or a distinct HTTP header. Confirm both are reachable directly before involving DNS.

2. Create the health checks first. Create an HTTP or HTTPS health check for each endpoint, pointing at a path that exercises the application rather than a static file, and configure it to accept only a 200 response. Set the interval to the fast option and the failure threshold to a low value so the canary is pulled quickly, and note the health-check IDs. Wait for both checks to report healthy in the console before continuing — a routing configuration built on checks that have never been observed healthy is not a valid test.

3. Create the weighted record set. In the test hosted zone, create two records with the same name and type, both using the weighted routing policy. Assign the primary a weight of 95 and the canary a weight of 5, and attach the corresponding health check to each record. Set the TTL to 60 seconds so the experiment completes in a reasonable time. If you are using alias records to AWS resources, note that the TTL is inherited from the target rather than set explicitly.

4. Verify the split. Query the name repeatedly from a client that does not cache — a loop of dig or Resolve-DnsName calls, or a short script that resolves the name many times — and count how often each endpoint's identifying response appears. With a 95/5 split you should see the canary roughly one time in twenty, and the exact ratio will vary because the selection is probabilistic per query. Record the observed ratio as your baseline.

5. Break the canary deliberately. Make the canary endpoint return a non-200 status, or stop the process behind it, so that its health check begins failing. Watch the health-check status transition from healthy to unhealthy, and note how long that takes given the interval and failure threshold you configured. This elapsed time is the detection half of your failover window.

6. Confirm the automatic rollback. Once the canary's health check is unhealthy, repeat the resolution loop. Every answer should now be the primary, because the canary record has been removed from the set and the remaining weight renormalized to 100 percent. Confirm this both by observing the answers and by checking that the canary record shows as unhealthy in the console. This is the behavior that makes weighted routing safe for canary releases: a bad canary removes itself.

7. Restore and observe recovery. Bring the canary back to a healthy state and watch the health check return to healthy after the success threshold is met. Confirm that the five percent split resumes. Note the recovery time separately from the detection time, because the success threshold adds to it.

8. Write down the failover window. Record the total elapsed time from breaking the canary to the last observed canary answer, and decompose it into detection time plus TTL. This number is the one you would put in a runbook, and it is the number an exam scenario is implicitly asking you to reason about when it gives you an interval, a threshold, and a TTL.

9. Extend to failover routing. As a second exercise, create a failover pair in the same zone with the primary and secondary records, attach a health check to the primary only, and break the primary. Observe that the secondary is returned, then break the secondary as well and observe that Route 53 still returns it — the demonstration that failover routing has no third option and that a robust design must monitor the standby separately.

Scenario Question Drills (20 min)

Q1. You want most users routed to the AWS region with the lowest network latency for them, automatically, with no manual tuning.

A. Weighted routing
B. Latency-based routing policy
C. Simple routing
D. Geolocation routing
Correct answer: B. Latency-based routing uses AWS's latency measurements between users and regions to route each request to the lowest-latency healthy endpoint, with no tunable bias.

Q2. A single Route 53 A record points at one EC2 instance. The team asks how to make it fail over automatically if the instance dies.

A. Attach a health check to the simple routing record
B. Simple routing does not support health-check failover — move to a policy with multiple records, or put a load balancer behind the record
C. Lower the TTL to 0
D. Enable multi-value answers on the simple record
Correct answer: B. Simple routing has exactly one record and no health-check-driven removal, so there is nothing to fail over to. Resilience requires a multi-record policy or a load balancer handling target health.

Q3. A weighted record set has two records with weights 95 and 5. The record with weight 5 fails its health check. What does Route 53 return?

A. Nothing, because the set is incomplete
B. The healthy record 95 percent of the time and nothing 5 percent of the time
C. The healthy record for all answers, because the failed record is removed and the remaining weight renormalizes to 100 percent
D. The failed record, because weights override health
Correct answer: C. Unhealthy records are excluded from the set and the surviving weights are renormalized, so a failed canary removes itself and the primary serves all traffic.

Q4. A European bank must ensure that customer DNS queries from EU countries are answered by an EU-hosted endpoint regardless of which region would be faster.

A. Latency-based routing
B. Geoproximity routing with a bias toward the EU region
C. Geolocation routing, mapping European countries to the EU endpoint
D. Multi-value answer routing
Correct answer: C. Geolocation routing pins queries to a mapped location regardless of measured latency, which is what a data-residency requirement demands. Latency routing optimizes for speed and geoproximity is a tunable distance heuristic, not a residency control.

Q5. A team wants to shift traffic gradually from us-east-1 to us-west-2 while both regions remain fully healthy, with the ability to pause or reverse the shift.

A. Failover routing with the primary in us-east-1
B. Geoproximity routing with a bias parameter adjusted over time
C. Latency-based routing
D. Simple routing with a low TTL
Correct answer: B. Geoproximity routing's bias parameter is the only mechanism that deliberately pulls traffic toward or away from a healthy region by a tunable amount, which is exactly a gradual regional migration.

Q6. A health check is configured against an ALB's DNS name over HTTPS. During an application outage where every target returns 503, the health check still reports healthy and no failover occurs. Why?

A. Route 53 health checks cannot monitor ALBs
B. The health check is monitoring the load balancer's reachability, not the application's health, and is not configured to treat 503 as unhealthy
C. The TTL is too high
D. The record set has too many records
Correct answer: B. The ALB keeps answering even when all targets fail, so the check must point at a deep-health endpoint and accept only 200 to reflect real application health.

Q7. A failover configuration has a health check on the primary record only. The primary fails, traffic moves to the secondary, and users still report errors. What is the most likely cause?

A. Failover routing requires health checks on both records to function
B. The secondary is also unhealthy, and failover routing has no third option so it returns the secondary anyway
C. The TTL is set too low
D. The secondary record needs a higher weight
Correct answer: B. Failover routing returns the secondary whenever the primary is unhealthy, without checking the secondary's own health. A robust design monitors the standby separately.

Q8. A team needs routing decisions based on a custom application metric — queue depth — rather than endpoint reachability. What is the mechanism?

A. A TCP health check against the queue endpoint
B. A CloudWatch-alarm-based health check, which turns any alarmable metric into a routing signal
C. Multi-value answer routing
D. A calculated health check combining three endpoint checks
Correct answer: B. Health checks can be backed by CloudWatch alarms, which is the escape hatch for conditions Route 53 cannot probe directly, such as queue depth or replication lag.

Q9. A multi-region design should be considered healthy as long as at least two of three regional endpoints are healthy. Which health-check construct expresses this?

A. Three independent endpoint health checks with no aggregation
B. A calculated health check combining the three child checks with boolean logic
C. A single TCP health check against one region
D. A CloudWatch composite alarm alone
Correct answer: B. Calculated health checks combine child checks with boolean logic, which is how you express "healthy if at least two of three are healthy."

Q10. A record has a 300-second TTL and its health check uses the standard 30-second interval with the default failure threshold of three. Roughly how long after an endpoint fails will all clients stop receiving its address?

A. About 30 seconds
B. About 90 seconds
C. About 90 seconds of detection plus up to 300 seconds of TTL caching
D. Immediately, because Route 53 pushes updates to resolvers
Correct answer: C. Effective failover time is detection time (interval times failure threshold, roughly 90 seconds) plus the TTL, because resolvers and clients cache the previous answer until it expires.

Q11. A weighted record set has three records, all with weight set to 0. What happens?

A. Route 53 returns nothing
B. The records are treated as equally weighted, so traffic is split evenly
C. Only the first record is returned
D. The configuration is rejected
Correct answer: B. All-zero weights are treated as equal weighting, which is a common accidental configuration when someone intends to disable a record.

Q12. A team wants redundancy across four web servers without paying for a load balancer, and the application is stateless with clients that retry on connection failure.

A. Simple routing with four A records
B. Multi-value answer routing with a health check on each record
C. Failover routing with four records
D. Geolocation routing
Correct answer: B. Multi-value answer routing returns up to eight healthy records per response and excludes unhealthy ones, giving crude client-side spreading without an ELB — appropriate for stateless clients that retry.

Q13. A health check is configured against an endpoint by IP address rather than by hostname, and it reports unhealthy even though the application responds correctly to browser requests. What is the likely cause?

A. Route 53 cannot health-check IP addresses
B. The check is not sending the Host header the application expects, so the application returns an error the check treats as unhealthy
C. The TTL is too low
D. The endpoint needs an alias record
Correct answer: B. Health checks by IP do not send the expected Host header, which commonly produces false unhealthy results for virtual-hosted applications.

Q14. A team declares a five-minute RTO, configures failover routing correctly, and the console shows the failover succeeded — but a subset of users remains broken for hours. What is the most likely explanation?

A. The health check threshold is too low
B. Clients with long-lived connections never re-resolve DNS and keep talking to the failed endpoint
C. The secondary record has no health check
D. The hosted zone quota was exceeded
Correct answer: B. DNS-based failover only affects new resolutions. Clients holding connections never re-resolve, which is the case for a network-layer mechanism such as Global Accelerator.

Q15. A scenario requires a fixed IP address that clients can allowlist, and failover that completes in seconds without depending on client DNS caching. Which mechanism fits?

A. Route 53 latency-based routing with a 60-second TTL
B. Route 53 failover routing with health checks
C. AWS Global Accelerator, which uses static anycast IPs and reroutes at the AWS network edge
D. Multi-value answer routing
Correct answer: C. Global Accelerator provides static anycast IPs for allowlisting and reroutes inside the AWS network without any DNS caching delay, unlike any Route 53 policy.

Peek into Tomorrow: What Happens When the Answer Is "Restore It"

Everything in this day assumed the data still exists somewhere and the problem is choosing which healthy copy to serve it from. Routing policies and health checks are excellent at that, and they are useless when the failure is not an unhealthy endpoint but a deleted table, a corrupted volume, or an account that has been compromised and had its snapshots removed. The open question this leaves is what the recovery path looks like when the correct answer is not "route around it" but "put it back" — and specifically, how you make that path trustworthy when the thing you are recovering from may have had administrative access to your backups.

That question has a shape that DNS cannot address. It requires a policy that decides what gets backed up and how often, a lifecycle that decides how long each copy survives, and a copy that lives somewhere the compromised account cannot reach. It also requires immutability, because a backup that an administrator can delete is not a backup against ransomware — it is a backup against hardware failure only. Tomorrow's material covers centralized, policy-based backup across EBS, RDS, DynamoDB, and EFS, cross-region and cross-account copy, and Vault Lock's write-once-read-many guarantee, which is the mechanism that makes the last line of defense actually hold.

Sources