Day 36 of 70 · Week 6
Day 36 / 70 Week 6 of 14 Phase 3: SRE Observability, Resilience & DR

Multi-Region Active-Passive & Pilot Light Patterns

🕑 ~58 min read · 3 services covered
Pilot Light Warm Standby Backup & Restore

Recap: Where We Left Off

Day 35 pushed the multi-region conversation to its logical extreme: serving live traffic from two or more regions simultaneously, with Global Accelerator anycast IPs or Route 53 geoproximity routing steering users to the nearest healthy endpoint and a multi-writer data tier underneath. That architecture buys the lowest possible RTO and RPO, but it also demands that every write path be conflict-aware — either through last-writer-wins semantics in DynamoDB Global Tables or through deliberate partitioning so that two regions never contend for the same key. The cost of that guarantee is real: you pay for full production capacity in every region, all the time, and you accept a permanent increase in operational complexity.

Today's material contrasts with that posture rather than extending it. Active-passive patterns accept a period of degraded or absent service in exchange for dramatically lower standing cost, and the engineering problem shifts from conflict resolution to failover orchestration. The question is no longer "how do two live regions agree on state" but "how quickly and reliably can a dormant region be brought to life, and how much data do we lose in the gap." That reframing is what makes Pilot Light, Warm Standby, and Backup & Restore a distinct family of designs rather than just cheaper versions of active-active.

Foundations You'll Need Today

Today's material is about what happens when an entire AWS region stops working, and about the different ways you can prepare for that. Before the strategies make sense, four underlying ideas need to be clear. None of them are complicated on their own, but the exam assumes all four are second nature, so it is worth making them explicit.

Regions and Availability Zones

AWS runs its infrastructure in geographically separate clusters called regions — Northern Virginia, Oregon, Ireland, Singapore, and so on. Each region is an independent island: it has its own power, cooling, networking, and control plane, and a failure in one region does not directly affect another. Inside each region are several Availability Zones, which are physically distinct data centers (or groups of them) with independent power and networking, connected to each other by low-latency private links. The practical consequence is a hierarchy of failure: a single server can fail, an entire Availability Zone can fail, or — much more rarely — an entire region can fail. Spreading your workload across multiple AZs in one region protects you against the middle case. It does nothing at all for the last case, because every AZ in that region goes down together. That distinction is the single most important thing to hold onto today, and it is the reason "Multi-AZ" and "multi-region" are not interchangeable terms.

Replication, RPO, and RTO

If you want a copy of your data somewhere else, you replicate it — you continuously copy changes from the primary location to a secondary one. Replication is almost always asynchronous, meaning the copy lags slightly behind the original, because waiting for the copy to confirm every write would slow the primary down. That lag is the source of the two numbers the rest of this day revolves around. RPO (Recovery Point Objective) is how much data you are willing to lose, expressed as time — an RPO of five minutes means that after a disaster, you accept losing up to the last five minutes of writes. RTO (Recovery Time Objective) is how long you are willing to be down — an RTO of four hours means the business can tolerate four hours of the service being unavailable. These are business decisions, not technical ones, and they are usually stated in the scenario. Every DR strategy in this day is really just a different answer to "how do we hit these two numbers for the least money."

DNS and Traffic Shifting

When a user types your domain name, their computer asks a DNS resolver to translate that name into an IP address. AWS's DNS service, Route 53, can be configured to answer that question differently depending on which endpoints are healthy — this is called failover routing. If the primary region's health check starts failing, Route 53 begins handing out the standby region's IP address instead. The catch is caching: resolvers and operating systems remember DNS answers for a period called the TTL (time to live), so a change does not reach every user instantly. A low TTL shortens that delay but increases DNS query volume. AWS Global Accelerator sidesteps the problem entirely by giving you a pair of static IP addresses that never change; traffic is rerouted at AWS's network edge rather than through DNS, so there is no caching delay to wait out. Both mechanisms appear in today's scenarios, and knowing why you would pick one over the other is part of the decision.

AMIs, Snapshots, and Infrastructure as Code

An AMI (Amazon Machine Image) is the template an EC2 instance is launched from — it bundles the operating system, the installed software, and the configuration into a single artifact. An EBS snapshot is a point-in-time copy of a disk volume, and an RDS snapshot is the equivalent for a managed database. The critical detail for today is that all three are regional resources: an AMI created in Northern Virginia simply does not exist in Oregon, and a launch template in Oregon that references it will fail. To use it in another region you must explicitly copy it there. Infrastructure as code means describing your environment — networks, servers, load balancers, permissions — in version-controlled configuration files rather than clicking through the console, so that the same description can be deployed into a second region on demand. Together, these are the raw materials of a disaster recovery plan: the snapshots and AMIs are what you restore from, and the infrastructure-as-code is what rebuilds everything around them.

With that grounding, here is why the four disaster recovery strategies exist as a spectrum, and what problem each one actually solves.

1. Why This Is on the Exam

Disaster recovery design is one of the most heavily weighted topics in the SAP-C02 blueprint, and it sits squarely inside the Resilient Architectures domain. The reason it appears so often is that DR is where cost, reliability, and business requirements collide most visibly. A question that asks you to choose between four architectures is almost always really asking whether you can read a stated RTO and RPO, translate them into a required recovery mechanism, and then pick the cheapest design that satisfies both constraints. Candidates who answer by picking the most resilient option fail these questions constantly, because the exam rewards the cheapest sufficient answer, not the strongest possible one.

The second reason this topic is exam-dense is that it forces you to reason about what "recovery" actually means for each layer of a stack. Compute can be recreated from an AMI or a container image in minutes. Data cannot be recreated at all if it was never replicated. Configuration, DNS, certificates, and IAM roles sit somewhere in between — reproducible in principle, but only if they were codified. A scenario that says "the primary region is unavailable" is testing whether you understand which of those layers is the actual bottleneck, because the DR strategy you choose is determined by the slowest layer to recover, not the fastest.

Finally, the exam uses DR scenarios to probe whether you understand the difference between a backup and a replica. A backup is a point-in-time artifact you restore from, and restoring it takes time proportional to its size. A replica is a continuously updated copy you can promote, and promoting it takes time proportional to detection and cutover, not data volume. Backup & Restore strategies are bounded by restore time; Pilot Light and Warm Standby strategies are bounded by detection and scaling time. Recognizing which bound applies to a given scenario is usually enough to eliminate two of the four answer choices immediately.

2. How the Spectrum Actually Works

The four canonical DR strategies are not four separate products. They are four points on a single continuum defined by how much of the standby environment is kept running before a disaster occurs. At one end, Backup & Restore keeps nothing running in the standby region except storage for the backups themselves. At the other end, Multi-Site Active-Active keeps a full production footprint running in every region and serves traffic from all of them. Pilot Light and Warm Standby occupy the middle, and the difference between them is precisely how much of the compute layer is pre-provisioned.

Pilot Light takes its name from the small flame that keeps a gas furnace ready to ignite. In AWS terms, the pilot light is the data tier. You replicate your databases, object storage, and configuration into the standby region continuously, so that the data is current to within seconds or minutes. But you do not run application servers there. The standby region holds AMIs, launch templates, Auto Scaling group definitions, and infrastructure-as-code that can be deployed on demand, but no running compute. When a disaster is declared, you launch the compute layer against the already-current data, attach it to the network, and cut DNS over. The recovery time is dominated by how long it takes to boot and warm the application tier, which is why Pilot Light RTOs are typically measured in tens of minutes rather than hours.

Warm Standby keeps a scaled-down but fully functional copy of the entire stack running in the standby region. The application servers exist, the load balancer exists, the Auto Scaling group exists — it is simply sized to handle a small fraction of production traffic, often the minimum capacity needed to keep the deployment healthy. On failover, you scale that group up to production size and shift traffic. Because the compute is already running and already warm, the recovery time is dominated by the scaling event and the DNS or traffic-shift propagation, which is why Warm Standby RTOs land in the single-digit-minutes range. The tradeoff is that you are paying for that idle capacity every hour of every day.

Backup & Restore is the simplest and cheapest strategy and the one most often misapplied. You take backups — snapshots, database dumps, object replication — and store them in the standby region or a dedicated backup account. Nothing else exists there. On disaster, you provision the entire environment from infrastructure-as-code, restore the data from the most recent backup, and cut over. The RTO here is genuinely hours to days, because it includes provisioning time plus restore time, and the RPO is whatever your backup interval was. This is the correct answer whenever the business genuinely tolerates that window, and the exam will sometimes present exactly that scenario to see whether you resist the urge to over-engineer.

3. The Core Decision Boundary

Every DR scenario question reduces to a single fork: what are the stated RTO and RPO, and which is the cheapest strategy on the spectrum that satisfies both? The trap is that candidates tend to optimize for one of the two numbers and ignore the other. A scenario that specifies an RPO of five minutes but an RTO of four hours is not asking for Warm Standby — it is asking for a strategy with frequent replication but slow recovery, which is Backup & Restore with a short backup interval, or Pilot Light with a deliberately unhurried cutover process. Conversely, a scenario with a generous RPO but a tight RTO is asking for pre-warmed compute with infrequent data replication.

The second half of the decision boundary is the cost constraint, which the exam usually states explicitly with phrases like "minimize cost" or "the business wants to avoid paying for idle capacity." When that phrase appears, it is a strong signal that the answer is the cheapest strategy that still meets the stated RTO and RPO, and that any answer involving always-on standby compute is a distractor. When the phrase is absent and the RTO is aggressive, the cost constraint is implicitly subordinate to the recovery requirement.

The table below is the one to internalize. Note that the RTO and RPO columns are ranges, not guarantees — the actual numbers depend on your replication frequency, your instance boot times, and how much of the recovery you have automated. The exam treats these as characteristic ranges rather than precise figures, and so should you.

StrategyStandby computeTypical RTOTypical RPORelative cost
Backup & RestoreNoneHours to daysHours (backup interval)Lowest
Pilot LightNone (data only)Tens of minutesSeconds to minutesLow
Warm StandbyScaled-down, runningMinutesSecondsModerate
Multi-Site Active-ActiveFull, in every regionNear zeroNear zeroHighest

4. Configuration Modes and Their Tradeoffs

Within each strategy there are configuration choices that move the RTO and RPO numbers, and understanding those knobs is what separates a memorized answer from a defensible one. For Backup & Restore, the dominant knob is backup frequency and retention. AWS Backup lets you define schedules as often as hourly, and cross-region copy jobs replicate those recovery points into the standby region. Shortening the interval improves RPO linearly and increases storage cost and API activity linearly. The RTO knob is separate: it is how much of the rebuild is automated. A runbook that requires a human to click through the console has an RTO floor set by human response time, which is why the exam favors answers that describe infrastructure-as-code and automated restore.

For Pilot Light, the knobs are replication lag and launch automation. Data replication is typically continuous — Aurora Global Database for relational workloads, DynamoDB Global Tables or cross-region replication for key-value, S3 Cross-Region Replication for objects — so RPO is governed by the replication mechanism's lag rather than by a schedule. The RTO knob is how quickly the compute layer can be brought up. Pre-baked AMIs, pre-created launch templates, and pre-provisioned networking reduce that to a scaling event. If the launch template references an AMI that only exists in the primary region, you have silently added an AMI copy step to your RTO, which is a classic exam gotcha.

For Warm Standby, the knobs are the standby fleet size and the scaling policy that grows it. A standby sized at ten percent of production will scale to full capacity faster than one sized at two percent, because fewer instances need to be added, but it costs more to keep running. The scaling policy itself matters: a target-tracking policy that reacts to CPU will lag behind a sudden traffic shift, whereas a scheduled or manual scaling action triggered by the failover runbook responds immediately. The exam tends to favor the explicit scaling action in failover scenarios, because it is deterministic.

Across all three, the traffic-shift mechanism is a shared knob with its own tradeoffs. Route 53 failover routing with health checks is the standard approach, but DNS TTL and client-side caching mean the shift is not instantaneous — a low TTL reduces propagation delay at the cost of more query volume. Global Accelerator avoids the DNS caching problem entirely by using static anycast IPs and shifting traffic at the network edge, which is why it appears in scenarios that demand fast, deterministic failover. The choice between them is a recurring exam decision point and is covered in more depth in the comparison section below.

5. Sizing, Limits and Quotas

DR designs fail on quotas more often than on architecture. The most common cause is that a standby region has never run at production scale, so its service quotas were never raised. EC2 On-Demand vCPU limits, Elastic IP quotas, load balancer counts, and RDS instance class availability all differ by region and by account history. A Pilot Light design that assumes it can launch two hundred instances in the standby region on demand will discover at the worst possible moment that the account's vCPU quota there is set to the default. The mitigation is to request quota increases in the standby region ahead of time and to verify them periodically, which is exactly the kind of readiness validation that Route 53 ARC formalizes.

Data transfer is the second sizing dimension. Cross-region replication of a large dataset is billed as data transfer out of the source region, and for continuous replication that cost accrues every hour. Aurora Global Database replication, S3 Cross-Region Replication, and DynamoDB global table replication all carry this cost, and it scales with write volume rather than with dataset size. A workload that writes heavily will pay substantially more for cross-region replication than one that is read-heavy, and that asymmetry sometimes makes Backup & Restore the economically correct choice even when a tighter RPO is technically achievable.

Restore throughput is the third dimension and the one most often underestimated. Restoring an RDS snapshot is not instantaneous; the time scales with the size of the snapshot and the storage type, and a multi-terabyte restore can take hours. S3 restores from Glacier classes have their own retrieval times, which range from milliseconds for Glacier Instant Retrieval to hours for Glacier Flexible Retrieval and up to twelve hours for Deep Archive. If your backup strategy stores recovery points in an archival class, your effective RTO includes that retrieval time, and the exam will sometimes present exactly that mismatch as the flaw in an otherwise reasonable design.

Finally, consider the limits on the failover mechanism itself. Route 53 health checks have a minimum evaluation interval and a required number of consecutive failures before an endpoint is marked unhealthy, which sets a floor on detection time. Auto Scaling group scaling activities are rate-limited by how quickly instances can launch and pass health checks. Every one of these contributes to the real RTO, and the sum is always larger than the individual pieces suggest.

6. Failure Modes and What They Look Like in Production

The most common failure mode in active-passive designs is not the disaster itself but the discovery, during the disaster, that the standby environment does not work. This manifests as a failover that succeeds at the infrastructure layer and fails at the application layer: instances launch, the load balancer reports healthy targets, and the application returns errors because a configuration value, a secret, a certificate, or a database credential was never replicated. The symptom is a green infrastructure dashboard alongside a red application error rate, and the first diagnostic move is to compare the standby environment's configuration against the primary's, ideally by diffing the infrastructure-as-code rather than by inspecting running resources.

The second failure mode is replication lag that is larger than assumed. Aurora Global Database replication is typically sub-second, but it degrades under heavy write load or network congestion, and the lag is visible in the AuroraGlobalDBReplicationLag metric. DynamoDB global table replication lag is visible in ReplicationLatency. If your stated RPO is five minutes and your actual lag during peak load is fifteen, your design does not meet its requirement, and you will only find out by monitoring the lag metric continuously rather than assuming the documented typical value holds. The first diagnostic move when a failover produces unexpected data loss is to pull the replication lag metric for the window preceding the event.

The third failure mode is a failover that works but never fails back. Failing back to the primary region is a separate, often harder operation than failing over, because the primary's data is now stale and must be resynchronized before it can resume serving. Teams that rehearse failover but not failback discover this during the incident. The symptom is a prolonged period of running in the standby region at higher cost and higher latency for the user base, and the diagnostic move is to check whether the failback runbook exists at all.

The fourth is the silent partial failure, where the primary region is degraded rather than down. Health checks may still pass, so automated failover never triggers, and the system limps along at reduced capacity. This is the hardest case to handle because it requires a human judgment call, and it is the reason the exam favors designs that include both automated health-check-based failover and a manual override.

7. The Operational and SRE Angle

From an SRE perspective, a DR strategy is a promise about recovery time, and an untested promise is not a promise at all. The operational discipline that makes active-passive designs trustworthy is regular, scheduled failover testing — a Game Day in which the primary region is deliberately taken out of service and the standby is required to carry production traffic. The value of the exercise is not that it proves the design works; it is that it finds the configuration drift, the expired certificate, and the quota that was never raised, while the stakes are low. Teams that run these exercises quarterly tend to have failovers that work; teams that run them never tend to have failovers that almost work.

The monitoring that supports this is specific. You need a replication lag alarm in every region, not just the primary, because lag is the leading indicator of RPO breach. You need a synthetic canary in the standby region that exercises the application end to end, because infrastructure health checks do not prove the application can serve traffic. You need an alarm on the standby environment's capacity and quota headroom, because a standby that cannot scale to production size is not a standby. And you need a dashboard that shows the current replication lag, the standby's readiness state, and the time since the last successful failover test, because that last number is the one that predicts whether the next failover will work.

The runbook shape follows from the strategy. A Pilot Light runbook has an explicit launch phase, a validation phase, and a traffic-shift phase, and each phase should have a defined success criterion and a defined abort path. A Warm Standby runbook is shorter — scale, validate, shift — but the validation phase is more important because the environment was already running and may have drifted. In both cases the runbook should be executable by someone who did not write it, at three in the morning, without access to the person who did. That is the operational test that matters.

Finally, the SLO implications are worth stating plainly. A DR strategy with a fifteen-minute RTO means that a regional failure consumes your entire monthly error budget for availability if your SLO is measured in nines. That is not necessarily wrong — regional failures are rare — but it should be a conscious decision rather than a surprise, and it should be reflected in how the SLO is defined and reported.

8. Edge Cases and Exam Gotchas

The single most common exam trap in this topic is the AMI and snapshot region scoping problem. AMIs, EBS snapshots, and RDS snapshots are regional resources. A launch template in the standby region that references an AMI ID from the primary region will fail, and the fix is to copy the AMI or snapshot into the standby region as part of the replication pipeline. Scenarios that describe a Pilot Light design and then mention that the standby launch fails are testing exactly this.

The second trap is confusing Multi-AZ with multi-region. Multi-AZ is a high-availability mechanism within a single region; it protects against an Availability Zone failure and does nothing for a regional failure. A scenario that says "the application must survive the loss of an entire AWS region" cannot be satisfied by Multi-AZ alone, no matter how many AZs are configured. This distinction appears in some form on nearly every practice exam.

The third trap is the assumption that a read replica is a DR solution. A cross-region read replica can be promoted, but promotion is a manual operation, the replica is asynchronous so there is data loss on promotion, and the promoted instance does not automatically inherit the primary's configuration, parameter groups, or security group associations in all cases. It is a component of a DR strategy, not a strategy by itself.

The fourth trap is the cost of the standby region's data transfer, which is easy to overlook when comparing strategies on compute cost alone. A Pilot Light design with continuous cross-region replication of a high-write database can cost more in transfer than a Warm Standby design costs in idle compute, which inverts the intuitive cost ordering. When a scenario emphasizes cost, check whether the write volume makes replication the dominant expense.

The fifth trap is the assumption that failover is automatic. Route 53 failover routing with health checks can automate the DNS shift, but the compute layer still has to be launched or scaled, and that step is often manual or triggered by a separate mechanism. A scenario that describes a fully automated failover needs both halves to be automated, and the exam will sometimes present a design where only the DNS half is.

9. This vs. the Strategies It Gets Confused With

The strategies in this family are most often confused with each other, and the distinguishing question is always the same: what is already running in the standby region before the disaster? If the answer is "nothing but backups," it is Backup & Restore. If it is "the data, but not the compute," it is Pilot Light. If it is "a scaled-down version of everything," it is Warm Standby. If it is "a full production footprint serving live traffic," it is Active-Active. Answering that question first collapses most scenario questions to a single candidate.

The second confusion is between Pilot Light and Warm Standby specifically, because both involve continuous data replication and both have RTOs measured in minutes rather than hours. The discriminator is whether the application tier is running. A scenario that mentions "a small Auto Scaling group running in the standby region" is describing Warm Standby. A scenario that mentions "AMIs and launch templates ready to deploy" is describing Pilot Light. The cost difference between them is the cost of that idle compute, and the RTO difference is the time it takes to launch and warm it.

The third confusion is between these strategies and the availability mechanisms that operate within a region. Multi-AZ, Auto Scaling, and load balancer health checks all improve availability without providing disaster recovery. They are complementary, not alternative, and a complete design uses both: Multi-AZ within each region for zone-level resilience, and one of these four strategies across regions for regional resilience.

Scenario signalPickWhy
RTO hours, RPO hours, minimize costBackup & RestoreNothing needs to be running; restore time is acceptable
RTO tens of minutes, RPO seconds, cost-sensitivePilot LightData replicated continuously, compute launched on demand
RTO minutes, RPO seconds, some idle budgetWarm StandbyCompute already running, only needs scaling
RTO near zero, RPO near zero, cost secondaryMulti-Site Active-ActiveBoth regions serve live traffic continuously
Survive loss of one AZ onlyMulti-AZ (not a DR strategy)Zone-level HA, not regional DR

Hands-on Lab: Mapping RTO/RPO to the Cheapest Sufficient Strategy (45 min)

The goal of this lab is to build the reflex the exam rewards: read a set of business requirements, derive the RTO and RPO they imply, and select the cheapest strategy that satisfies both. You will work through five workload profiles, justify each choice in writing, and then verify one of them by actually building a Pilot Light data replication path.

Step 1 — Build the requirement table. Create a table with five rows, one per workload profile below, and columns for stated business requirement, derived RTO, derived RPO, chosen strategy, and justification. The five profiles are: (a) an internal HR system used during business hours, where a day of downtime is tolerable and losing a day of data is acceptable; (b) a customer-facing e-commerce checkout service where the business has stated that four hours of downtime is the maximum tolerable and losing more than five minutes of orders is unacceptable; (c) a real-time bidding platform where any downtime directly loses revenue and the business has approved budget for full redundancy; (d) a batch analytics pipeline that runs nightly and can be re-run from source data if a region is lost; (e) a regulatory archive that must be retained for seven years and is queried a few times a month.

Step 2 — Derive, don't assume. For each profile, write the RTO and RPO as explicit numbers before choosing a strategy. The discipline here is to resist jumping to a strategy based on the workload's perceived importance. Profile (d) sounds important but has a naturally generous RTO because the pipeline can be re-run; profile (b) sounds routine but has a tight RPO because orders are money. The exam tests exactly this inversion.

Step 3 — Choose and justify. For each profile, select the cheapest strategy on the spectrum that meets both numbers, and write two sentences of justification. If your chosen strategy is Pilot Light or Warm Standby, state explicitly what is running in the standby region and what is not. If your chosen strategy is Backup & Restore, state the backup interval that satisfies the RPO and the automation that satisfies the RTO.

Step 4 — Build one Pilot Light replication path. Take profile (b) and implement the data half of a Pilot Light design. Create an S3 bucket in a primary region and enable Cross-Region Replication to a bucket in a standby region, with replication metrics enabled. Write a small script or use the console to upload a set of objects, then confirm that they appear in the standby bucket and that the replication metrics show zero lag once the transfer completes. This is the pilot light: the data is current, and nothing else exists in the standby region.

Step 5 — Identify the missing compute half. Now write down, as a numbered list, everything that would have to happen to serve traffic from the standby region for profile (b). Include the AMI or container image availability, the launch template, the network configuration, the load balancer, the DNS change, and the database promotion. For each item, note whether it is already replicated or would have to be created at failover time. The length of that list is your real RTO, and the items that are not replicated are the ones that will break during an actual incident.

Step 6 — Add the readiness check. Finally, write a single CloudWatch alarm or synthetic canary that would tell you, before a disaster, whether the standby region is actually ready. For the S3 replication path you built, that is a replication lag alarm. For the compute half, it would be a canary that exercises the application in the standby region. Note in your write-up which of these you could implement today and which would require the compute layer to exist first — that gap is precisely the problem the next day's material addresses.

Scenario Question Drills (20 min)

Q1. A workload can tolerate an RTO of 15 minutes and an RPO of 5 minutes, but the business wants to minimize standing infrastructure cost. Which DR pattern fits best?

A. Backup and Restore
B. Pilot Light
C. Multi-Site Active-Active
D. Multi-AZ within a single region
Correct answer: B. Pilot Light keeps only core data continuously replicated and minimal infrastructure running, scaling up the rest on failover — matching a ~15-minute RTO at much lower cost than warm standby or active-active.

Q2. A Pilot Light design has been built for a three-tier application. During a failover test, instances launch successfully in the standby region but the application fails to start. What is the most likely cause?

A. The standby region does not support the instance type
B. The launch template references an AMI ID that exists only in the primary region, because AMIs are regional resources
C. Route 53 health checks have not yet marked the primary unhealthy
D. The Auto Scaling group has no scaling policy attached
Correct answer: B. AMIs and EBS snapshots are regional. A launch template in the standby region referencing a primary-region AMI ID will fail, and the fix is to copy the AMI into the standby region as part of the replication pipeline.

Q3. Which statement best distinguishes Warm Standby from Pilot Light?

A. Warm Standby replicates data; Pilot Light does not
B. Warm Standby keeps a scaled-down but running copy of the application tier; Pilot Light keeps only the data tier current and launches compute on demand
C. Warm Standby is cheaper than Pilot Light
D. Pilot Light has a lower RTO than Warm Standby
Correct answer: B. Both replicate data continuously. The discriminator is whether the application tier is already running in the standby region — Warm Standby yes, Pilot Light no.

Q4. A company states that its application must survive the loss of an entire AWS region. The current architecture runs across three Availability Zones in a single region with an RDS Multi-AZ instance. What is the gap?

A. Three AZs is insufficient; a fourth is required
B. Multi-AZ provides zone-level HA only and does nothing for a regional failure; a cross-region strategy is required
C. RDS Multi-AZ must be replaced with read replicas
D. Nothing — Multi-AZ across three AZs already survives a regional failure
Correct answer: B. Multi-AZ is a within-region high-availability mechanism. Surviving a regional failure requires one of the cross-region DR strategies, and Multi-AZ is complementary to it rather than a substitute.

Q5. A team's stated RPO is five minutes, but during peak write load the Aurora Global Database replication lag metric shows fifteen minutes. What does this mean for the design?

A. Nothing — the documented typical lag is sub-second, so the metric is wrong
B. The design does not meet its RPO requirement under peak load, and the lag metric must be monitored continuously rather than assumed
C. The RPO should be redefined as fifteen minutes
D. Aurora Global Database cannot be used for DR
Correct answer: B. Replication lag degrades under heavy write load. A design's real RPO is its worst observed lag, not the documented typical value, which is why a lag alarm is a required part of the monitoring.

Q6. A Pilot Light design assumes it can launch 200 EC2 instances in the standby region during failover. The failover test fails partway through with an instance launch error. What is the most likely cause?

A. The instances are in the wrong subnet
B. The standby region's EC2 On-Demand vCPU quota was never raised, because the region has never run at production scale
C. The AMI is encrypted
D. The security group does not allow outbound traffic
Correct answer: B. Service quotas are per-region and per-account. A standby region that has never run at production scale will have default quotas, and they must be raised ahead of time and verified periodically.

Q7. A Backup & Restore strategy stores recovery points in S3 Glacier Deep Archive to minimize cost. The business requires an RTO of four hours. What is the flaw?

A. Deep Archive cannot be used for backups
B. Deep Archive retrieval can take up to twelve hours, so the effective RTO exceeds the four-hour requirement
C. Deep Archive does not support cross-region replication
D. There is no flaw; retrieval time is not part of RTO
Correct answer: B. Retrieval time is part of the RTO. Storing recovery points in an archival class with multi-hour retrieval adds that time to the recovery window, which can silently violate the stated requirement.

Q8. A team has a documented DR runbook that has never been executed. What is the recommended next step before relying on it?

A. Trust the documentation as-is, since it was written by the architects
B. Run a scheduled failover exercise (Game Day) to validate that the runbook and automated recovery actually work, and to surface configuration drift and quota gaps
C. Increase backup frequency only
D. Skip testing to avoid production risk
Correct answer: B. Untested runbooks are unverified assumptions. A Game Day exercises the real failure mode against real infrastructure and finds the drift, expired certificates, and quota gaps while the stakes are low.

Q9. A scenario requires sub-minute failover that is not subject to client-side DNS caching delays. Which traffic-shift mechanism fits?

A. Route 53 failover routing with a 60-second TTL
B. AWS Global Accelerator, which uses static anycast IPs and shifts traffic at the AWS network edge
C. A single-region Network Load Balancer
D. CloudFront with a long cache TTL
Correct answer: B. Global Accelerator avoids DNS caching entirely by using static anycast IPs and rerouting at the AWS network edge, giving deterministic failover independent of client resolver behavior.

Q10. A cross-region read replica is promoted to primary during a regional failure. Which statement about this approach is accurate?

A. It provides zero data loss because replication is synchronous
B. It is a component of a DR strategy, not a complete strategy: promotion is manual, replication is asynchronous so some data is lost, and the promoted instance may not inherit all primary configuration
C. It automatically updates Route 53 to point at the new primary
D. It is equivalent to Multi-AZ failover
Correct answer: B. A read replica can be promoted, but promotion is manual, replication is asynchronous so there is data loss on promotion, and configuration inheritance is not guaranteed. It is one piece of a DR design.

Q11. A Pilot Light design replicates a high-write database continuously across regions. When comparing total cost against a Warm Standby design, what is often overlooked?

A. Cross-region data transfer cost, which scales with write volume and can exceed the cost of idle standby compute
B. The cost of the AMI copies
C. The cost of Route 53 health checks
D. Nothing; Pilot Light is always cheaper
Correct answer: A. Continuous cross-region replication is billed as data transfer out of the source region and scales with write volume. For a high-write workload this can invert the intuitive cost ordering between Pilot Light and Warm Standby.

Q12. A scenario describes a fully automated failover for a Pilot Light design. Which two halves must both be automated for that claim to be true?

A. Backup creation and backup retention
B. The traffic shift (DNS or Global Accelerator) and the compute launch or scaling in the standby region
C. The AMI copy and the security group rules
D. Only the DNS shift needs to be automated
Correct answer: B. Route 53 failover routing can automate the DNS shift, but the compute layer still has to be launched or scaled. A design that automates only the DNS half is not fully automated.

Q13. A nightly batch analytics pipeline can be re-run from source data if a region is lost. Which DR strategy is the cheapest sufficient choice?

A. Multi-Site Active-Active
B. Warm Standby
C. Backup & Restore, because the workload can be re-run and therefore tolerates a long RTO
D. Pilot Light with continuous replication
Correct answer: C. The ability to re-run the pipeline from source gives a naturally generous RTO, which makes Backup & Restore the cheapest sufficient strategy. The exam tests whether you resist over-engineering based on perceived importance.

Q14. The primary region is degraded but not down: latency is elevated and error rates are climbing, yet health checks still pass. What is the operational implication?

A. Automated failover will trigger correctly and no action is needed
B. Automated health-check-based failover will not trigger, so the design needs a manual override and a human judgment path
C. The health check interval should be increased
D. This situation cannot occur in AWS
Correct answer: B. Silent partial failure is the hardest case because health checks still pass. The design must include a manual failover override, and the decision to use it is a human judgment call.

Q15. A team fails over to the standby region successfully but then struggles for days to return to the primary. What was missing from their preparation?

A. A larger standby fleet
B. A rehearsed failback runbook, including resynchronizing the primary's now-stale data before it can resume serving
C. A second standby region
D. More frequent backups
Correct answer: B. Failback is a separate and often harder operation than failover, because the primary's data is stale and must be resynchronized. Teams that rehearse failover but not failback discover this during the incident.

Peek into Tomorrow

Everything in today's material rests on an assumption we have not yet examined: that the standby region is actually capable of taking over. Pilot Light and Warm Standby both describe what should be running in the standby region, but neither describes how you would know, before a disaster, whether that description is still true. Configuration drifts. Quotas get consumed by other workloads. An AMI gets deregistered. A certificate expires. The standby environment that passed its last failover test six months ago may be quietly broken today, and the first time you find out is the moment you need it.

There is also a subtler problem with the failover mechanism itself. Route 53 health-check-based failover depends on the Route 53 control plane, which is a global service but still a dependency. If the failover decision and the failover execution both live in the same place as the thing that failed, you have a circular dependency. Tomorrow's material addresses both gaps directly: readiness checks that continuously validate a standby region can actually take over, and routing controls that live on a control plane independent of any single region, giving deterministic and auditable failover rather than ad hoc health-check behavior. The question today leaves open is what "ready" should mean operationally — and how you make that judgment without a human in the loop.

Sources