Week Synthesis — SRE & DR Scenario Drills
Recap — Week 6 in One Frame
Week 6 opened with the DR pattern spectrum — Backup & Restore, Pilot Light, Warm Standby, and Multi-Site Active-Active — and the observation that the choice is a cost/RTO trade rather than a technical preference. From there the week layered the mechanisms that make each point on that spectrum defensible: Route 53 ARC's readiness checks and routing controls for deterministic, auditable failover; AWS Backup with Vault Lock WORM immutability and cross-account copy for ransomware-resistant recovery; FIS experiment templates with mandatory stop conditions tied to CloudWatch alarms for controlled chaos; and X-Ray's service map for isolating the slow hop in a distributed call chain. The Well-Architected Reliability and Operational Excellence pillars supplied the vocabulary that ties them together — quantified RTO/RPO targets, tested recovery procedures, operations as code, and blameless post-incident review. This synthesis extends that arc rather than adding a new service: the exam rarely asks about one of these in isolation, it asks which combination satisfies a stated RTO, RPO, and budget.
Foundations You'll Need Today
Today's material is written in the vocabulary of disaster recovery, and that vocabulary is built on a handful of ideas that are easy to nod along to without actually being able to use them. Before the decision matrix makes sense, it's worth pinning down what these terms mean in plain language, because almost every question on this page is really a question about one of them.
RTO and RPO: The Two Numbers That Define a Disaster
When something breaks, two separate questions have to be answered, and people routinely blur them together. The first is "how long until we're serving customers again?" — that's the Recovery Time Objective, or RTO. It's a duration, measured from the moment of failure to the moment the system is back in business, and it's what determines whether you need a standby environment sitting there ready or whether you can afford to rebuild from scratch. The second question is "how much data are we willing to lose?" — that's the Recovery Point Objective, or RPO. It's also a duration, but it measures backwards from the failure: an RPO of five minutes means the most recent five minutes of data may be gone, because the last copy you have is five minutes old. A useful way to hold them apart is that RTO is about time-to-service and RPO is about time-to-last-good-copy. They are set by the business, not by engineering, and they are the constraints that every DR pattern on this page is trying to satisfy.
Regions and Availability Zones: Two Different Sizes of Failure
AWS runs its infrastructure in geographic groupings called Regions — Northern Virginia, Ireland, Singapore, and so on — and each Region is deliberately isolated from the others, with its own power, cooling, and network. Inside a Region there are Availability Zones, typically three or more, which are physically separate data centers within a short distance of each other. The distinction matters because the two failure domains are different sizes. Losing an Availability Zone is a routine event that a well-built application should absorb without anyone noticing, which is why "Multi-AZ" is the baseline for production databases. Losing an entire Region is rare but catastrophic, and it's the scenario that the whole DR pattern spectrum exists to handle. When this page says a workload is "active-active across two regions", it means the application is genuinely running and serving traffic in both, not that one is a copy of the other.
DNS, TTL, and Health Checks: How Traffic Finds the Healthy Copy
DNS is the phone book of the internet: it translates a human-readable name like app.example.com into the numeric IP address a client actually connects to. Route 53 is AWS's DNS service, and because it controls that translation, it can be used as a traffic director — return the IP of the healthy region and clients will go there. The catch is caching. Every DNS answer comes with a Time To Live, or TTL, which tells resolvers and client machines how long they may keep using that answer before asking again. A TTL of sixty seconds means that after a failover, some clients may keep dialing the dead region for up to a minute. A health check is the probe Route 53 runs against an endpoint to decide whether it's healthy enough to be handed out; when the probe fails, the record is withdrawn. So DNS-based failover works, but its speed is bounded by TTL and by how aggressively clients cache — which is exactly the problem the Global Accelerator and ARC discussion later on is solving.
Replication Lag: Why "Replicated" Doesn't Mean "Zero Data Loss"
Keeping a copy of your data in a second region means continuously shipping changes from the primary to the standby. If that shipping happens synchronously, the primary waits for the standby to confirm each write before telling the application it succeeded — which means no data loss, but also means every write is slower and the standby's availability now affects the primary. If it happens asynchronously, the primary confirms the write immediately and the standby catches up a moment later, which is fast and resilient but leaves a gap. That gap is the replication lag, and it is the thing that determines your real RPO: if the primary dies while the standby is thirty seconds behind, you have lost thirty seconds of writes. This is why the page treats replication lag as something to monitor and alarm on rather than an implementation detail — it is the RPO, measured live.
Distributed Tracing: Following One Request Across Many Services
In a modern application, a single user action — loading a checkout page, say — might pass through a load balancer, an API service, an authentication service, a database, and a payment provider before it returns. Each of those is a separate piece of software, and each can be slow or broken independently. Traditional monitoring tells you that the overall request is slow, but not which of the five steps is responsible, which is like knowing a package is late without knowing which depot it's stuck in. Distributed tracing solves this by tagging a request with a unique identifier at the front door and having every service it touches record its own start and end time under that identifier. The result is a single timeline showing each hop and how long it took — the "call chain" this page refers to — so a latency regression resolves to a specific service instead of a guess. That is what X-Ray produces, and it's why it appears in the observability section rather than alongside the metrics.
With that grounding, here's why the DR pattern spectrum exists and what problem each point on it actually solves.
The DR Pattern Decision Matrix
The single most reliable way to lose points on the resilience domain is to answer a DR scenario with the pattern that sounds most impressive rather than the one the stated RTO and RPO actually require. Every DR question on SAP-C02 is a constraint-satisfaction problem in disguise: the stem gives you a recovery time objective, a recovery point objective, and usually a cost signal ("minimize cost", "budget is not a constraint", "the business cannot tolerate more than X minutes"), and exactly one pattern on the spectrum satisfies all three. The work is to translate the stem's language into the spectrum's language before you look at the options.
The translation is mechanical once you internalize what each pattern actually keeps running. Backup & Restore keeps nothing running in the standby region — only backups exist, and recovery means provisioning infrastructure from templates and restoring data, which is why its RTO is measured in hours to days. Pilot Light keeps core data continuously replicated and a minimal skeleton of infrastructure (often just the database and network) but no application capacity, so recovery means scaling up the compute tier and pointing traffic at it — tens of minutes. Warm Standby keeps a scaled-down but fully functional copy of the entire stack running, so recovery is a scale-up plus a traffic shift — minutes. Multi-Site Active-Active serves live traffic from two or more regions simultaneously, so there is no recovery step at all, only the removal of the failed region from rotation — seconds, bounded by health-check detection and DNS or anycast propagation. The exam pattern is that the stem's RTO number maps almost directly onto one of those four bands, and the distractors are the adjacent bands.
| Pattern | What runs in standby | Typical RTO | Typical RPO | Cost posture |
|---|---|---|---|---|
| Backup & Restore | Nothing; backups only | Hours to days | Hours (last backup) | Lowest standing cost |
| Pilot Light | Core data replicated, minimal infra | Tens of minutes | Minutes (replication lag) | Low |
| Warm Standby | Scaled-down full stack | Minutes | Seconds to minutes | Moderate |
| Multi-Site Active-Active | Full production capacity, serving traffic | Seconds | Near zero | Highest |
Two refinements matter for exam accuracy. First, RPO is governed by the data replication mechanism, not by the compute pattern — a Warm Standby with asynchronous cross-region replication still has a non-zero RPO equal to the replication lag, and a Pilot Light with synchronous replication can have a better RPO than a Warm Standby with async. Second, the pattern names describe compute posture; the data layer is chosen separately, which is why the same scenario can pair Pilot Light compute with Aurora Global Database storage and still be internally consistent.
Failover Control: Route 53 ARC vs Health Checks vs Global Accelerator
Once a pattern is chosen, the next question is what actually executes the failover, and this is where the week's services separate cleanly. Route 53 health checks with failover routing are the default mechanism and are adequate for most workloads: the health check probes an endpoint, and when it fails, Route 53 stops returning that record. The weakness is that the failover decision is only as good as the health check's view of the world, and the mechanism itself lives inside a single service in a single region's control plane. For a workload where a wrong or delayed failover is itself an outage, that is not enough.
Route 53 Application Recovery Controller addresses both weaknesses. Readiness checks continuously validate that the standby region can actually take over — that its Auto Scaling Group has the capacity, that its database replica is caught up, that its configuration matches production — so you discover a broken standby during a readiness audit rather than during an incident. Routing controls sit on a control plane deliberately distributed across regions and partitions, so the failover mechanism survives the failure it is responding to, and every routing change is auditable. The exam pattern is that any stem containing "the failover mechanism must not depend on a single region", "we need proof the standby is ready", or "auditable failover" is pointing at ARC, while a stem that just needs traffic to move away from an unhealthy endpoint is pointing at ordinary health checks.
Global Accelerator occupies a third position: it fails over at the network layer using static anycast IPs and AWS's global edge, which sidesteps the client-side DNS caching that can make Route 53-only failover slow for some clients. It is the right answer when the stem emphasizes sub-second failover, TCP/UDP workloads, or clients that cache DNS aggressively — and it is orthogonal to ARC, since a mature active-active design often uses Global Accelerator for the data path and ARC for the control decision.
| Mechanism | Failover layer | Validates standby readiness | Survives regional control-plane loss | Pick when… |
|---|---|---|---|---|
| Route 53 health checks + failover routing | DNS | No | No | Standard HA, DNS TTL acceptable |
| Route 53 ARC | DNS, on a distributed control plane | Yes (readiness checks) | Yes | Critical systems needing auditable, deterministic failover |
| Global Accelerator | Network (anycast) | No | Yes (edge-based) | Sub-second failover, TCP/UDP, DNS-caching clients |
Observability: Which Signal Answers Which Question
The observability services from this week are frequently confused with each other on the exam because they all "monitor" something, but each answers a structurally different question. CloudWatch metrics and alarms answer "is a known quantity outside its expected range" — they are threshold or anomaly-band detectors over numeric time series, and they are the only one of the group that can drive automated action directly through alarm actions. Composite alarms refine that by requiring correlated signals, so a page fires on high latency AND high error rate together rather than on either alone. Synthetics canaries answer a different question entirely: "does the user-facing flow work from outside our infrastructure", which catches DNS, CDN, and front-end failures that server-side metrics cannot see because the servers are healthy.
X-Ray answers "where inside the request did the time go", which is the only question of the four that requires distributed context. A service map built from traces shows per-hop latency and error rates across a call chain, so a p99 regression in a twelve-service application resolves to a specific downstream hop rather than to a guess. The exam pattern is that a stem describing a symptom visible in aggregate but not attributable to a component is an X-Ray question, while a stem describing a user-visible failure with healthy backend metrics is a Synthetics question. Logs Insights and subscription filters sit alongside these as the ad-hoc query and the real-time forwarding mechanism respectively, and they are the answer when the stem asks for correlation across log lines or for org-wide log aggregation.
| Signal | Question it answers | Blind spot | Typical exam trigger phrase |
|---|---|---|---|
| CloudWatch metrics + composite alarms | Is a known metric out of band, and do multiple signals agree? | Cannot see user-facing flows or unknown failure modes | "reduce alert noise", "page only on real incidents" |
| Synthetics canaries | Does the end-to-end user flow work from outside? | Does not explain why it failed | "customers report the page is broken but servers are healthy" |
| X-Ray | Which hop in the call chain is slow or failing? | Sampled; not a complete request census | "p99 latency rising, unclear which service" |
| Logs Insights / subscription filters | What do the log lines say, and where should they be aggregated? | Not a real-time alerting path by itself | "centrally searchable logs", "SIEM ingestion" |
Chaos Engineering and Backup: Proving the Design Works
The last two pieces of the week are the ones that convert a design from a claim into a verified fact, and they are complementary rather than overlapping. FIS and Game Days test the recovery path: they inject a real failure — instance termination, AZ loss, dependency failure — and observe whether the alarms fire, the runbooks are accurate, and the automated recovery completes inside the target RTO. AWS Backup protects the data that recovery depends on, and its distinguishing features are about surviving the failure of the account itself rather than the failure of a region. Cross-account copy into an isolated backup account means a compromised workload account cannot delete its own backups, and Vault Lock makes the retention policy immutable so that even a root user cannot shorten it before expiry.
The exam pattern here is a clean split on the stem's verb. If the stem says "test", "validate", "verify the runbook", or "prove the RTO is achievable", the answer is FIS or a Game Day. If the stem says "protect", "recover from ransomware", "immutable", or "cannot be deleted by an administrator", the answer is AWS Backup with cross-account copy and Vault Lock. A stem that says both — "we need to prove we can recover from a ransomware event within four hours" — is testing whether you recognize that the two services are used together, with Backup providing the immutable copy and FIS or a Game Day validating the restore procedure against the clock.
One operational detail worth carrying into the exam: FIS stop conditions are mandatory in the sense that a well-designed experiment always has them, and they are tied to CloudWatch alarms so the experiment aborts automatically when production stability degrades. A stem that describes chaos testing without a safety mechanism is describing a design flaw, and the correct answer will add stop conditions rather than remove the experiment.
Hands-on Lab / Practical Action (45 min)
Build a one-page DR decision worksheet and then pressure-test it against twenty scenarios. Start by writing the four patterns across the top of a table — Backup & Restore, Pilot Light, Warm Standby, Multi-Site Active-Active — and the four constraint columns down the side: stated RTO, stated RPO, cost signal, and failover-control requirement. Then work through the scenarios below, filling in the row for each one before you look at any answer key.
For each scenario, force yourself to write two sentences: which pattern satisfies the constraints, and which single constraint eliminated each of the other three. That second sentence is the part that transfers to the exam, because the distractors are always the adjacent patterns and the reason they are wrong is always a specific number in the stem. A scenario that says "RTO of fifteen minutes, minimize standing infrastructure cost" eliminates active-active on cost and eliminates backup/restore on RTO, leaving Pilot Light — and being able to say that out loud is worth more than memorizing the pattern list.
Next, layer the failover-control question on top. For each scenario you assigned to Warm Standby or Active-Active, decide whether ordinary Route 53 health checks are sufficient or whether the stem's language demands ARC readiness checks and routing controls. The tell is whether the stem treats a failed failover as catastrophic in its own right; if it does, ARC is in the answer. Then do the same for the observability layer: for each scenario, name the one signal you would alarm on and the one signal you would use to diagnose a failure after the fact. Metrics and composite alarms for the first, X-Ray for the second, Synthetics when the stem's symptom is user-visible but backend-healthy.
Finish by writing the backup and chaos clauses. For any scenario with a compliance, ransomware, or immutability requirement, add the AWS Backup configuration — cross-account copy plus Vault Lock — and state explicitly why a same-account backup is insufficient. For any scenario that claims an RTO, add the FIS experiment or Game Day that would prove it, including the stop condition tied to a CloudWatch alarm. The deliverable is a single page where every DR pattern is justified by a number from the stem and every claim of recoverability is backed by a named test. Keep it; it is the artifact you will re-read on Day 60 when the mock exam analysis sends you back to Domain 2.
Mixed Scenario Quiz — 25 Questions
Q1. A workload requires near-zero RTO/RPO and budget is not the primary constraint. Which DR strategy and supporting service pairing fits?
Q2. A workload can tolerate an RTO of 15 minutes and an RPO of 5 minutes, but the business wants to minimize standing infrastructure cost. Which DR pattern fits best?
Q3. A critical financial system needs failover control that itself will not fail if an entire AWS region goes down, plus proof the standby region is actually ready to serve traffic. What should be used?
Q4. A company needs ransomware-resilient backups that cannot be deleted even by a compromised administrator account. What should they configure?
Q5. A team wants to safely test EC2 instance failure in production without risking an uncontrolled outage. What FIS feature guarantees the experiment halts if things go wrong?
Q6. A microservices application has growing p99 latency but it is unclear which of twelve services is the bottleneck. What AWS service pinpoints this?
Q7. Internal CloudWatch metrics show healthy servers, but customers report the login page is broken. What monitoring gap does this reveal?
Q8. On-call engineers are fatigued by alarms firing on isolated latency spikes that self-resolve. What reduces noise while still catching real incidents?
Q9. An active-active application needs sub-second failover at the network layer when a region's health checks fail, independent of DNS TTL and client caching. What fits?
Q10. A team has documented DR runbooks but has never tested them under simulated failure. What is the recommended next step before relying on them?
Q11. Which best exemplifies the Reliability pillar's failure management best practice?
Q12. A team manually SSHes into servers to apply emergency patches, occasionally causing configuration drift. Which Operational Excellence practice addresses this?
Q13. An organization wants every account's application logs centrally searchable in a dedicated logging account in near-real time. What is the mechanism?
Q14. A workload requires near-zero RTO/RPO and budget is not the primary constraint. Which DR strategy and supporting service pairing fits?
Q15. A team wants to prove that a regional failover completes inside the four-hour RTO the business has committed to, and that the standby database is caught up before the switch. Which pairing covers both halves?
Q16. A workload's RPO requirement is five minutes and its data is replicated asynchronously to a standby region. What does this imply about the achievable RPO?
Q17. A team wants to test how their application behaves when a dependency becomes slow rather than unavailable, without touching production data. What is the appropriate approach?
Q18. A company wants a single dashboard that shows, for a failing checkout request, which downstream call consumed the most time. Which service provides this directly?
Q19. A compliance requirement states that backups must be retained for seven years and that no administrator may shorten that period. Which configuration satisfies this?
Q20. A team is choosing between Warm Standby and Multi-Site Active-Active for a workload with a two-minute RTO and a stated goal of minimizing cost. Which is correct and why?
Q21. An application's error rate spikes only during traffic peaks, and the team wants to be paged only when both latency and errors degrade together. What should they configure?
Q22. A team needs to validate that their multi-AZ application survives the loss of an entire Availability Zone, and they want the exercise to abort automatically if customer-facing error rates climb. What should they use?
Q23. A team's DR plan relies on restoring from backups, but they have never measured how long a restore actually takes. Which two actions best close this gap?
Q24. A workload runs active-active across two regions. Which combination of services is required to make the data layer consistent with that posture?
Q25. A team has a two-region active-passive design with Route 53 failover routing. During a real regional event, failover took far longer than expected because some clients kept hitting the failed region. What is the most likely cause and the best mitigation?
Preview — Day 43
Everything in this week assumed the workload was already running on AWS and the question was how to keep it running. The open question that Week 7 opens with is the one that comes before all of it: for a portfolio of workloads that are not on AWS yet, how do you decide what each one's move actually is? The 7 Rs — Retire, Retain, Rehost, Relocate, Repurchase, Replatform, Refactor — are the vocabulary for that decision, and the exam treats them as a classification problem with real consequences: a data center lease expiring in six months pushes most of a portfolio toward Rehost via MGN, while an application with unresolved licensing blockers is a Retain rather than a rushed migration. The trap is that the Rs sound like a maturity ladder, with Refactor at the top, when they are actually a set of trade-offs against deadline, cost, and technical debt. Day 43 also introduces Migration Hub as the tracking layer across tools, which matters because a portfolio migration is a program of work, not a single cutover.