Day 42 of 70 · Week 6
Day 42 / 70 Week 6 of 14 Phase 3: SRE Observability, Resilience & DR

Week Synthesis — SRE & DR Scenario Drills

🕑 ~28 min read · 4 services covered
FIS Route 53 ARC AWS Backup X-Ray

Recap — Week 6 in One Frame

Week 6 opened with the DR pattern spectrum — Backup & Restore, Pilot Light, Warm Standby, and Multi-Site Active-Active — and the observation that the choice is a cost/RTO trade rather than a technical preference. From there the week layered the mechanisms that make each point on that spectrum defensible: Route 53 ARC's readiness checks and routing controls for deterministic, auditable failover; AWS Backup with Vault Lock WORM immutability and cross-account copy for ransomware-resistant recovery; FIS experiment templates with mandatory stop conditions tied to CloudWatch alarms for controlled chaos; and X-Ray's service map for isolating the slow hop in a distributed call chain. The Well-Architected Reliability and Operational Excellence pillars supplied the vocabulary that ties them together — quantified RTO/RPO targets, tested recovery procedures, operations as code, and blameless post-incident review. This synthesis extends that arc rather than adding a new service: the exam rarely asks about one of these in isolation, it asks which combination satisfies a stated RTO, RPO, and budget.

Foundations You'll Need Today

Today's material is written in the vocabulary of disaster recovery, and that vocabulary is built on a handful of ideas that are easy to nod along to without actually being able to use them. Before the decision matrix makes sense, it's worth pinning down what these terms mean in plain language, because almost every question on this page is really a question about one of them.

RTO and RPO: The Two Numbers That Define a Disaster

When something breaks, two separate questions have to be answered, and people routinely blur them together. The first is "how long until we're serving customers again?" — that's the Recovery Time Objective, or RTO. It's a duration, measured from the moment of failure to the moment the system is back in business, and it's what determines whether you need a standby environment sitting there ready or whether you can afford to rebuild from scratch. The second question is "how much data are we willing to lose?" — that's the Recovery Point Objective, or RPO. It's also a duration, but it measures backwards from the failure: an RPO of five minutes means the most recent five minutes of data may be gone, because the last copy you have is five minutes old. A useful way to hold them apart is that RTO is about time-to-service and RPO is about time-to-last-good-copy. They are set by the business, not by engineering, and they are the constraints that every DR pattern on this page is trying to satisfy.

Regions and Availability Zones: Two Different Sizes of Failure

AWS runs its infrastructure in geographic groupings called Regions — Northern Virginia, Ireland, Singapore, and so on — and each Region is deliberately isolated from the others, with its own power, cooling, and network. Inside a Region there are Availability Zones, typically three or more, which are physically separate data centers within a short distance of each other. The distinction matters because the two failure domains are different sizes. Losing an Availability Zone is a routine event that a well-built application should absorb without anyone noticing, which is why "Multi-AZ" is the baseline for production databases. Losing an entire Region is rare but catastrophic, and it's the scenario that the whole DR pattern spectrum exists to handle. When this page says a workload is "active-active across two regions", it means the application is genuinely running and serving traffic in both, not that one is a copy of the other.

DNS, TTL, and Health Checks: How Traffic Finds the Healthy Copy

DNS is the phone book of the internet: it translates a human-readable name like app.example.com into the numeric IP address a client actually connects to. Route 53 is AWS's DNS service, and because it controls that translation, it can be used as a traffic director — return the IP of the healthy region and clients will go there. The catch is caching. Every DNS answer comes with a Time To Live, or TTL, which tells resolvers and client machines how long they may keep using that answer before asking again. A TTL of sixty seconds means that after a failover, some clients may keep dialing the dead region for up to a minute. A health check is the probe Route 53 runs against an endpoint to decide whether it's healthy enough to be handed out; when the probe fails, the record is withdrawn. So DNS-based failover works, but its speed is bounded by TTL and by how aggressively clients cache — which is exactly the problem the Global Accelerator and ARC discussion later on is solving.

Replication Lag: Why "Replicated" Doesn't Mean "Zero Data Loss"

Keeping a copy of your data in a second region means continuously shipping changes from the primary to the standby. If that shipping happens synchronously, the primary waits for the standby to confirm each write before telling the application it succeeded — which means no data loss, but also means every write is slower and the standby's availability now affects the primary. If it happens asynchronously, the primary confirms the write immediately and the standby catches up a moment later, which is fast and resilient but leaves a gap. That gap is the replication lag, and it is the thing that determines your real RPO: if the primary dies while the standby is thirty seconds behind, you have lost thirty seconds of writes. This is why the page treats replication lag as something to monitor and alarm on rather than an implementation detail — it is the RPO, measured live.

Distributed Tracing: Following One Request Across Many Services

In a modern application, a single user action — loading a checkout page, say — might pass through a load balancer, an API service, an authentication service, a database, and a payment provider before it returns. Each of those is a separate piece of software, and each can be slow or broken independently. Traditional monitoring tells you that the overall request is slow, but not which of the five steps is responsible, which is like knowing a package is late without knowing which depot it's stuck in. Distributed tracing solves this by tagging a request with a unique identifier at the front door and having every service it touches record its own start and end time under that identifier. The result is a single timeline showing each hop and how long it took — the "call chain" this page refers to — so a latency regression resolves to a specific service instead of a guess. That is what X-Ray produces, and it's why it appears in the observability section rather than alongside the metrics.

With that grounding, here's why the DR pattern spectrum exists and what problem each point on it actually solves.

The DR Pattern Decision Matrix

The single most reliable way to lose points on the resilience domain is to answer a DR scenario with the pattern that sounds most impressive rather than the one the stated RTO and RPO actually require. Every DR question on SAP-C02 is a constraint-satisfaction problem in disguise: the stem gives you a recovery time objective, a recovery point objective, and usually a cost signal ("minimize cost", "budget is not a constraint", "the business cannot tolerate more than X minutes"), and exactly one pattern on the spectrum satisfies all three. The work is to translate the stem's language into the spectrum's language before you look at the options.

The translation is mechanical once you internalize what each pattern actually keeps running. Backup & Restore keeps nothing running in the standby region — only backups exist, and recovery means provisioning infrastructure from templates and restoring data, which is why its RTO is measured in hours to days. Pilot Light keeps core data continuously replicated and a minimal skeleton of infrastructure (often just the database and network) but no application capacity, so recovery means scaling up the compute tier and pointing traffic at it — tens of minutes. Warm Standby keeps a scaled-down but fully functional copy of the entire stack running, so recovery is a scale-up plus a traffic shift — minutes. Multi-Site Active-Active serves live traffic from two or more regions simultaneously, so there is no recovery step at all, only the removal of the failed region from rotation — seconds, bounded by health-check detection and DNS or anycast propagation. The exam pattern is that the stem's RTO number maps almost directly onto one of those four bands, and the distractors are the adjacent bands.

PatternWhat runs in standbyTypical RTOTypical RPOCost posture
Backup & RestoreNothing; backups onlyHours to daysHours (last backup)Lowest standing cost
Pilot LightCore data replicated, minimal infraTens of minutesMinutes (replication lag)Low
Warm StandbyScaled-down full stackMinutesSeconds to minutesModerate
Multi-Site Active-ActiveFull production capacity, serving trafficSecondsNear zeroHighest

Two refinements matter for exam accuracy. First, RPO is governed by the data replication mechanism, not by the compute pattern — a Warm Standby with asynchronous cross-region replication still has a non-zero RPO equal to the replication lag, and a Pilot Light with synchronous replication can have a better RPO than a Warm Standby with async. Second, the pattern names describe compute posture; the data layer is chosen separately, which is why the same scenario can pair Pilot Light compute with Aurora Global Database storage and still be internally consistent.

Failover Control: Route 53 ARC vs Health Checks vs Global Accelerator

Once a pattern is chosen, the next question is what actually executes the failover, and this is where the week's services separate cleanly. Route 53 health checks with failover routing are the default mechanism and are adequate for most workloads: the health check probes an endpoint, and when it fails, Route 53 stops returning that record. The weakness is that the failover decision is only as good as the health check's view of the world, and the mechanism itself lives inside a single service in a single region's control plane. For a workload where a wrong or delayed failover is itself an outage, that is not enough.

Route 53 Application Recovery Controller addresses both weaknesses. Readiness checks continuously validate that the standby region can actually take over — that its Auto Scaling Group has the capacity, that its database replica is caught up, that its configuration matches production — so you discover a broken standby during a readiness audit rather than during an incident. Routing controls sit on a control plane deliberately distributed across regions and partitions, so the failover mechanism survives the failure it is responding to, and every routing change is auditable. The exam pattern is that any stem containing "the failover mechanism must not depend on a single region", "we need proof the standby is ready", or "auditable failover" is pointing at ARC, while a stem that just needs traffic to move away from an unhealthy endpoint is pointing at ordinary health checks.

Global Accelerator occupies a third position: it fails over at the network layer using static anycast IPs and AWS's global edge, which sidesteps the client-side DNS caching that can make Route 53-only failover slow for some clients. It is the right answer when the stem emphasizes sub-second failover, TCP/UDP workloads, or clients that cache DNS aggressively — and it is orthogonal to ARC, since a mature active-active design often uses Global Accelerator for the data path and ARC for the control decision.

MechanismFailover layerValidates standby readinessSurvives regional control-plane lossPick when…
Route 53 health checks + failover routingDNSNoNoStandard HA, DNS TTL acceptable
Route 53 ARCDNS, on a distributed control planeYes (readiness checks)YesCritical systems needing auditable, deterministic failover
Global AcceleratorNetwork (anycast)NoYes (edge-based)Sub-second failover, TCP/UDP, DNS-caching clients

Observability: Which Signal Answers Which Question

The observability services from this week are frequently confused with each other on the exam because they all "monitor" something, but each answers a structurally different question. CloudWatch metrics and alarms answer "is a known quantity outside its expected range" — they are threshold or anomaly-band detectors over numeric time series, and they are the only one of the group that can drive automated action directly through alarm actions. Composite alarms refine that by requiring correlated signals, so a page fires on high latency AND high error rate together rather than on either alone. Synthetics canaries answer a different question entirely: "does the user-facing flow work from outside our infrastructure", which catches DNS, CDN, and front-end failures that server-side metrics cannot see because the servers are healthy.

X-Ray answers "where inside the request did the time go", which is the only question of the four that requires distributed context. A service map built from traces shows per-hop latency and error rates across a call chain, so a p99 regression in a twelve-service application resolves to a specific downstream hop rather than to a guess. The exam pattern is that a stem describing a symptom visible in aggregate but not attributable to a component is an X-Ray question, while a stem describing a user-visible failure with healthy backend metrics is a Synthetics question. Logs Insights and subscription filters sit alongside these as the ad-hoc query and the real-time forwarding mechanism respectively, and they are the answer when the stem asks for correlation across log lines or for org-wide log aggregation.

SignalQuestion it answersBlind spotTypical exam trigger phrase
CloudWatch metrics + composite alarmsIs a known metric out of band, and do multiple signals agree?Cannot see user-facing flows or unknown failure modes"reduce alert noise", "page only on real incidents"
Synthetics canariesDoes the end-to-end user flow work from outside?Does not explain why it failed"customers report the page is broken but servers are healthy"
X-RayWhich hop in the call chain is slow or failing?Sampled; not a complete request census"p99 latency rising, unclear which service"
Logs Insights / subscription filtersWhat do the log lines say, and where should they be aggregated?Not a real-time alerting path by itself"centrally searchable logs", "SIEM ingestion"

Chaos Engineering and Backup: Proving the Design Works

The last two pieces of the week are the ones that convert a design from a claim into a verified fact, and they are complementary rather than overlapping. FIS and Game Days test the recovery path: they inject a real failure — instance termination, AZ loss, dependency failure — and observe whether the alarms fire, the runbooks are accurate, and the automated recovery completes inside the target RTO. AWS Backup protects the data that recovery depends on, and its distinguishing features are about surviving the failure of the account itself rather than the failure of a region. Cross-account copy into an isolated backup account means a compromised workload account cannot delete its own backups, and Vault Lock makes the retention policy immutable so that even a root user cannot shorten it before expiry.

The exam pattern here is a clean split on the stem's verb. If the stem says "test", "validate", "verify the runbook", or "prove the RTO is achievable", the answer is FIS or a Game Day. If the stem says "protect", "recover from ransomware", "immutable", or "cannot be deleted by an administrator", the answer is AWS Backup with cross-account copy and Vault Lock. A stem that says both — "we need to prove we can recover from a ransomware event within four hours" — is testing whether you recognize that the two services are used together, with Backup providing the immutable copy and FIS or a Game Day validating the restore procedure against the clock.

One operational detail worth carrying into the exam: FIS stop conditions are mandatory in the sense that a well-designed experiment always has them, and they are tied to CloudWatch alarms so the experiment aborts automatically when production stability degrades. A stem that describes chaos testing without a safety mechanism is describing a design flaw, and the correct answer will add stop conditions rather than remove the experiment.

Hands-on Lab / Practical Action (45 min)

Build a one-page DR decision worksheet and then pressure-test it against twenty scenarios. Start by writing the four patterns across the top of a table — Backup & Restore, Pilot Light, Warm Standby, Multi-Site Active-Active — and the four constraint columns down the side: stated RTO, stated RPO, cost signal, and failover-control requirement. Then work through the scenarios below, filling in the row for each one before you look at any answer key.

For each scenario, force yourself to write two sentences: which pattern satisfies the constraints, and which single constraint eliminated each of the other three. That second sentence is the part that transfers to the exam, because the distractors are always the adjacent patterns and the reason they are wrong is always a specific number in the stem. A scenario that says "RTO of fifteen minutes, minimize standing infrastructure cost" eliminates active-active on cost and eliminates backup/restore on RTO, leaving Pilot Light — and being able to say that out loud is worth more than memorizing the pattern list.

Next, layer the failover-control question on top. For each scenario you assigned to Warm Standby or Active-Active, decide whether ordinary Route 53 health checks are sufficient or whether the stem's language demands ARC readiness checks and routing controls. The tell is whether the stem treats a failed failover as catastrophic in its own right; if it does, ARC is in the answer. Then do the same for the observability layer: for each scenario, name the one signal you would alarm on and the one signal you would use to diagnose a failure after the fact. Metrics and composite alarms for the first, X-Ray for the second, Synthetics when the stem's symptom is user-visible but backend-healthy.

Finish by writing the backup and chaos clauses. For any scenario with a compliance, ransomware, or immutability requirement, add the AWS Backup configuration — cross-account copy plus Vault Lock — and state explicitly why a same-account backup is insufficient. For any scenario that claims an RTO, add the FIS experiment or Game Day that would prove it, including the stop condition tied to a CloudWatch alarm. The deliverable is a single page where every DR pattern is justified by a number from the stem and every claim of recoverability is backed by a named test. Keep it; it is the artifact you will re-read on Day 60 when the mock exam analysis sends you back to Domain 2.

Mixed Scenario Quiz — 25 Questions

Q1. A workload requires near-zero RTO/RPO and budget is not the primary constraint. Which DR strategy and supporting service pairing fits?

A. Backup & Restore with AWS Backup
B. Pilot Light with manual DNS updates
C. Multi-Site Active-Active with Route 53 ARC and DynamoDB Global Tables/Aurora Global Database
D. Warm standby without health checks
Correct answer: C. Near-zero RTO/RPO demands live traffic serving from multiple regions (active-active) with deterministic, tested failover control (ARC) and multi-region-write-capable data services.

Q2. A workload can tolerate an RTO of 15 minutes and an RPO of 5 minutes, but the business wants to minimize standing infrastructure cost. Which DR pattern fits best?

A. Backup and Restore
B. Pilot Light
C. Multi-Site Active-Active
D. No DR strategy is needed
Correct answer: B. Pilot Light keeps only core data continuously replicated and minimal infrastructure running, scaling up the rest on failover — matching a ~15-minute RTO at much lower cost than warm standby or active-active.

Q3. A critical financial system needs failover control that itself will not fail if an entire AWS region goes down, plus proof the standby region is actually ready to serve traffic. What should be used?

A. Route 53 simple health-check failover only
B. Route 53 Application Recovery Controller with readiness checks and a routing control cluster
C. CloudWatch Alarms triggering a Lambda failover
D. Global Accelerator alone
Correct answer: B. ARC's routing control cluster is deliberately distributed across regions and partitions for resilience of the failover mechanism itself, and readiness checks validate standby capacity before you rely on it.

Q4. A company needs ransomware-resilient backups that cannot be deleted even by a compromised administrator account. What should they configure?

A. Standard EBS snapshots only
B. AWS Backup with a cross-account copy into an isolated account, in a vault with Backup Vault Lock enabled
C. S3 versioning only
D. Increase snapshot frequency
Correct answer: B. Cross-account isolation prevents a compromised primary account from touching backups, and Vault Lock makes the retention policy immutable — even the root user cannot delete locked backups before expiry.

Q5. A team wants to safely test EC2 instance failure in production without risking an uncontrolled outage. What FIS feature guarantees the experiment halts if things go wrong?

A. IAM permission boundaries
B. Stop conditions tied to CloudWatch alarms
C. Increasing the blast radius
D. Manual monitoring only
Correct answer: B. FIS stop conditions automatically abort a running experiment the moment a linked CloudWatch alarm enters ALARM state, capping the blast radius of chaos testing.

Q6. A microservices application has growing p99 latency but it is unclear which of twelve services is the bottleneck. What AWS service pinpoints this?

A. CloudWatch Logs Insights alone
B. AWS X-Ray, using the service map to isolate the slow hop
C. VPC Flow Logs
D. AWS Config
Correct answer: B. X-Ray's distributed tracing and service map visualize per-hop latency across the full call chain, directly identifying the bottleneck service that aggregate metrics cannot isolate.

Q7. Internal CloudWatch metrics show healthy servers, but customers report the login page is broken. What monitoring gap does this reveal?

A. Missing X-Ray tracing
B. No outside-in synthetic monitoring of the actual user flow
C. Missing VPC Flow Logs
D. Insufficient EC2 instance count
Correct answer: B. Server-side health metrics do not verify end-to-end user experience; Synthetics canaries probe from outside the infrastructure, catching failures (DNS, CDN, front-end bugs) invisible to internal metrics.

Q8. On-call engineers are fatigued by alarms firing on isolated latency spikes that self-resolve. What reduces noise while still catching real incidents?

A. Delete the latency alarm
B. A composite alarm requiring both the latency alarm AND the error-rate alarm to be in ALARM state
C. Lower the alarm threshold
D. Increase the evaluation period to 24 hours
Correct answer: B. Composite alarms let you require correlated signals before paging, cutting single-metric false positives while preserving sensitivity to genuine multi-symptom incidents.

Q9. An active-active application needs sub-second failover at the network layer when a region's health checks fail, independent of DNS TTL and client caching. What fits?

A. Route 53 latency-based routing alone
B. AWS Global Accelerator, which uses static anycast IPs and reroutes at the AWS network edge
C. CloudFront alone
D. A single-region NLB
Correct answer: B. Global Accelerator uses anycast IPs and AWS's global network for near-instant failover, avoiding client-side DNS caching delays inherent to Route 53-only failover.

Q10. A team has documented DR runbooks but has never tested them under simulated failure. What is the recommended next step before relying on them?

A. Trust the documentation as-is
B. Run a Game Day exercise using FIS to simulate the failure and validate the runbook and automated recovery actually work
C. Increase backup frequency only
D. Skip testing to avoid production risk
Correct answer: B. Untested runbooks are unverified assumptions; a Game Day exercises the real failure mode against real infrastructure to confirm RTO/RPO targets are actually achievable.

Q11. Which best exemplifies the Reliability pillar's failure management best practice?

A. Using the largest possible instance type
B. Automatically testing recovery procedures via Game Days and setting quantified RTO/RPO targets
C. Manually reviewing logs weekly
D. Avoiding all managed services
Correct answer: B. Failure management is about anticipating failure, testing recovery (Game Days), and having quantified, validated RTO/RPO — not just provisioning bigger resources.

Q12. A team manually SSHes into servers to apply emergency patches, occasionally causing configuration drift. Which Operational Excellence practice addresses this?

A. Increase server count
B. Perform operations as code using SSM Automation documents and runbooks instead of manual SSH changes
C. Disable CloudTrail logging
D. Add more IAM users
Correct answer: B. Operations as code — codifying operational procedures such as SSM Automation — eliminates ad hoc manual changes and the drift and error they introduce.

Q13. An organization wants every account's application logs centrally searchable in a dedicated logging account in near-real time. What is the mechanism?

A. Manually export logs nightly via S3
B. CloudWatch Logs subscription filters streaming to Kinesis Data Firehose in the central logging account
C. CloudTrail only
D. Increase log retention in each account
Correct answer: B. Subscription filters push log events in near-real time to a destination (commonly Kinesis Firehose or Streams) in a centralized account, enabling org-wide log aggregation.

Q14. A workload requires near-zero RTO/RPO and budget is not the primary constraint. Which DR strategy and supporting service pairing fits?

A. Backup & Restore with AWS Backup
B. Pilot Light with manual DNS updates
C. Multi-Site Active-Active with Route 53 ARC and DynamoDB Global Tables or Aurora Global Database
D. Warm standby without health checks
Correct answer: C. Near-zero RTO/RPO demands live traffic serving from multiple regions with deterministic, tested failover control and multi-region-write-capable data services.

Q15. A team wants to prove that a regional failover completes inside the four-hour RTO the business has committed to, and that the standby database is caught up before the switch. Which pairing covers both halves?

A. AWS Backup Vault Lock plus CloudWatch Logs Insights
B. Route 53 ARC readiness checks plus an FIS experiment or Game Day that exercises the failover
C. X-Ray service map plus Synthetics canaries
D. Composite alarms plus Trusted Advisor
Correct answer: B. Readiness checks continuously validate that the standby can take over, and FIS or a Game Day proves the recovery path completes inside the committed RTO — the two halves of a verified DR claim.

Q16. A workload's RPO requirement is five minutes and its data is replicated asynchronously to a standby region. What does this imply about the achievable RPO?

A. RPO is zero because replication is continuous
B. RPO is bounded by the replication lag, which must be monitored and kept under five minutes
C. RPO is irrelevant once Multi-AZ is enabled
D. RPO is determined by the backup schedule only
Correct answer: B. With asynchronous replication the RPO equals the replication lag at the moment of failure, so the lag itself becomes an SLO that must be alarmed on — not an implementation detail.

Q17. A team wants to test how their application behaves when a dependency becomes slow rather than unavailable, without touching production data. What is the appropriate approach?

A. Take the dependency offline in production during business hours
B. Use FIS to inject latency or network stress in a controlled experiment with stop conditions tied to CloudWatch alarms
C. Rely on the next real incident to reveal the behavior
D. Add a composite alarm and wait
Correct answer: B. FIS supports network latency and stress actions as first-class experiment types, and stop conditions bound the blast radius so the experiment aborts if production stability degrades.

Q18. A company wants a single dashboard that shows, for a failing checkout request, which downstream call consumed the most time. Which service provides this directly?

A. CloudWatch metrics with high-resolution alarms
B. AWS X-Ray, using subsegments and annotations to drill into the specific downstream call
C. AWS Config conformance packs
D. VPC Flow Logs
Correct answer: B. X-Ray subsegments break a trace into individual downstream calls, and annotations let you filter traces by business attributes — the mechanism for attributing latency to a specific call rather than a service.

Q19. A compliance requirement states that backups must be retained for seven years and that no administrator may shorten that period. Which configuration satisfies this?

A. A lifecycle policy on the backup vault that administrators can edit
B. AWS Backup Vault Lock in compliance mode, which makes the retention policy immutable even to the root user
C. S3 Object Lock on a bucket that administrators own
D. A tag-based retention policy enforced by convention
Correct answer: B. Vault Lock in compliance mode enforces WORM immutability on the retention policy itself, so it cannot be shortened or removed before expiry — the property that makes it suitable for regulatory retention.

Q20. A team is choosing between Warm Standby and Multi-Site Active-Active for a workload with a two-minute RTO and a stated goal of minimizing cost. Which is correct and why?

A. Active-Active, because it has the lowest RTO
B. Warm Standby, because a two-minute RTO is achievable with a scaled-down running stack at materially lower cost than full duplicate capacity
C. Backup and Restore, because cost is the only constraint
D. Neither; the RTO is unachievable without active-active
Correct answer: B. The cheapest pattern that satisfies the stated RTO is the correct answer; a two-minute RTO is within Warm Standby's band, and active-active would be over-provisioning against the cost constraint.

Q21. An application's error rate spikes only during traffic peaks, and the team wants to be paged only when both latency and errors degrade together. What should they configure?

A. Two independent alarms with the same SNS topic
B. A composite alarm combining the latency and error-rate alarms with AND logic
C. A single alarm on CPU utilization
D. An anomaly detection band on request count
Correct answer: B. Composite alarms express the AND/OR relationship between underlying alarm states, which is exactly the "both signals degraded together" condition the team wants to page on.

Q22. A team needs to validate that their multi-AZ application survives the loss of an entire Availability Zone, and they want the exercise to abort automatically if customer-facing error rates climb. What should they use?

A. A manual failover performed by an on-call engineer
B. An FIS experiment template with an AZ availability action and a stop condition tied to the error-rate alarm
C. A CloudWatch dashboard review
D. A Route 53 health check with a shorter interval
Correct answer: B. FIS provides AZ availability disruption as an action, and the stop condition tied to the error-rate alarm is what makes the experiment safe to run against production.

Q23. A team's DR plan relies on restoring from backups, but they have never measured how long a restore actually takes. Which two actions best close this gap?

A. Increase backup frequency and add a second vault
B. Run a restore test on a schedule and record the elapsed time against the committed RTO
C. Add a composite alarm on backup job failures
D. Enable cross-region copy and assume the restore time is unchanged
Correct answer: B. Restore duration is the dominant term in a Backup & Restore RTO and is only knowable by measuring it; scheduled restore tests turn the RTO claim into a verified number.

Q24. A workload runs active-active across two regions. Which combination of services is required to make the data layer consistent with that posture?

A. A single-region Aurora cluster with read replicas
B. A multi-region-write-capable data service such as DynamoDB Global Tables, or Aurora Global Database with a deliberate single-writer region and conflict-aware partitioning
C. RDS Multi-AZ with a cross-region snapshot copy
D. ElastiCache for Redis as the system of record
Correct answer: B. Active-active compute requires a data layer that can accept writes in more than one region; DynamoDB Global Tables does this natively with last-writer-wins, while Aurora Global Database requires you to keep writes in one region or partition the write paths.

Q25. A team has a two-region active-passive design with Route 53 failover routing. During a real regional event, failover took far longer than expected because some clients kept hitting the failed region. What is the most likely cause and the best mitigation?

A. The health check interval was too short; lengthen it
B. Client-side DNS caching and TTL behavior delayed the switch; front the endpoints with Global Accelerator anycast IPs, or use ARC routing controls for a deterministic control-plane decision
C. The standby region lacked capacity; add more instances
D. The failover record was misconfigured; switch to simple routing
Correct answer: B. DNS-based failover is bounded by resolver and client caching, which is exactly the failure mode Global Accelerator's anycast IPs and ARC's distributed routing controls are designed to remove.

Preview — Day 43

Everything in this week assumed the workload was already running on AWS and the question was how to keep it running. The open question that Week 7 opens with is the one that comes before all of it: for a portfolio of workloads that are not on AWS yet, how do you decide what each one's move actually is? The 7 Rs — Retire, Retain, Rehost, Relocate, Repurchase, Replatform, Refactor — are the vocabulary for that decision, and the exam treats them as a classification problem with real consequences: a data center lease expiring in six months pushes most of a portfolio toward Rehost via MGN, while an application with unresolved licensing blockers is a Retain rather than a rushed migration. The trap is that the Rs sound like a maturity ladder, with Refactor at the top, when they are actually a set of trade-offs against deadline, cost, and technical debt. Day 43 also introduces Migration Hub as the tracking layer across tools, which matters because a portfolio migration is a program of work, not a single cutover.

Sources