Designing Game Days & Chaos Experiments
Recap: From Experiment Templates to Exercises
Day 33 established the mechanics of AWS Fault Injection Service: experiment templates that declare an action, a target selection mode, and a blast-radius percentage, plus mandatory stop conditions tied to CloudWatch alarms that abort the run the moment production stability degrades. That lesson covered the primitives — EC2 and ECS task termination, CPU and network stress, and the AZ failover experiment that forces an RDS or Aurora failover — and treated each as an isolated capability you can invoke on demand.
Today extends that material rather than restating it. The primitives are the vocabulary; a Game Day is the sentence. An experiment template answers "what can I break and how do I stop it," while a Game Day answers "who is in the room, what are we trying to prove, what counts as success, and what do we change afterward." The distinction matters because the exam rarely asks you to recite an FIS action name. It asks you to recognize a scenario where a team has documented recovery procedures but has never validated them, and to select the practice that converts those assumptions into measured facts. That practice is the Game Day, and it is built on exactly the stop-condition discipline Day 33 described.
Foundations You'll Need Today
Today's lesson is about deliberately breaking things on purpose, which only makes sense once you know what the things are and what "broken" means for each of them. Five ideas carry the whole discussion, and none of them require hands-on experience to understand — just a clear picture of what each one is for.
Availability Zones: the unit of failure
An AWS Region is a geographic area, and inside each Region AWS operates several physically separate data centers grouped into Availability Zones, usually three or more. The zones have independent power, cooling, and network connections, so a fire, flood, or power failure in one zone does not take down the others. This separation is the entire reason AWS can promise high availability: you spread your application across multiple zones, and if one zone dies, the others keep serving. When today's lesson talks about "losing an AZ," it means exactly that — one of those physically separate groups of data centers becomes unreachable, and everything you placed inside it stops responding at once. That is the failure mode most AWS resilience design is built around, which is why it is the most common thing a Game Day simulates.
Load balancers and Auto Scaling Groups: the two things that react
An Application Load Balancer sits in front of your application and distributes incoming requests across a pool of servers, called targets. Crucially, it continuously health-checks those targets and stops sending traffic to any that fail — a process called draining. An Auto Scaling Group is the mechanism that keeps a desired number of servers running: it watches how many are healthy and launches replacements when some disappear. Together they form the automatic recovery path that most AWS architectures rely on. When today's hypothesis says the load balancer will "drain the affected targets within 30 seconds" and the scaling group will "launch replacement capacity within 5 minutes," it is describing these two mechanisms doing their jobs. A Game Day exists to find out whether they actually do, and how fast.
Multi-AZ databases and failover
A database is the hardest part of an application to recover, because unlike a stateless web server it holds data that cannot simply be recreated. AWS offers a Multi-AZ configuration for its managed databases: a second copy of the database runs in a different Availability Zone, kept continuously in sync, and if the primary copy fails, AWS promotes the standby to become the new primary. That promotion is called a failover, and it takes time — typically under two minutes, but not zero. Applications must also reconnect to the new primary, which is why failover is a common source of surprises. Today's exercises treat database failover as one of the failure modes worth measuring, because the gap between the documented failover time and the real one is often where an RTO claim quietly breaks.
RTO and RPO: the two numbers everything is measured against
Recovery Time Objective, or RTO, is how long the business can tolerate the service being down before the damage becomes unacceptable — "we must be back within fifteen minutes." Recovery Point Objective, or RPO, is how much data the business can afford to lose, expressed as time — "we can accept losing the last five minutes of transactions." These are business decisions, not technical ones, and they are the yardstick for everything in this lesson. Without a stated RTO there is no way to say whether a recovery was fast enough, which is why today's lab insists you write one down before designing the exercise. A Game Day is fundamentally a test of whether the architecture actually meets the RTO and RPO it claims.
CloudWatch alarms and stop conditions: the safety net
Amazon CloudWatch is AWS's monitoring service: it collects numeric measurements from your resources, called metrics, and an alarm is a rule that watches one metric and changes state when it crosses a threshold you set — for example, "alert me if the error rate exceeds 0.5% for two consecutive minutes." An alarm can be in one of three states: OK, ALARM, or INSUFFICIENT_DATA, the last meaning it has not yet received enough measurements to judge. A stop condition is simply an alarm wired into a failure-injection experiment so that if the alarm enters ALARM, the experiment halts automatically. This is what makes it defensible to break things in production: the moment the damage exceeds what you agreed to accept, the experiment stops itself rather than waiting for a human to notice. Keep the INSUFFICIENT_DATA state in mind, because today's lesson treats an alarm stuck in that state as a safety net that is not actually there.
With that grounding, here is why Game Days exist and what problem they actually solve.
1. Why Game Days Are on the Exam
The SAP-C02 exam tests design judgment under constraint, and one of the most reliable constraints it uses is the gap between what an architecture claims and what it has actually been proven to do. A team can draw a three-AZ, multi-region diagram with automatic failover and still have a recovery time objective that exists only in a slide deck. The exam's resilience domain is built around that gap, and Game Days are the mechanism AWS documents for closing it. When a question describes a workload with a stated RTO, a documented runbook, and no evidence the runbook has ever been executed, the correct answer is almost always to test the recovery path rather than to add more redundancy.
This maps most directly to the Resilient Architectures domain, but it also touches Operational Excellence, because a Game Day is fundamentally an operations-as-code practice: the exercise is scripted, the success criteria are quantified in advance, and the output is a set of corrective actions rather than a vague sense that things went fine. The Well-Architected Reliability pillar names failure management as one of its four areas, and within failure management the guidance is explicit that recovery procedures must be tested, not merely written. A question that offers "increase backup frequency" or "add a second standby region" as an answer to an untested-runbook scenario is testing whether you understand that more infrastructure does not substitute for validated procedure.
There is a second, subtler exam angle. Game Days are the point where several earlier topics converge: the composite alarms from Day 29 determine whether the exercise produces signal or noise, the synthetic canaries from Day 30 determine whether you detect the failure from the customer's perspective, and the FIS stop conditions from Day 33 determine whether the exercise stays safe. A scenario that asks how to validate an SLO end to end is often really asking whether you would run a controlled failure and measure the detection and recovery path, rather than whether you would add another dashboard.
2. How a Game Day Actually Runs
A Game Day is a scheduled, time-boxed exercise in which a team deliberately introduces a realistic failure into a system and observes whether the documented detection and recovery mechanisms behave as designed. The word "scheduled" carries weight: unlike an ad hoc chaos experiment run by one engineer to see what happens, a Game Day has a defined start and end, named participants, a written hypothesis, and pre-agreed success criteria. The hypothesis is the part teams skip and the part that makes the exercise useful. "We believe that if we lose one Availability Zone, the Application Load Balancer will drain the affected targets within 60 seconds, the Auto Scaling Group will replace capacity in the surviving zones, and customer-facing error rate will stay below 0.5% for the duration" is a hypothesis. "Let's kill an AZ and see what happens" is not.
Mechanically, the exercise runs in four phases. In the planning phase the team selects a failure mode, writes the hypothesis, defines the success criteria as measurable thresholds, identifies the stop conditions that will abort the run, and notifies anyone whose systems might be affected. In the execution phase the failure is injected — typically through an FIS experiment template, though a Game Day can also be run by manually failing over a database or revoking a credential — while participants watch the same dashboards an on-call engineer would watch. In the observation phase the team records what actually happened against what was predicted: time to detection, time to first alarm, time to automated recovery, time to manual intervention if any, and the peak customer-visible impact. In the retrospective phase the deltas between prediction and reality become backlog items.
The critical structural property is that the failure injection is bounded by the same stop conditions that govern any FIS experiment. If the error-rate alarm breaches its threshold, the experiment halts automatically rather than waiting for a human to notice. This is what makes it defensible to run a Game Day against production, which is where the exercise has value. A Game Day run only in a staging environment validates the runbook against staging's topology, staging's data volume, and staging's traffic patterns — none of which match production. The exam consistently favors testing in production with bounded blast radius over testing in an environment that does not reproduce the failure mode.
3. The Core Decision Boundary: What Are You Trying to Prove?
Every Game Day design question reduces to a single fork: is the exercise validating a detection path, a recovery path, or a decision path? The three are different exercises with different failure modes, and conflating them produces an exercise that proves nothing. A detection exercise asks whether the monitoring stack notices the failure and pages the right person within the target window. A recovery exercise asks whether the automated or documented remediation actually restores service within the RTO. A decision exercise asks whether the humans in the loop make the right call — fail over or wait, escalate or absorb — under time pressure and incomplete information.
The fork matters because the success criteria are completely different. A detection exercise can succeed even if recovery is slow, as long as the alarm fired and the page reached the on-call engineer. A recovery exercise can succeed even if detection was late, as long as the system returned to health within the RTO. A decision exercise is the hardest to score because the "correct" call depends on information that may not have been available at the time, and the retrospective has to evaluate the decision against the information the responder actually had rather than against the outcome.
In practice most Game Days target recovery, because that is where the RTO claim lives and where the gap between documentation and reality is widest. But the exam will present scenarios where the real weakness is detection — a workload with excellent automated failover and no alarm that fires when it happens — and the correct answer is to exercise the detection path first. The table below maps the fork to its observable signals.
| Exercise type | Primary question | Success signal | Common failure |
|---|---|---|---|
| Detection | Did we notice, and did the right person get paged? | Alarm state change and page delivery within the target window | Alarm exists but routes to a queue nobody watches |
| Recovery | Did service return within the RTO? | Measured time from injection to healthy state | Runbook step references a console path that no longer exists |
| Decision | Did the responder choose the right action given available information? | Documented rationale matching the runbook's decision tree | Runbook has no decision tree, only a happy path |
| Degradation | Does the system shed load gracefully instead of failing hard? | Reduced functionality with no hard errors | No queue or backpressure, so the failure cascades |
4. Scoping the Exercise: Blast Radius, Environment, and Timing
The knobs that determine whether a Game Day is safe and useful are blast radius, environment, and timing, and each trades safety against fidelity. Blast radius is the fraction of the fleet or the number of targets the injection touches. A single-task termination proves that the scheduler replaces a task; it does not prove that the service survives losing a third of its capacity. A full AZ failure proves the latter but consumes real capacity and real customer impact. The right setting is the smallest blast radius that still exercises the mechanism you are testing — if the hypothesis is about AZ-level failover, you need AZ-level scope, and no smaller injection will validate it.
Environment is the second knob, and the tradeoff is stark. Staging is safe and reproducible but rarely reproduces production's data volume, traffic distribution, dependency graph, or configuration drift. Production is faithful but carries real risk, which is precisely why the stop conditions matter. The defensible position, and the one the exam rewards, is to run in production with a bounded blast radius and automated abort conditions, and to reserve staging for rehearsing the mechanics of the injection itself before the real exercise. Teams that only ever test in staging tend to discover during a real incident that the runbook's assumptions about production topology were never true.
Timing is the third knob and the one most often mishandled. Running a Game Day during peak traffic maximizes fidelity but also maximizes the cost of a mistake. Running it at 3 a.m. minimizes customer impact but also means the on-call engineer who would respond to a real incident is asleep, so the exercise tests the automated path only and says nothing about the human path. The common compromise is to run during business hours with the full response team present and available, accepting a small amount of customer-visible risk in exchange for testing the complete detection-to-recovery chain including the humans. The table below summarizes the tradeoffs.
| Knob | Low setting | High setting | What it buys you |
|---|---|---|---|
| Blast radius | One task or instance | Full AZ or region | Fidelity to the failure mode you claim to survive |
| Environment | Staging | Production | Real topology, data volume, and configuration drift |
| Timing | Off-hours | Peak business hours | Whether the human response path works, not just the automated one |
| Duration | Minutes | Hours | Whether slow-burn failures (connection pool exhaustion, disk fill) surface |
5. Sizing the Exercise and the Numbers That Bound It
Game Day sizing is really RTO and RPO arithmetic made concrete. If a workload's stated RTO is fifteen minutes, the exercise must be long enough to observe whether recovery completes inside that window, which means the observation period has to extend past the RTO rather than stopping the moment the injection ends. A common mistake is to declare success when the alarm clears, without measuring the full path from injection to steady-state health. The measurement that matters is wall-clock time from the moment the failure is introduced to the moment the service is serving normal traffic at normal error rates, and that number is what gets compared against the RTO.
The FIS experiment itself is bounded by its stop conditions, and those thresholds should be set tighter than the SLO breach point, not equal to it. If the service's error-rate SLO allows 1% errors over a five-minute window, the stop condition should trip well before that — the exercise is meant to reveal weakness, not to consume the error budget. Similarly, the blast radius percentage in the target selection should be the minimum that exercises the mechanism. FIS supports targeting by tag, by resource ID, and by percentage of a matched set, and the percentage mode is the one that makes an exercise reproducible across a fleet whose size changes between runs.
Duration is bounded by the failure mode. Fast failures — instance termination, task kill, AZ failover — complete their recovery within minutes and the exercise can be short. Slow failures — memory leaks, connection pool exhaustion, disk fill, certificate expiry — need the exercise to run long enough for the symptom to develop, which can mean hours. The table below lists the failure modes most commonly exercised and the observation window each requires.
| Failure mode | Typical observation window | What you are measuring |
|---|---|---|
| Single task or instance termination | Minutes | Scheduler replacement time and load balancer drain behavior |
| AZ failure (subnet or AZ-scoped injection) | 10-30 minutes | Cross-AZ failover, capacity replacement, connection re-establishment |
| Database failover (Multi-AZ or Aurora) | Minutes | Endpoint DNS propagation and application reconnect logic |
| Dependency latency injection | 15-60 minutes | Timeout configuration, retry storms, circuit breaker behavior |
| Resource exhaustion (disk, connections) | Hours | Whether alarms fire before the resource is fully consumed |
6. Failure Modes of the Exercise Itself
The most common way a Game Day fails is that it proves nothing because the injection never actually reached the system under test. A target selection that matches zero resources, an IAM role that lacks permission to perform the action, or an experiment scoped to a tag that no longer exists all produce a run that reports success while nothing happened. The symptom is an exercise where every metric stayed flat and the retrospective concludes the system is resilient, when in fact the failure was never introduced. The first diagnostic move is always to confirm the injection landed: check the FIS experiment's action log and verify the target count is what you expected before interpreting any other signal.
The second failure mode is the opposite — the injection lands but the stop conditions do not fire when they should, because the alarm they reference is misconfigured, evaluating the wrong metric, or in INSUFFICIENT_DATA state. An alarm that has never had enough datapoints to leave INSUFFICIENT_DATA cannot trigger a stop condition, which means the safety net is absent precisely during the exercise that needs it. Before any Game Day, the stop-condition alarms should be verified as being in OK state with a healthy evaluation history, not merely present in the template.
The third failure mode is a recovery that works but for the wrong reason. If the service recovered because a human manually scaled the fleet rather than because the Auto Scaling Group did it, the exercise has validated the human, not the automation, and the runbook's claim about automated recovery is still unproven. Distinguishing these requires the observation phase to record who or what performed each recovery action, which is why the exercise should be run with the response team watching rather than silently. The table below maps symptoms to first diagnostic moves.
| Symptom | Likely cause | First diagnostic move |
|---|---|---|
| All metrics flat during the exercise | Injection never reached a target | Check FIS action log and target resource count |
| Stop condition never tripped despite visible impact | Alarm in INSUFFICIENT_DATA or wrong metric | Inspect alarm state history and metric namespace |
| Recovery succeeded but no automation ran | Human intervention masked the gap | Review CloudTrail for manual API calls during the window |
| Recovery succeeded in staging, failed in production | Topology or data-volume divergence | Compare target group membership and connection counts |
7. The SRE Angle: Game Days as Error-Budget Spend
From an SRE perspective a Game Day is a deliberate expenditure of error budget in exchange for information, and it should be budgeted and scheduled like any other planned risk. The exercise consumes some fraction of the service's allowed unreliability, and the return is a measured recovery time, a validated runbook, and a list of corrective actions. Framing it this way makes the go/no-go decision tractable: if the service has ample error budget remaining, the exercise is cheap; if the budget is nearly exhausted, the exercise should be deferred or run at a smaller blast radius, because the team cannot afford the additional unreliability on top of whatever is already consuming the budget.
The observability requirements for a Game Day are the same as for an incident, which is a useful forcing function. If the team cannot tell during the exercise whether the service is healthy, they will not be able to tell during a real incident either, and that discovery is itself a valuable output. The minimum instrumentation is a customer-facing availability signal (synthetic canary or load balancer error rate), a latency signal at the percentile that matters, a saturation signal for the constrained resource, and an alarm on each that is known to be in OK state before the exercise begins. Composite alarms are particularly useful here because they let the exercise's stop condition require correlated signals rather than tripping on a single noisy metric.
The runbook shape that emerges from a well-run Game Day is a decision tree rather than a checklist. The exercise reveals the branch points — at what error rate do we fail over, at what latency do we shed load, who has authority to declare the incident — and those branches get written down with the thresholds that were validated. The retrospective output is not "the system worked" but a set of specific deltas: the alarm fired four minutes later than predicted, the runbook's failover step referenced a console page that has moved, the connection pool needed a manual restart that no step mentioned. Each delta becomes a backlog item with an owner.
8. Edge Cases and Exam Gotchas
The single most-tested gotcha is the substitution trap: a scenario describes an untested recovery procedure and offers "increase backup frequency," "add a second standby region," or "enable Multi-AZ" as answers. None of these validate the procedure. More redundancy changes the recovery characteristics but does not prove the runbook works, and the exam expects you to select the answer that tests rather than the answer that adds. A related trap offers "run the exercise in a staging environment" as the safe choice; it is safer but it does not reproduce production's failure mode, and the exam generally prefers bounded production testing.
The second gotcha concerns stop conditions. A question may describe an FIS experiment with no stop condition, or with a stop condition referencing an alarm that has never fired, and ask what the risk is. The answer is that the experiment has no automated abort and will continue injecting failure regardless of impact, which is exactly the scenario Day 33's stop-condition discipline exists to prevent. A third gotcha is the difference between a Game Day and a chaos experiment run ad hoc: the Game Day's defining features are the pre-written hypothesis, the quantified success criteria, and the scheduled participation of the response team. An experiment without those is a test, not a Game Day, and the exam distinguishes them.
Finally, watch for scenarios that conflate the exercise with the remediation. A Game Day that reveals a gap does not fix the gap; the corrective actions do. An answer choice that says "run a Game Day to resolve the issue" is wrong if the issue is a known defect — the Game Day would only confirm what is already known. The exercise is for discovering unknown gaps and validating claimed capabilities, not for closing defects you have already identified.
9. Game Days vs. the Practices They Get Confused With
Game Days sit in a family of resilience practices that are easy to conflate: chaos experiments, load tests, disaster recovery drills, and penetration tests. The distinctions matter on the exam because each answers a different question and produces a different artifact. A chaos experiment is the injection mechanism — it answers "what happens if I break this." A Game Day wraps a chaos experiment in a hypothesis, a schedule, and a response team, and answers "does our documented recovery actually work." A load test answers "does the system hold up under expected and peak traffic," which is a capacity question rather than a failure question, though the two are often run together.
A DR drill is the closest relative and the most commonly confused. A DR drill exercises the failover to a standby region or environment and is typically a larger, less frequent, more disruptive event. A Game Day is usually scoped to a single failure mode within a single region and can be run far more often. The relationship is that Game Days build the muscle memory and validate the components that a DR drill then assembles into a full failover. A penetration test is unrelated in purpose — it probes for exploitable vulnerabilities rather than for recovery capability — and should not be substituted for resilience testing.
The practical rule is to pick the practice that matches the claim you need to validate. If the claim is "we recover from an AZ loss in under fifteen minutes," run a Game Day with an AZ-scoped injection. If the claim is "we can operate from our secondary region," run a DR drill. If the claim is "we handle Black Friday traffic," run a load test. The table below makes the selection explicit.
| Practice | Question it answers | Typical frequency | Pick it when… |
|---|---|---|---|
| Chaos experiment | What happens if I break this? | Continuous or on demand | You need to know the blast radius of a specific fault |
| Game Day | Does our documented recovery work? | Quarterly per critical service | A runbook or RTO claim has never been validated |
| DR drill | Can we operate from the standby region? | Annually or semi-annually | The claim is region-level failover, not component-level |
| Load test | Does it hold under peak traffic? | Before major launches | The risk is capacity, not failure recovery |
| Penetration test | Can an attacker exploit this? | Annually or after major change | The risk is security, not availability |
Hands-On Lab: A Full AZ-Failure Game Day (60 min)
In this lab you will design and document a complete Game Day exercise that simulates the loss of one Availability Zone for a three-AZ application, with success criteria defined before the exercise runs. The deliverable is a written Game Day plan plus the FIS experiment template that would execute it. You do not need to run the injection against a live production workload to complete the lab; the design artifacts are the point, and running it against a sandbox account is a bonus.
Step 1 — Define the system under test. Choose or describe a three-AZ application with an Application Load Balancer, an Auto Scaling Group spanning all three AZs, and a Multi-AZ database. Write down the current stated RTO and RPO for this workload. If no RTO exists, that is your first finding: a Game Day cannot have success criteria without a target to measure against, so set a provisional RTO and note that it needs business sign-off.
Step 2 — Write the hypothesis. State it as a falsifiable prediction with numbers. For example: "If we lose AZ-a entirely, the ALB will stop routing to targets in AZ-a within 30 seconds, the Auto Scaling Group will launch replacement capacity in AZ-b and AZ-c within 5 minutes, the database will fail over within 90 seconds, and customer-facing error rate will remain below 0.5% for the duration of the exercise." Every clause must be measurable from a dashboard you already have.
Step 3 — Define success criteria and stop conditions separately. Success criteria are what you are trying to prove: recovery within the RTO, error rate under the threshold, no data loss beyond the RPO. Stop conditions are what aborts the exercise: error rate exceeding a hard ceiling, latency exceeding a hard ceiling, or any signal that the impact is escaping the intended blast radius. Set the stop conditions tighter than the success thresholds so the exercise aborts before it consumes the error budget.
Step 4 — Build the FIS experiment template. Create an experiment template with an action that simulates AZ impairment — for example, network disruption scoped to the subnets in one AZ, or termination of the instances tagged with that AZ. Set the target selection to the resources in the chosen AZ only, and attach the stop conditions from Step 3 as CloudWatch alarm ARNs. Verify each referenced alarm is currently in OK state with a healthy evaluation history before proceeding.
Step 5 — Define the observation plan. List the exact dashboards and metrics the team will watch, and assign one person to record timestamps: injection start, first alarm, page delivered, first automated recovery action, service healthy. Assign a second person to watch CloudTrail for manual API calls so you can distinguish automated recovery from human intervention.
Step 6 — Write the retrospective template. Before running anything, create the document that will capture the deltas: predicted versus actual for each hypothesis clause, every runbook step that was wrong or missing, every alarm that fired late or not at all, and a list of corrective actions with owners. The exercise is only as valuable as this document.
Step 7 — Dry run and then execute. Rehearse the injection mechanics in a sandbox account first to confirm the template targets the right resources and the stop conditions are wired correctly. Then schedule the production exercise during business hours with the response team present, notify stakeholders, and run it. Record everything. Do not skip the dry run; the most common Game Day failure is an injection that never landed.
Scenario Question Drills (20 min)
Q1. A team has documented DR runbooks but has never tested them under simulated failure. What is the recommended next step before relying on them?
Q2. A Game Day is run and every dashboard stays flat for the full duration. The retrospective concludes the system is highly resilient. What is the most likely explanation?
Q3. An FIS experiment template used for a Game Day has no stop conditions configured. What is the risk?
Q4. A service's error-rate SLO allows 1% errors over five minutes. Where should the Game Day's stop-condition alarm threshold be set?
Q5. During a Game Day the service recovers within the RTO, but CloudTrail shows an engineer manually scaled the Auto Scaling Group. What has the exercise actually validated?
Q6. A team wants to validate that their monitoring detects an AZ failure and pages the correct on-call engineer within two minutes. Which exercise type are they running?
Q7. A Game Day is run only in a staging environment because production is considered too risky. What is the primary limitation of this approach?
Q8. A stop-condition alarm referenced by an FIS experiment has never had enough datapoints to leave INSUFFICIENT_DATA. What is the consequence during a Game Day?
Q9. A workload's stated RTO is fifteen minutes. How long should the Game Day's observation period run?
Q10. Which practice answers the question "can we operate from our standby region?"
Q11. A team's error budget is nearly exhausted for the quarter. What is the appropriate response to a scheduled Game Day?
Q12. A Game Day reveals that the runbook's failover step references a console page that no longer exists. What is the correct next action?
Q13. Which combination best describes the defining features that distinguish a Game Day from an ad hoc chaos experiment?
Q14. A team already knows their connection pool configuration is defective and will exhaust under load. They propose running a Game Day to "resolve" the issue. What is wrong with this plan?
Q15. A Game Day is scheduled at 3 a.m. to minimize customer impact. What capability does this timing choice fail to exercise?
Peek into Tomorrow
Everything in today's exercise assumed a failure you could recover from within a single region: lose an AZ, watch the surviving zones absorb the load, confirm the RTO. That assumption is comfortable, and it is also the one most likely to be wrong in the scenarios the exam cares about most. The open question a Game Day cannot answer is what happens when the failure is not a component but the region itself, and the business cannot tolerate a failover window at all — not fifteen minutes, not five, not even the ninety seconds a database promotion takes.
Answering that question forces a different architecture, not a better runbook. If two regions must serve live traffic simultaneously, then both are writing to the same logical dataset, and the moment two writers touch the same item you need a conflict resolution rule that the application can live with. It also changes how traffic is steered: DNS-based failover inherits client caching delays that a Game Day would never surface, which is why the next lesson looks at Global Accelerator's anycast IPs and Route 53 geoproximity routing as network-layer alternatives. The uncomfortable part is that active-active does not eliminate the failure modes you tested today — it multiplies them, because now every failure exists in two places at once and the replication path between them is itself a dependency.
Sources
- AWS Fault Injection Service — User Guide
- AWS FIS — Stop Conditions
- AWS Well-Architected Framework — Reliability Pillar
- Reliability Pillar — Test Reliability (Game Days)
- AWS Well-Architected Framework — Operational Excellence Pillar
- Amazon CloudWatch — Using Alarms
- Amazon CloudWatch — Composite Alarms
- AWS Cloud Operations Blog — Running Game Days with AWS FIS