Day 33 of 70 · Week 5
Day 33 / 70 Week 5 of 14 Phase 3: SRE Observability, Resilience & DR

AWS Fault Injection Service (FIS) — Chaos Engineering Fundamentals

🕑 ~58 min read · 3 services covered
AWS FIS Experiment Templates Stop Conditions

Recap: From Reading Logs to Breaking Things On Purpose

Day 32 ended on a deceptively comfortable note. Logs Insights gave us a query language over CloudWatch Logs, and subscription filters gave us a way to stream those logs in near-real time into a centralized logging account for org-wide search, retention, and SIEM ingestion. That combination is genuinely powerful, and it is also entirely passive. Every signal we built yesterday waits for something to go wrong and then describes it after the fact. A subscription filter forwarding ERROR-level events to a central account tells you that a service failed; it does not tell you whether the system was designed to survive that failure, how long recovery actually took, or whether the alarm that was supposed to page someone fired at all.

FIS extends that observability investment into something active. The alarms, metrics, and log pipelines from Days 29 through 32 stop being a rear-view mirror and become the safety harness for deliberately injecting failure into a running system. The centralized logging account becomes the place you go to prove what happened during an experiment, and the composite alarms from Day 29 become the stop conditions that keep the experiment from becoming an incident. Today is where the SRE phase stops measuring and starts testing.

Foundations You'll Need Today

Today's topic is about deliberately breaking things to see whether they heal. That only makes sense if you already know what the things are and how the system notices they broke. Four ideas carry most of the weight in this day, and none of them are explained anywhere else in the file, so we'll build them here first.

CloudWatch Alarms: The Thing That Decides When to Stop

Every AWS service emits numbers over time — how many requests it handled, how long they took, how many failed. CloudWatch is the service that collects those numbers, and a metric is one of those streams of numbers, like "the count of 5xx errors on this load balancer, sampled once per minute." A metric on its own is just a graph. An alarm is a rule you attach to a metric that says "if this number crosses this threshold, flip into a state called ALARM." You configure two things that matter here: the period (how wide each sample is — one minute, five minutes) and the evaluation periods (how many consecutive samples must breach before the alarm actually flips). An alarm with a five-minute period and three evaluation periods needs roughly fifteen minutes of sustained badness before it fires. That delay is not a bug; it exists so a single blip doesn't page someone at 3 a.m. But it becomes very important today, because FIS uses alarms as its emergency brake, and an alarm that takes fifteen minutes to fire cannot brake a two-minute experiment.

IAM Roles and Why "Passing" One Is a Separate Permission

An IAM role is a set of permissions that isn't attached to a person. Instead, something assumes the role temporarily — a person, an EC2 instance, or another AWS service — and gets those permissions for the duration. This is how AWS avoids giving services permanent, standing power over your account. FIS works exactly this way: you create a role that is allowed to, say, stop ECS tasks in one specific cluster, and you tell FIS "use this role when you run the experiment." FIS itself has no ability to touch your tasks; it borrows the role you scoped. There's a second, easily-missed piece: the person or pipeline that launches the experiment needs a permission called iam:PassRole, which is the explicit right to hand a role to a service. Without it, the launch fails even though the role itself is perfectly configured. This trips people up constantly, and the exam knows it.

Tags: Labels That Machines Can Act On

A tag is just a key-value label you stick on an AWS resource — chaos-eligible=true, Environment=prod, Owner=payments-team. What makes tags more than documentation is that AWS services can query them. Instead of naming forty specific instance IDs, you can say "act on everything tagged chaos-eligible=true," and the service figures out the list for you. This idiom shows up all over AWS — Auto Scaling groups use tags to decide which instances they manage, backup plans use tags to decide what to back up, and FIS uses tags to decide what to break. The catch, and it's the catch this day keeps returning to, is when that list gets resolved. If FIS resolves the tag at the moment the experiment starts, then a tag someone added to the wrong resource last week silently joins the blast radius today. Tags are powerful precisely because they're dynamic, and dangerous for the same reason.

ECS Tasks and Auto Scaling Groups: The Things We'll Be Breaking

Two compute patterns come up repeatedly in the examples. An ECS service runs your containerized application as a set of identical copies called tasks, and you tell it a desired count — "keep three of these running at all times." If one dies, the service notices and starts a replacement. An Auto Scaling group does the same job for plain EC2 virtual machines: you give it a desired capacity and a health check, and it replaces instances that fail. Both are self-healing by design, which is exactly why they're good chaos targets — the interesting question isn't "did the thing die," it's "how long did the system take to notice and recover, and was that inside the recovery time we promised our users?" That gap between "the design says it self-heals" and "we have measured it self-healing" is the entire reason FIS exists.

With that grounding, here's why FIS exists and what problem it actually solves: you now have alarms that can detect failure, roles that can be scoped to limit damage, tags that can define a blast radius, and self-healing systems whose recovery time is a claim rather than a fact. FIS is the service that lets you test that claim safely.

1. Why FIS Is on the Exam

SAP-C02 has a resilience domain that is weighted heavily, and within that domain the exam repeatedly probes a specific gap: the difference between an architecture that is theoretically resilient and one that has been demonstrated to be resilient. Almost every candidate can draw a multi-AZ design on a whiteboard. Far fewer can describe how they would prove that the design actually fails over correctly, that the failover completes inside the stated RTO, and that the monitoring catches the failure rather than the customers. FIS is the AWS-native answer to that second question, and the exam uses it as a marker for whether you think about resilience as a design artifact or as a continuously verified property.

The service also shows up because it forces a conversation about blast radius and safety controls that maps cleanly onto the Well-Architected Reliability pillar's failure management area. A scenario that asks you to validate recovery procedures without risking production is describing FIS with stop conditions. A scenario that asks how you would test that an Auto Scaling group replaces unhealthy instances is describing an FIS action targeting EC2 with an instance-termination fault. A scenario that asks how you would validate that a multi-AZ RDS failover completes within the RTO is describing the RDS reboot-with-failover action. The exam rarely asks "what is FIS" directly; it asks you to select the right tool for a validation goal, and FIS is the only AWS service whose entire purpose is that goal.

There is a third reason it appears, and it is the one that separates a passing answer from a confident one. FIS is frequently the correct answer in scenarios where the tempting answer is "build a custom Lambda that terminates instances on a schedule." That custom approach works, but it has no stop conditions, no IAM-scoped experiment roles, no audit trail of what was targeted, and no way to express a fault as a reusable template. The exam rewards the managed service because the managed service encodes the safety model. Understanding why the safety model matters is the real content of this day.

2. How FIS Actually Works

FIS is built around four objects, and understanding how they compose is most of the mechanism. An action is a single fault to inject: terminate an EC2 instance, stop an ECS task, inject CPU stress on a percentage of instances, add latency to network traffic, reboot an RDS instance with failover, or fail over an Aurora cluster. A target is the set of resources the action applies to, selected either by resource ID, by resource tag, or by a resource filter that resolves at runtime. An experiment template is the reusable definition that binds actions to targets along with the parameters, the IAM role FIS assumes, and the stop conditions. An experiment is a single execution of that template, with its own run ID, its own resolved target list, and its own log of what actually happened.

The target resolution step is where most of the operational subtlety lives. When you specify targets by tag, FIS resolves the tag to a concrete list of resource IDs at the moment the experiment starts, not when the template was authored. That means a template written months ago against a tag like chaos-eligible=true will target whatever carries that tag today. This is deliberate and it is the mechanism that makes FIS safe to keep in a repository: you control the blast radius by controlling tag membership, and you can shrink the eligible set to a single canary instance by removing the tag from everything else. The exam likes this pattern because it is the same tag-driven scoping used for Auto Scaling groups and backup plans, so it rewards candidates who recognize the idiom.

Actions within an experiment run in parallel by default, and each action can carry its own duration and its own parameters. A CPU stress action, for example, takes a percentage and a duration; a network latency action takes a delay in milliseconds and a target interface. FIS assumes an IAM role that you specify in the template, and that role's permissions define the outer bound of what the experiment can touch. This is a deliberate design choice: the service does not have standing permission to terminate your instances, it borrows a role you scoped for the purpose. If the role cannot describe or act on a resource, the experiment fails at that action rather than silently skipping it, which is the behavior you want when the whole point is to know exactly what was tested.

Finally, every experiment writes its state transitions and action outcomes to CloudWatch Logs and emits metrics, and it can optionally publish an event to EventBridge. That gives you a durable record of what was injected, when, against which resolved targets, and whether each action succeeded or was aborted by a stop condition. For an audit or a post-incident review, that record is the difference between "we think we tested failover last quarter" and "here is the experiment run ID and the exact instance that was terminated."

3. The Core Decision Boundary: What Are You Trying to Prove?

Every FIS scenario on the exam reduces to one question: what claim about the system are you trying to validate? The answer determines the action, the target, and — critically — the stop condition. Candidates who skip straight to picking an action tend to pick the most dramatic one, which is usually wrong. The right move is to name the claim first, then work backwards. "Our ECS service recovers from a single task failure within its target RTO" is a claim that maps to a task-termination action against a small percentage of tasks, with a stop condition on the service's error rate. "Our multi-AZ database survives an AZ loss" is a claim that maps to an RDS reboot-with-failover action, with a stop condition on connection errors. The action is downstream of the claim.

The second half of the boundary is the safety envelope. FIS experiments are only as safe as the stop conditions attached to them, and the exam tests whether you understand that stop conditions are not optional in spirit even when the API allows you to omit them. A stop condition is a CloudWatch alarm ARN; when that alarm enters ALARM state during the experiment, FIS halts the experiment and begins rolling back what it can. The rollback is not magic — terminating an instance cannot be un-terminated — but halting the remaining actions and stopping the stress injection limits how much further damage accumulates. The practical rule is that the stop condition should be the alarm that would page a human if this were a real incident, because that is exactly the threshold at which you want the experiment to stop pretending.

The table below maps common validation claims to the action and stop condition that serve them. This is the decision boundary in tabular form, and it is worth internalizing because exam scenarios are usually a paraphrase of one of these rows.

Claim to validateFIS actionTypical targetStop condition
Service survives loss of a single taskECS task terminationTag-scoped task set, count 1Service 5xx rate alarm
ASG replaces unhealthy instancesEC2 instance terminationTag-scoped instances, percentage 10-20%Healthy host count alarm
Database survives AZ lossRDS reboot with failoverSingle DB identifierConnection failure alarm
App degrades gracefully under CPU pressureCPU stressTag-scoped instances, percentage 50%p99 latency alarm
Timeouts and retries behave under latencyNetwork latency injectionTag-scoped instances, delay in msError rate alarm
Global failover completes inside RTOAurora cluster failoverCluster identifierReplication lag or write error alarm

4. Configuration Modes and Their Tradeoffs

The first real knob is target selection mode, and it determines how much of your blast radius is fixed at authoring time versus resolved at run time. Explicit resource IDs are the most predictable: the template names exactly which instances or tasks it will hit, and nothing else can ever be affected. That predictability is also the limitation, because the template goes stale the moment the fleet changes — a terminated instance ID is a permanently broken experiment. Tag-based selection inverts the tradeoff. The template stays valid as the fleet churns, and you control scope by controlling which resources carry the tag, but you have to be disciplined about tag hygiene because a misapplied tag silently widens the blast radius. Resource filters sit in between, letting you express conditions like "instances in this ASG with this state" that resolve dynamically but with more structure than a bare tag.

The second knob is the action's scope parameter, which for most instance-level actions is expressed as a count or a percentage. A percentage is almost always the right choice for a fleet because it scales with the fleet size and keeps the experiment proportional as you grow. A count is right when the claim is specifically about a fixed number of failures — "we survive the loss of exactly one task" — or when the fleet is small enough that a percentage would round to something meaningless. The exam will sometimes describe a scenario where a percentage-based experiment on a two-instance fleet would take out 50% of capacity, which is a hint that the candidate should be thinking about whether the claim is about proportional degradation or absolute failure count.

The third knob is duration and rollback behavior. Stress and latency actions have a duration after which the fault stops on its own, which is the safest configuration because the experiment self-terminates even if nothing else goes wrong. Termination actions have no duration — the instance is gone — so the recovery mechanism is the system's own self-healing rather than FIS. This distinction matters for how you write the experiment's success criteria: a stress experiment succeeds if the system holds its SLO for the duration, while a termination experiment succeeds if the system restores capacity within the RTO. Confusing the two leads to experiments that "pass" because the fault ended rather than because the system recovered.

The fourth knob is the IAM role. FIS assumes a role you name in the template, and that role's policy is the hard ceiling on what the experiment can do. The best practice is a dedicated role per experiment family with permissions scoped to exactly the actions and resource types involved, plus the iam:PassRole permission granted to whoever launches the experiment rather than baked into the role itself. A role that can terminate any EC2 instance in the account is a role that turns a typo in a target filter into an outage. The exam occasionally presents a scenario where the experiment fails with an access-denied error, and the answer is almost always that the FIS role lacks a describe or act permission on the resolved target type.

5. Sizing, Limits and Quotas

FIS quotas are modest by design, and the numbers are worth knowing because they constrain how you structure a large chaos program. An experiment template can contain a limited number of actions and targets, and the service enforces a cap on concurrent running experiments per account and per region. The practical consequence is that you cannot run a hundred simultaneous experiments across a large estate; you run a small number of well-scoped experiments, often sequentially, and you use tags to widen the target set within a single experiment rather than multiplying experiments. This is a deliberate constraint that pushes teams toward fewer, more meaningful experiments rather than a spray of small ones.

Target resolution has its own practical ceiling. When you select targets by tag, FIS resolves the tag to a list of resource IDs, and very large target sets increase the time the experiment spends in the resolving and starting phases before any fault is injected. For a fleet in the thousands, the experiment's setup time becomes visible in the run timeline, and the first action may not fire until resolution completes. This is rarely a correctness problem, but it does mean that a "quick" experiment against a huge tag set is not actually quick, and the stop conditions need to be evaluated against the window in which the fault is actually active rather than the window in which the experiment is nominally running.

Stop condition evaluation is bounded by CloudWatch alarm behavior, which is the constraint most candidates miss. A CloudWatch alarm evaluates its metric over a configured number of periods and requires a number of breaching periods before it transitions to ALARM. An alarm with a one-minute period and three evaluation periods takes roughly three minutes of sustained breach before it fires. If your experiment injects a fault for two minutes, the stop condition may never trigger even if the system is genuinely failing, because the alarm never had time to transition. The fix is to configure the stop-condition alarm with a shorter evaluation window than the experiment's fault duration, or to lengthen the fault duration so the alarm has room to fire. This is a real design coupling between the experiment and the alarm, and it is exactly the kind of detail the exam probes with a scenario about an experiment that "did not stop when it should have."

Finally, FIS is a regional service. Experiments, templates, and their logs live in the region where they run, and an experiment targeting resources in another region is not a thing. For multi-region validation you run separate experiments per region, and if you want a coordinated multi-region failure you orchestrate that at a higher layer — which is precisely the gap that tomorrow's Game Day discussion fills.

6. Failure Modes and What They Look Like in Production

The most common failure mode is not the experiment going wrong; it is the experiment going right and revealing that the system was never as resilient as the design diagram claimed. The symptom is a service that does not recover within its RTO after a single task termination, or an Auto Scaling group that takes longer than expected to replace an instance because the launch template's bootstrap script is slow. These are findings, not bugs in FIS, and the correct response is to treat them as production defects with the same priority as any other reliability gap. The first diagnostic move is to read the experiment's action log to confirm the fault actually landed on the intended target, then read the service's own metrics over the experiment window to see where recovery stalled.

The second failure mode is a stop condition that fires too late or not at all, which turns a controlled experiment into an uncontrolled incident. The symptom is an experiment that runs to completion while the service is visibly degraded, with no automatic halt. The cause is almost always an alarm whose evaluation window is longer than the fault duration, or an alarm that is watching the wrong metric — for example, an alarm on host-level CPU when the actual user-facing symptom is elevated latency. The first diagnostic move is to check the alarm's state history for the experiment window and confirm whether it ever entered ALARM. If it did not, the alarm configuration is the defect, not the experiment.

The third failure mode is target resolution picking up more than intended. This happens when a tag used for chaos eligibility is also applied to resources that were never meant to be in scope, often because the tag was added by an automated process or copied from a template. The symptom is an experiment that reports a target count far higher than expected, or that touches a resource in an environment you did not intend. The first diagnostic move is to inspect the resolved target list in the experiment details before the fault is injected — FIS exposes this — and to treat any surprise in that list as a reason to abort. The durable fix is to make the chaos tag exclusive and to enforce its application through a Config rule or an SCP rather than relying on humans to remember.

The fourth failure mode is permission drift. The FIS role works today and fails next quarter because someone tightened a policy or renamed a resource type. The symptom is an experiment that fails immediately with an access-denied error on one action while others succeed. The first diagnostic move is to read the failed action's error and compare the required permission against the role's policy. Because FIS fails loudly rather than skipping, this is a benign failure mode — it costs you a test run, not an outage — but it does mean that experiment templates need to be exercised regularly enough that permission drift is caught before you need the experiment to work.

7. The Operational and SRE Angle

From an SRE perspective, FIS is the mechanism that converts an SLO from a dashboard number into a tested property. The operational pattern is to tie each experiment to a specific SLO and to a specific alarm that represents that SLO being breached, then run the experiment on a schedule and treat a stop-condition trip as a finding. If the experiment trips the stop condition, the system did not hold its SLO under the injected fault, and that is a reliability defect to be fixed before the next experiment. If the experiment completes without tripping, you have evidence — with a run ID and a timestamp — that the system held its SLO under that fault. Over time, that evidence is what lets you raise confidence in the RTO and RPO numbers you publish.

The monitoring shape follows directly. You want the experiment's own metrics and logs alongside the service's metrics, so that a reviewer can see the fault window and the service's response on the same timeline. FIS emits experiment state transitions and action outcomes, and it can publish to EventBridge, which means you can route experiment events into the same centralized logging account that Day 32 built. That gives you a single place where an experiment's start, its resolved targets, its action outcomes, and its stop-condition trips are all recorded next to the application logs they affected. For an audit, that is the complete story.

The runbook shape is a short, repeatable procedure: confirm the experiment template's target set is what you expect, confirm the stop-condition alarm is in OK state and has a sane evaluation window, launch the experiment, watch the experiment timeline and the service's golden signals together, and record the outcome against the SLO. The runbook should also name the abort path — how to stop a running experiment manually if the stop condition does not fire — because the stop condition is a safety net, not a substitute for a human watching the first few runs. New experiments should run in a non-production environment first, then against a single canary in production, then against a wider tag set once the canary runs clean.

One operational detail worth internalizing: FIS experiments are not free of side effects on your monitoring. A CPU stress experiment will move your CPU utilization metrics, and if you have anomaly-detection alarms on those metrics, the experiment will trip them. That is usually desirable — it proves the alarm works — but it also means that experiment windows should be annotated on dashboards and excluded from SLO calculations where the SLO is meant to reflect real user traffic. The cleanest approach is to tag experiment windows in your observability tooling so that post-incident reviews can distinguish injected faults from organic ones.

8. Edge Cases and Exam Gotchas

The single most-tested gotcha is that stop conditions are CloudWatch alarms, not FIS-native thresholds. You cannot express "stop if error rate exceeds 5%" directly in FIS; you create a CloudWatch alarm that does that and reference its ARN. Candidates who describe a stop condition as a percentage or a rule rather than an alarm ARN are describing something the service does not do. The corollary is that the alarm's evaluation period and datapoints-to-alarm settings are part of the stop condition's behavior, and a stop condition that never fires is usually an alarm configuration problem rather than an FIS problem.

The second gotcha is that FIS does not roll back destructive actions. Terminating an instance is not reversible; the recovery is the system's own self-healing. This means the experiment's success criterion for a destructive action is about recovery time, not about the fault ending. A scenario that asks how to test that an ASG replaces a terminated instance is asking about the ASG's health-check and replacement behavior, with FIS as the trigger. A scenario that asks how to test graceful degradation under load is asking about a stress action with a duration, where the fault does end on its own.

The third gotcha is the distinction between FIS and a custom automation. A Lambda that terminates instances on a schedule can inject the same fault, but it has no stop conditions, no resolved-target audit trail, no reusable template, and no scoped experiment role. When a scenario emphasizes safety, auditability, or repeatability, FIS is the answer; when a scenario emphasizes a bespoke fault that no FIS action supports, a custom automation may be the only option, and the exam will usually signal that by describing a fault type outside FIS's action catalog.

The fourth gotcha is scope creep through tags. Because tag-based targeting resolves at run time, a tag that is applied too broadly silently widens the blast radius. The exam may describe an experiment that affected more instances than intended, and the answer is usually about tag hygiene and enforcement rather than about FIS configuration. The fifth gotcha is regional scope: FIS experiments are regional, and a multi-region failure requires either separate experiments per region or an orchestration layer above FIS. The sixth is that the FIS role needs iam:PassRole granted to the principal launching the experiment, not to the role itself, which is a common source of access-denied errors in lab environments.

9. FIS vs. the Services It Gets Confused With

FIS is most often confused with three things: custom fault-injection automation, load testing tools, and the observability services it depends on. The distinction with custom automation is the safety model — FIS encodes stop conditions, scoped roles, and resolved-target auditing that a hand-rolled Lambda does not. The distinction with load testing is intent: load testing asks whether the system can handle a given volume of traffic, while FIS asks whether the system behaves correctly when a component fails. A load test that saturates a service and a chaos experiment that kills a task are different questions, and the exam will sometimes present a scenario where the correct answer is a load test rather than FIS because the goal is capacity validation rather than failure validation.

The distinction with observability services is that FIS consumes them rather than replacing them. CloudWatch alarms are the stop conditions; CloudWatch Logs is where experiment outcomes land; X-Ray traces show how the request path behaved during the fault; Synthetics canaries show whether the user-facing flow stayed healthy. FIS without those services is a fault injector with no safety net and no evidence. The exam rewards candidates who describe FIS as the active layer on top of a passive observability stack, which is exactly the relationship this day has to Day 32.

OptionWhat it doesPick it when…
AWS FISManaged fault injection with templates, scoped roles, and alarm-based stop conditionsYou need safe, auditable, repeatable failure validation against real infrastructure
Custom Lambda automationBespoke fault injection with no built-in safety modelThe fault type is outside FIS's action catalog and you accept building the safety controls yourself
Load testing toolsGenerates traffic volume to validate capacity and throughputThe question is whether the system handles a given load, not whether it survives a component failure
CloudWatch alarmsDetects and notifies on metric thresholdsYou need detection and paging — FIS uses these as stop conditions rather than replacing them
X-Ray / SyntheticsTraces request paths and probes user flowsYou need to see how the system behaved during the fault, not to inject the fault

Hands-on Lab: Terminating an ECS Task Under a Stop Condition

Goal: run a controlled FIS experiment that terminates a single ECS task in a running service, with a stop condition tied to the service's error-rate alarm, and verify the service self-heals within its target RTO.

Prerequisites: an ECS service running on Fargate or EC2 with at least two tasks, an Application Load Balancer in front of it, a CloudWatch alarm on the ALB's 5xx count or the service's error rate, and permissions to create IAM roles, FIS templates, and CloudWatch alarms.

  1. Tag the ECS tasks that are eligible for chaos with a dedicated tag, for example chaos-eligible=true. Apply it to the service's tasks only, and confirm no other resource in the account carries that tag. This tag is the blast-radius control for the whole lab.
  2. Create the stop-condition alarm. Point it at the ALB's HTTPCode_Target_5XX_Count metric (or your service's equivalent error metric), set the period to one minute, set the evaluation periods to one, and set the datapoints to alarm to one. A one-minute, one-datapoint alarm transitions to ALARM quickly, which is what you want for a short experiment. Confirm the alarm is in OK state before proceeding.
  3. Create the FIS execution role. The role needs permissions to describe and stop ECS tasks in the target cluster, plus permissions to describe the resources FIS resolves during target selection. Keep the policy scoped to the specific cluster and task definition rather than granting ecs:*.
  4. Create the experiment template. Add one action of type aws:ecs:stop-task. Set the target to the ECS task ARNs selected by the chaos-eligible=true tag, with a selection mode of COUNT and a count of one. Attach the stop-condition alarm ARN from step 2. Attach the execution role from step 3. Give the template a name that encodes the claim it validates, for example ecs-single-task-loss-rto.
  5. Before launching, open the template's target preview and confirm the resolved target list contains exactly the tasks you expect. If the list is larger than expected, stop and fix the tag before running anything.
  6. Launch the experiment. Watch three things on the same timeline: the FIS experiment timeline (resolving, running, completed), the ECS service's running-task count, and the ALB's healthy-host count. Note the timestamp when the task is stopped and the timestamp when the service returns to its desired task count.
  7. Compute the recovery time as the difference between those two timestamps. Compare it against the RTO you have documented for the service. If recovery exceeds the RTO, you have found a real defect — investigate the task's startup time, the health-check grace period, and the ALB's deregistration delay.
  8. Deliberately trip the stop condition to prove it works. Temporarily lower the alarm threshold so that normal traffic breaches it, launch the experiment again, and confirm that FIS halts the experiment and reports the stop condition as the reason. Restore the threshold afterwards.
  9. Record the outcome. Capture the experiment run ID, the resolved target, the recovery time, and whether the stop condition fired. Store this alongside the service's SLO documentation so the next reviewer can see the evidence rather than the claim.

What to look for: the most common finding is that recovery is slower than expected because the ALB's deregistration delay keeps the terminated task's connections draining longer than the service's own startup time. That is a configuration defect worth fixing, and it is exactly the kind of thing a design review would never surface.

Scenario Question Drills

Q1. A team wants to safely test EC2 instance failure in production without risking an uncontrolled outage. What FIS feature guarantees the experiment halts if things go wrong?

A. IAM permission boundaries
B. Stop conditions tied to CloudWatch alarms
C. Increasing the blast radius
D. Manual monitoring only
Correct answer: B. FIS stop conditions automatically abort a running experiment the moment a linked CloudWatch alarm enters ALARM state, capping the blast radius of chaos testing.

Q2. An experiment template targets instances by the tag chaos-eligible=true. Six months later the experiment is launched and affects far more instances than the original author intended. What is the most likely cause?

A. FIS ignores tags and always targets the whole account
B. Tag-based targets resolve at run time, so any resource that acquired the tag since authoring is now in scope
C. The experiment template cached the original target list and reused it
D. Stop conditions widened the target set
Correct answer: B. Tag-based target selection resolves to concrete resource IDs at experiment start, not at template authoring, so tag hygiene and enforcement are the real blast-radius control.

Q3. A team configures a stop condition on an alarm with a five-minute period and three evaluation periods, then runs an experiment that injects a fault for two minutes. The experiment completes without stopping even though the service degraded. Why?

A. FIS stop conditions only work on EC2 actions
B. The alarm needs roughly fifteen minutes of sustained breach to transition to ALARM, far longer than the fault window
C. Stop conditions must be SNS topics, not alarms
D. The experiment role lacked permission to read the alarm
Correct answer: B. A stop condition is only as fast as its alarm's evaluation window; the alarm must be able to transition to ALARM within the fault duration for the stop condition to be meaningful.

Q4. Which FIS action is appropriate for validating that an application degrades gracefully under sustained CPU pressure without the fault persisting indefinitely?

A. EC2 instance termination
B. CPU stress with a configured duration
C. RDS reboot with failover
D. Aurora cluster failover
Correct answer: B. Stress actions take a duration and stop on their own, which is the right shape for validating graceful degradation; termination actions have no duration and rely on the system's own recovery.

Q5. A team wants to prove that their multi-AZ RDS deployment survives an Availability Zone loss within the documented RTO. Which FIS action and stop condition pairing fits?

A. EC2 instance termination with a CPU alarm as the stop condition
B. RDS reboot with failover, with a connection-failure alarm as the stop condition
C. Network latency injection with a healthy-host alarm
D. ECS task termination with a 5xx alarm
Correct answer: B. The RDS reboot-with-failover action forces the standby to take over, and the stop condition should be the alarm that represents the user-facing symptom of a database outage.

Q6. An FIS experiment fails immediately with an access-denied error on one action while other actions succeed. What is the most likely cause?

A. The stop condition alarm is in ALARM state
B. The FIS execution role lacks the permission required for that action's resource type
C. FIS cannot run more than one action per experiment
D. The target tag was applied to too many resources
Correct answer: B. FIS assumes a role you specify, and that role's policy is the hard ceiling on what the experiment can do; a missing permission fails that action loudly rather than skipping it.

Q7. A team wants to validate that their Auto Scaling group replaces unhealthy instances within the target RTO. Which FIS action and success criterion are correct?

A. CPU stress, with success measured by the fault ending on its own
B. EC2 instance termination, with success measured by the ASG restoring capacity within the RTO
C. Network latency injection, with success measured by the alarm staying in OK
D. ECS task termination, with success measured by the task count returning to desired
Correct answer: B. Termination is destructive and not rolled back by FIS, so the success criterion is the system's own recovery time, not the fault ending.

Q8. Which statement best describes the relationship between FIS and the observability services built in Days 29-32?

A. FIS replaces CloudWatch alarms with its own native threshold engine
B. FIS consumes CloudWatch alarms as stop conditions and writes experiment outcomes to CloudWatch Logs, acting as the active layer on top of a passive observability stack
C. FIS is independent of observability and requires no alarms
D. FIS only works with X-Ray and not with CloudWatch
Correct answer: B. FIS does not replace observability; it uses alarms as stop conditions and logs as its audit trail, which is why the observability investment from the previous days is a prerequisite.

Q9. A scenario asks how to validate that a service handles a bespoke failure mode that no FIS action supports. What is the correct approach?

A. Force the closest FIS action to approximate the fault
B. Build custom fault-injection automation, accepting that you must implement the safety controls yourself
C. Skip the validation entirely
D. Use a load testing tool instead
Correct answer: B. FIS covers a catalog of common faults; a fault outside that catalog requires custom automation, and the tradeoff is that you lose FIS's built-in stop conditions and audit trail.

Q10. A team wants to run a chaos experiment against a fleet of 2,000 instances but only affect a small, proportional slice. Which target configuration is most appropriate?

A. Explicitly list 2,000 resource IDs in the template
B. Tag-based selection with a percentage-based scope so the blast radius scales with the fleet
C. A count of 1,000 to keep it simple
D. Run 2,000 separate experiments
Correct answer: B. Percentage-based scope keeps the experiment proportional as the fleet grows, and tag-based selection keeps the template valid as instances churn.

Q11. An experiment targeting resources in us-east-1 is authored, but the team also wants to validate the same fault in eu-west-1. What is the correct approach?

A. Add the eu-west-1 resources to the same template's target list
B. Run a separate experiment in eu-west-1, since FIS is a regional service
C. Use a global FIS template that spans regions
D. FIS cannot target resources outside the management account
Correct answer: B. FIS experiments, templates, and logs are regional; multi-region validation requires separate experiments per region or an orchestration layer above FIS.

Q12. A team's experiment trips its stop condition and halts. What is the correct interpretation of this outcome?

A. The experiment failed and should be deleted
B. The system did not hold its SLO under the injected fault, which is a reliability defect to fix before the next experiment
C. The stop condition is misconfigured and should be removed
D. FIS is not suitable for this workload
Correct answer: B. A stop-condition trip is a finding, not a failure of the tooling; it means the system breached the threshold you defined as unacceptable, and that gap should be fixed.

Q13. Which permission must be granted to the principal that launches an FIS experiment, in addition to the permissions on the FIS execution role itself?

A. cloudwatch:PutMetricData
B. iam:PassRole for the FIS execution role
C. ec2:TerminateInstances directly
D. logs:CreateLogGroup
Correct answer: B. The launching principal needs iam:PassRole for the execution role; the role itself carries the permissions FIS uses to act on targets.

Q14. A team wants to prove that their alarm actually fires during a real failure, not just that the alarm exists. How should they use FIS to do this?

A. Manually set the alarm to ALARM state in the console
B. Inject a fault that should breach the alarm's metric and confirm the alarm transitions to ALARM during the experiment window
C. Delete and recreate the alarm
D. Increase the alarm's evaluation period
Correct answer: B. Injecting a fault that should breach the metric is the only way to validate the alarm end-to-end, including its metric, threshold, and evaluation window.

Q15. A team runs a CPU stress experiment and finds that their anomaly-detection alarms on CPU utilization trip during the experiment window, polluting their SLO dashboard. What is the best practice?

A. Disable anomaly detection permanently
B. Annotate experiment windows in observability tooling so injected faults can be distinguished from organic ones and excluded from SLO calculations
C. Stop running CPU stress experiments
D. Raise the alarm thresholds so they never trip
Correct answer: B. Experiment windows should be annotated so post-incident reviews and SLO calculations can distinguish injected faults from real user-impacting events.

Peek into Tomorrow

Everything covered here assumes a single, well-scoped experiment against a single fault type, run by a team that is watching. That is the right starting point, but it leaves an obvious question unanswered: what does it look like to validate an entire recovery strategy rather than one component? A single ECS task termination proves that the service replaces a task. It does not prove that the organization's runbooks are correct, that the on-call rotation knows what to do when the pager fires, or that the documented RTO for a full Availability Zone loss is anything other than a number someone wrote in a wiki. Those are different claims, and they need a different structure than a single experiment template.

Tomorrow's topic is the scheduled Game Day — a simulated AZ outage or region loss exercise that turns DR assumptions into tested facts. The open question today leaves is how you compose multiple FIS experiments, a runbook, and a set of success criteria into a single exercise that a team can actually run and learn from, and how you decide in advance what "the exercise succeeded" means when the failure you are simulating is large enough that partial degradation is expected. That is a design problem about exercises, not about fault injection, and it is where the SRE phase gets genuinely hard.

Sources