Day 29 of 70 · Week 5
Day 29 / 70 Week 5 of 14 Phase 3: SRE Observability, Resilience & DR

CloudWatch Metrics, Alarms & Composite Alarms

🕑 ~58 min read · 4 services covered
CloudWatch Metrics Alarms Composite Alarms Anomaly Detection

Recap: From Data Selection to Signal Selection

Day 28 closed Phase 2 by locking in the database and storage selection matrix — the relational vs key-value vs cache fork, the consistency model trade-offs that sit behind it, and the global replication decision between Aurora Global DB and DynamoDB Global Tables. That synthesis was deliberately a selection exercise: given a stated latency, scale, and consistency requirement, name the data service and defend it. Phase 2 as a whole was built the same way. Compute, containers, and data stores were all presented as a set of forks with a defensible answer on each side, and the exam rewards candidates who can walk into a scenario and eliminate options on architecture grounds rather than on service trivia.

Phase 3 extends that same selection discipline into a domain where the forks are harder to see, because the artifacts are not resources you provision but signals you interpret. Observability, resilience, and disaster recovery all hinge on a prior question: what does the system look like when it is healthy, and what does it look like when it is not? CloudWatch metrics and alarms are the first layer of that answer, and they are the layer every other Phase 3 topic — canaries, tracing, chaos experiments, DR failover — ultimately reports into. The selection matrix you built yesterday chose what to store. Today's matrix chooses what to watch.

Foundations You'll Need Today

Today's topic is monitoring, and monitoring has its own vocabulary that the rest of this curriculum uses as if it were obvious. Before the architecture discussion starts, here are the four ideas everything else in this day is built on.

Metrics, Namespaces, and Dimensions

A metric is just a number that something reports over time — how many requests a service handled in the last minute, how much disk space is free, how long a request took. AWS services report these numbers automatically, and you can also have your own code report custom ones. The problem this solves is that without a running record of these numbers, you have no way to know whether a system is healthy or degrading until a user tells you it is broken.

Two pieces of vocabulary matter here. A namespace is the grouping a metric belongs to — AWS services each publish into their own namespace, and your custom metrics go into one you name yourself. A dimension is a further qualifier that identifies which specific thing the number is about, such as which particular server or which particular load balancer. The reason this matters is that a metric with a different dimension is treated as a completely separate metric, not as another data point on the same line. If you are watching five servers, you have five separate metrics, and an alarm attached to one of them knows nothing about the other four.

Statistics and Periods

Raw numbers arrive constantly, and no human wants to look at every single one. So when you ask a monitoring system a question about a metric, you ask it in terms of a statistic over a period. The period is the time window — "the last five minutes" — and the statistic is how you want those five minutes summarized: the average, the highest value, the total, or a percentile.

The percentile is the one worth pausing on, because it is the least intuitive and the most important for today. If you line up every request by how long it took and pick the one that sits at the ninety-ninth position out of a hundred, that is the p99 — the value that ninety-nine percent of requests came in under. It exists because averages hide problems: if ninety-nine requests take fifty milliseconds and one takes ten seconds, the average looks perfectly healthy while one real user waited ten seconds. A percentile lets you ask about the slow tail directly instead of having it averaged away.

Alarms, Thresholds, and Alarm States

An alarm is a rule that watches one metric and changes its own status when that metric crosses a line you drew. The line is the threshold, and the status is the alarm state. The reason alarms exist is that a metric sitting on a dashboard only helps someone who is already looking at the dashboard; an alarm is what turns a number into something that reaches a person or triggers an automated response.

There are three states, not two, and the third one is the trap. OK means the metric is inside its normal range. ALARM means it has crossed the threshold. INSUFFICIENT_DATA means the alarm has no numbers to evaluate at all — which happens when a metric stops being reported, for instance because the thing reporting it was deleted or crashed. The reason this is a trap is that a stopped metric is often the most serious problem of all, and by default it does not raise an alarm; it just sits quietly in that third state. Keep this in mind, because it comes back repeatedly in today's material.

Who Publishes, and Who Gets Told

Two roles get conflated constantly. The publisher is whatever produces the metric — an AWS service, or your own application code. The notification channel is whatever receives the alarm when it fires, which on AWS is usually a simple pub/sub topic that fans the message out to email, chat, or another piece of automation. They are separate concerns, and separating them is what lets you change who gets paged without touching the thing being measured. It also means that if the publisher dies, nothing is being measured at all — and no notification channel in the world can tell you about a number that was never sent.

With that grounding, here is why CloudWatch alarms exist, what problem they actually solve, and where the design decisions get genuinely hard.

1. Why This Is on the Exam

CloudWatch alarms appear on SAP-C02 in two distinct guises, and confusing them is the most common way candidates lose points on otherwise straightforward questions. The first guise is as a component inside a larger architecture: an alarm that triggers an Auto Scaling policy, an alarm that acts as a stop condition on a Fault Injection Service experiment, an alarm that drives a Route 53 health check or a Lambda-based remediation. In those questions the alarm is not the subject of the question — it is a required piece of plumbing, and the exam is testing whether you know that the alarm must exist and what it must be attached to. The second guise is as the subject itself: a scenario describes alert fatigue, a noisy pager, or a monitoring gap, and the correct answer is a specific alarm configuration rather than a different service.

The reason this topic carries weight is that it sits at the intersection of the Resilient Architectures domain and the Operational Excellence concerns that thread through every scenario. A design that has no alarms is not merely incomplete; it cannot satisfy a stated RTO, because nothing detects the failure that starts the recovery clock. A design that has too many alarms is also wrong, because the exam increasingly tests whether you understand that an alerting system which pages on every transient spike is functionally equivalent to one that pages on nothing — the on-call engineer stops trusting it. Composite alarms and anomaly detection bands exist precisely to answer that second failure mode, and they are the two features most likely to appear as the distinguishing detail between two otherwise plausible answers.

There is also a quieter reason. CloudWatch is the default integration point for nearly every other AWS service's health signal, which means a question about ECS task health, RDS storage, Lambda error rates, or NAT Gateway port exhaustion is often really a question about whether you know which metric to alarm on and at what granularity. The exam does not expect you to memorize every metric name, but it does expect you to know the shape of the answer: which namespace, which statistic, which period, and which of the alarm states you would act on.

2. How Metrics and Alarms Actually Work

A CloudWatch metric is a time-ordered series of data points identified by a namespace, a metric name, and a set of dimensions. Dimensions are the part that trips people up, because they are not labels in the Kubernetes sense — they are part of the metric's identity. A metric published as AWS/EC2 with dimension InstanceId=i-abc123 is a different metric from the same name with dimension InstanceId=i-def456, and there is no query that aggregates across them unless you explicitly ask for it. This matters for alarms because an alarm is bound to exactly one metric (or one metric math expression), and the dimensions you choose determine whether that alarm watches a single instance, a single target group, or an aggregate.

Data points arrive at a resolution determined by the publisher. Most AWS service metrics are standard resolution, published at one-minute granularity, and retained for fifteen months with progressively coarser aggregation as they age. Custom metrics can be published at high resolution, down to one-second granularity, and high-resolution alarms can evaluate on those shorter periods. The retention and resolution rules are worth internalizing because they bound what an alarm can possibly detect: an alarm whose period is five minutes cannot react to a thirty-second spike, no matter how the threshold is set, because the spike is averaged into the five-minute datapoint before the alarm ever sees it.

An alarm evaluates a statistic over a period against a threshold, and it does so repeatedly over a number of evaluation periods. The DatapointsToAlarm and EvaluationPeriods pair is the mechanism that controls sensitivity: with three evaluation periods and a datapoints-to-alarm of three, the metric must breach on three consecutive periods before the alarm transitions to ALARM. With three evaluation periods and a datapoints-to-alarm of one, a single breaching period is enough. This is the knob that most candidates overlook, and it is frequently the difference between a design that flaps and one that does not.

Alarms have three states, not two. OK and ALARM are the obvious pair, but INSUFFICIENT_DATA is a real state that occurs when the metric has no datapoints for the evaluation window — a newly created alarm, a metric that stopped publishing because the resource was terminated, or a sparse custom metric. Treating INSUFFICIENT_DATA as equivalent to OK is a genuine production hazard, because a resource that has stopped reporting entirely will sit in INSUFFICIENT_DATA and never page anyone. The TreatMissingData setting exists to let you choose whether missing data is treated as breaching, not-breaching, or ignored, and the correct choice depends on whether absence of data means "healthy and idle" or "dead."

3. The Core Decision Boundary: Single-Metric vs Composite

The fork that most scenario questions hinge on is whether a single metric threshold is sufficient to describe the condition you actually care about, or whether the condition is inherently a conjunction of signals. A single-metric alarm answers a narrow question: is this one number outside its expected range? That is the right tool when the metric is a direct proxy for the thing you care about — disk space consumed, queue depth, replication lag. It is the wrong tool when the thing you care about is a user-visible symptom that can be caused by several different underlying conditions, because then any single metric will either fire on conditions that are not actually user-impacting, or fail to fire on conditions that are.

Composite alarms exist to express that conjunction. A composite alarm takes a rule expression over the states of other alarms — ALARM(latencyAlarm) AND ALARM(errorRateAlarm) — and enters ALARM only when the expression evaluates true. The important architectural property is that composite alarms do not evaluate metrics at all; they evaluate alarm states. That means they compose recursively, and it also means they inherit the evaluation semantics of their children, including the children's own periods and datapoints-to-alarm settings. A composite alarm cannot be more sensitive than the alarms feeding it.

The practical consequence is that composite alarms are the correct answer whenever a scenario describes alert fatigue, a noisy pager, or a requirement to page only on genuine incidents. They are also the correct answer when a scenario describes a condition that is only meaningful in combination — high latency alone might be a slow dependency, but high latency together with elevated error rate is a service degradation. The table below lays out the boundary.

Condition you are describingRight toolWhy
A single resource crossing a hard limit (disk full, queue depth, replication lag)Single-metric alarm with a static thresholdThe metric is a direct proxy; no correlation needed
A user-visible symptom with multiple possible causesComposite alarm over the candidate cause alarmsExpresses the conjunction; suppresses single-cause noise
A metric whose normal range varies by time of day or traffic patternAnomaly detection band alarmStatic thresholds cannot track a moving baseline
A condition that should page only if it persistsSingle-metric alarm with multiple evaluation periodsPersistence is a temporal property, not a correlation
A condition that should page only if two independent subsystems agreeComposite alarm with ANDIndependent signals reduce false-positive rate multiplicatively

4. Configuration Modes and Their Tradeoffs

Once you have decided what condition you are describing, the configuration knobs determine how faithfully the alarm tracks it. The first and most consequential choice is the statistic. Average is the default and the most misleading for latency, because a service where ninety-nine percent of requests complete in fifty milliseconds and one percent take ten seconds has an average that looks healthy. Maximum catches the tail but is hypersensitive to a single outlier. Percentile statistics — p90, p99, p99.9 — are the correct choice for latency SLOs, and the exam expects you to reach for them when a scenario mentions tail latency or a small fraction of slow requests. The tradeoff is that percentiles require enough datapoints to be statistically meaningful, so they are a poor fit for low-traffic metrics.

The second choice is the period, which sets the resolution of the evaluation. A shorter period detects faster but is noisier, because a single bad minute can breach a one-minute alarm. A longer period is smoother but introduces detection latency proportional to its length. The interaction with the statistic matters: a one-minute maximum is far more volatile than a five-minute average, and choosing both a short period and a maximum statistic produces an alarm that will flap on any transient blip. The usual production pattern is to pair a percentile statistic with a period long enough to be stable, then use evaluation periods to require persistence.

The third choice is the threshold type. Static thresholds are simple, auditable, and correct when the metric has a known acceptable range. Anomaly detection bands replace the static number with a model that learns the metric's normal pattern — including daily and weekly seasonality — and alarms when the metric falls outside a configurable number of standard deviations from the predicted band. The tradeoff is legibility: a static threshold can be explained in a runbook in one sentence, while an anomaly band's behavior depends on training data and can be difficult to reason about during an incident. Anomaly detection is the right answer when the metric's baseline genuinely moves with traffic, and the wrong answer when the metric has a hard physical limit that a static threshold expresses more clearly.

The fourth choice is the action. An alarm can notify an SNS topic, trigger an Auto Scaling policy, stop or start an instance, or invoke an SSM automation. The design question is whether the action is a notification or a remediation. Remediation actions should be paired with a notification, because an automated action that fires silently is indistinguishable from a bug. And any alarm that drives an automated remediation should have a corresponding alarm on the remediation's own success, or you will eventually discover that your self-healing system has been failing to heal for weeks.

5. Sizing, Limits and Quotas

The numbers that matter for this topic are mostly about resolution and retention, because those bound what an alarm can detect. Standard-resolution metrics are published at one-minute granularity. High-resolution custom metrics can be published at one-second granularity, and alarms on those metrics can evaluate with periods as short as ten seconds or thirty seconds. The practical implication is that sub-minute detection requires a custom metric — you cannot get a ten-second alarm on a standard EC2 metric no matter how you configure it.

Retention follows a tiered schedule. Datapoints at one-minute resolution are retained for fifteen days, five-minute datapoints for sixty-three days, one-hour datapoints for four hundred and fifty-five days, and beyond that the data is aggregated further. This matters for capacity planning and for post-incident review: if you want to compare this month's latency against the same month last year, you are looking at hourly aggregates, not the original minute-level data.

Alarm evaluation has its own limits. Each alarm evaluates one metric or one metric math expression, and metric math lets you combine up to a bounded number of metrics into a derived series — useful for ratios like error rate as a percentage of total requests, which is a far better alarm target than raw error count. Composite alarms can reference other composite alarms, but the rule expression has a bounded size, so deeply nested hierarchies eventually hit a ceiling. In practice the useful depth is two or three levels: leaf alarms on individual metrics, a middle layer expressing service-level conditions, and a top layer expressing customer-visible impact.

PropertyValueConsequence for alarm design
Standard metric resolution1 minuteFastest standard alarm period is 60s
High-resolution custom metric resolution1 secondEnables 10s/30s alarm periods
1-minute datapoint retention15 daysRecent-incident forensics only
5-minute datapoint retention63 daysQuarterly trend comparison
1-hour datapoint retention455 daysYear-over-year comparison
Alarm statesOK, ALARM, INSUFFICIENT_DATAMissing data is a distinct state, not OK

6. Failure Modes and What They Look Like in Production

The first failure mode is the alarm that never fires because the metric stopped publishing. This is the INSUFFICIENT_DATA trap, and it is the most dangerous because it is silent. A Lambda function that has been deleted, an EC2 instance that has been terminated, a custom metric publisher that has crashed — all of these produce a metric that simply stops, and an alarm with the default missing-data treatment will sit in INSUFFICIENT_DATA indefinitely without notifying anyone. The symptom is a monitoring gap discovered during an incident, and the first diagnostic move is to check the alarm's state history rather than its current state. The fix is to set TreatMissingData to breaching for metrics where absence means failure, and to add a separate alarm on the publisher's own health.

The second failure mode is the flapping alarm. This presents as a pager that fires and clears repeatedly over a short window, and it is almost always a configuration problem rather than a real instability. The usual causes are a period too short for the metric's natural variance, a maximum statistic on a noisy metric, or a threshold set at the edge of the normal range. The diagnostic move is to pull the metric's history over the last week and look at where the threshold sits relative to the distribution — if the threshold is inside the normal band, the alarm is measuring noise. The fix is to widen the period, switch to a percentile statistic, or require multiple evaluation periods.

The third failure mode is the alarm that fires correctly but on the wrong thing. This happens when the alarm watches an infrastructure metric that is a poor proxy for the user-visible symptom. CPU utilization on a fleet behind a load balancer is the classic example: the fleet can be at thirty percent CPU while every request is failing, because the failure is in a downstream dependency. The symptom is an incident where the alarms that fired were not the alarms that mattered, and the diagnostic move is to trace backward from the user-visible symptom to the metric that would have detected it earliest. The fix is usually to add a composite alarm at the service level rather than to remove the infrastructure alarms.

The fourth failure mode is alarm fatigue itself, which is a failure mode of the monitoring system rather than of the application. It presents as engineers acknowledging pages without investigating, or muting alarms during incidents. It is caused by a pager whose signal-to-noise ratio has degraded, and the fix is structural: composite alarms to express conjunctions, anomaly detection for metrics with moving baselines, and a periodic review of which alarms actually resulted in action. An alarm that has never once led to a human doing something is a candidate for deletion or for demotion to a dashboard.

7. The Operational and SRE Angle

From an SRE perspective, alarms are the mechanism that converts an SLO into an actionable signal. An SLO is a target — ninety-nine point nine percent of requests succeed within three hundred milliseconds over a thirty-day window — and the error budget is the amount of failure that target permits. The alarm's job is to fire when the burn rate against that budget is high enough that a human needs to intervene before the budget is exhausted. This framing produces a different alarm design than the infrastructure-centric one: instead of alarming on CPU, you alarm on the ratio of slow requests to total requests, and you alarm on the rate at which the error budget is being consumed rather than on the absolute error count.

Multi-window burn-rate alerting is the pattern that follows from this. A fast-burn alert uses a short window and a high burn rate to catch acute outages — the kind that would exhaust a monthly budget in hours. A slow-burn alert uses a longer window and a lower burn rate to catch chronic degradation that would exhaust the budget over days. Both are needed, because a single window either misses slow burns or pages too often on fast ones. In CloudWatch terms this is implemented with metric math over the request and error counts, producing a burn-rate series that a static threshold alarm can watch.

The runbook shape follows directly. Every alarm that pages a human should have a runbook entry that names the metric, the threshold, the likely causes in order of probability, and the first diagnostic command for each. The runbook should also state what the alarm does not tell you — an alarm on elevated error rate does not distinguish between a code defect and a dependency outage, and the runbook should say so explicitly so the responder does not waste the first five minutes assuming the wrong cause. Alarms whose runbook entry is empty are, in practice, alarms that will be muted during the next incident.

Finally, the alarm configuration itself should be version-controlled. Alarms created by hand in the console drift, and the drift is invisible until an incident. Defining alarms in CloudFormation, Terraform, or CDK means the threshold, the period, and the missing-data treatment are all reviewable in a pull request, and it means a deleted alarm is restored by the next deployment rather than discovered missing during an outage.

8. Edge Cases and Exam Gotchas

The single most-tested gotcha is that composite alarms evaluate alarm states, not metrics. A question that asks you to reduce alert noise by combining two conditions is pointing at a composite alarm; a question that asks you to alarm on a ratio of two metrics is pointing at metric math on a single alarm. These are different features and the exam distinguishes them. If the scenario describes two independent conditions that must both be true, it is composite. If it describes one derived number, it is metric math.

The second gotcha is the INSUFFICIENT_DATA state. Any question that describes a resource being terminated, a publisher failing, or a metric going silent is testing whether you know that the default behavior does not page. The correct answer is almost always to set the missing-data treatment explicitly, and to add an alarm on the publisher itself.

The third gotcha is the statistic. A scenario that mentions a small percentage of users experiencing slow responses is describing a tail-latency problem, and the correct statistic is a percentile, not an average. A scenario that mentions a hard capacity limit is describing a maximum or a sum, not an average. The exam rarely names the statistic in the question; it describes the symptom and expects you to infer it.

The fourth gotcha is the relationship between alarms and automated actions. An alarm that triggers an Auto Scaling policy is not the same as an alarm that notifies a human, and a scenario that requires both will have an answer that configures both. Similarly, an alarm used as a stop condition on a Fault Injection Service experiment must be a real alarm with a real threshold — the exam will sometimes offer a plausible-sounding but non-existent "stop condition" resource as a distractor.

The fifth gotcha is the difference between an alarm and a health check. Route 53 health checks are a separate mechanism with their own evaluation logic, and while an alarm can be used as a health check input, they are not interchangeable. A scenario about DNS failover is asking about health checks; a scenario about paging a human is asking about alarms.

9. This vs. the Services It Gets Confused With

CloudWatch alarms are frequently confused with three other mechanisms, and the exam exploits the confusion. The first is CloudWatch Logs metric filters, which derive a metric from log events rather than from a published metric. A metric filter is the right tool when the signal you care about only exists in log content — a specific error string, a stack trace, a business event — and the resulting metric can then be alarmed on normally. The distinction is that a metric filter is a metric source, not an alarm type.

The second is EventBridge rules, which react to events rather than to metric thresholds. An EventBridge rule is the right tool for reacting to a discrete occurrence — an EC2 instance state change, an API call recorded by CloudTrail, a scheduled tick — while an alarm is the right tool for reacting to a continuous measurement crossing a boundary. A scenario that describes reacting to a state change is asking about EventBridge; a scenario that describes reacting to a threshold breach is asking about alarms.

The third is AWS Health and Trusted Advisor, which surface account-level and best-practice findings rather than metric conditions. These are complementary rather than competing: a Trusted Advisor check might tell you that an alarm is missing, but it does not replace the alarm.

MechanismInputPick it when…
CloudWatch alarmA metric or metric math expressionThe condition is a measurement crossing a boundary
Composite alarmThe states of other alarmsThe condition is a conjunction of independent signals
Logs metric filterLog event contentThe signal only exists in log text
EventBridge ruleAn eventThe condition is a discrete occurrence, not a level
Route 53 health checkEndpoint reachabilityThe condition drives DNS failover

Hands-on Lab: A Composite Alarm That Pages Only on Real Incidents

The goal of this lab is to build the alerting layer for a small service and then deliberately make it noisy, so you can see the difference a composite alarm makes. You will need an AWS account with permission to create CloudWatch alarms, SNS topics, and a Lambda function. Budget roughly forty-five minutes.

Step 1 — Create the notification channel. Create an SNS topic named svc-alerts and subscribe your email to it. Confirm the subscription. Every alarm you create in this lab will publish to this topic, so you have a single place to observe what fires.

Step 2 — Produce a metric worth alarming on. Deploy a small Lambda function behind a function URL that returns a 200 response after a configurable delay, and have it emit a custom metric in the Lab/Service namespace with two dimensions: Service=checkout and Stage=prod. Publish two metrics: RequestLatency in milliseconds and ErrorCount as a count. Use high-resolution publishing so you can evaluate on short periods.

Step 3 — Build the naive alarms. Create a latency alarm with a static threshold of 300 milliseconds on the average statistic, a one-minute period, and one evaluation period. Create an error alarm with a threshold of one error, same period and evaluation settings. Point both at the SNS topic. Now drive traffic that produces occasional slow requests but no errors, and watch the latency alarm flap. This is the failure mode the composite alarm exists to fix.

Step 4 — Build the composite alarm. Create a composite alarm whose rule expression is ALARM("latency-alarm") AND ALARM("error-alarm"). Point it at the same SNS topic, and remove the SNS action from the two child alarms so only the composite pages. Re-run the same traffic pattern. The composite should stay silent through the latency-only spikes.

Step 5 — Verify the composite fires on a real incident. Now drive traffic that produces both elevated latency and errors simultaneously. Confirm the composite transitions to ALARM and that the notification arrives. Check the composite's state history in the console to see the child alarm transitions that led to it.

Step 6 — Test the missing-data behavior. Delete the Lambda function, or stop it from publishing, and observe what happens to both child alarms. With the default missing-data treatment they will move to INSUFFICIENT_DATA and the composite will not fire. Change TreatMissingData to breaching on the error alarm and confirm the composite now fires when the publisher goes silent. This is the single most valuable thing to internalize from the lab.

Step 7 — Add an anomaly detection band. Replace the static latency threshold with an anomaly detection band on the same metric, and drive a traffic pattern with a clear daily shape. Observe how the band tracks the baseline rather than sitting at a fixed number, and note in your own words when you would prefer the band and when you would prefer the static threshold.

Step 8 — Write the runbook entry. For the composite alarm, write a short runbook entry naming the metric, the threshold, the two child conditions, the three most likely causes, and the first diagnostic command for each. Then delete the alarms and recreate them from a CloudFormation template so the configuration is version-controlled.

Scenario Question Drills

Q1. On-call engineers are fatigued by alarms firing on isolated latency spikes that self-resolve. What reduces noise while still catching real incidents?

A. Delete the latency alarm
B. A composite alarm requiring both the latency alarm AND the error-rate alarm to be in ALARM state
C. Lower the alarm threshold
D. Increase the evaluation period to 24 hours
Correct answer: B. Composite alarms let you require correlated signals before paging, cutting single-metric false positives while preserving sensitivity to genuine multi-symptom incidents.

Q2. A team wants to alarm on the percentage of requests that fail, derived from two separate metrics: total requests and error count. What should they configure?

A. A composite alarm combining the two metrics
B. A single alarm using a metric math expression that divides errors by total requests
C. Two independent alarms with the same threshold
D. A CloudWatch Logs metric filter
Correct answer: B. Composite alarms evaluate alarm states, not metrics. A derived ratio is a single metric produced by metric math, and it is alarmed on with one alarm.

Q3. An alarm on a custom metric has been sitting in INSUFFICIENT_DATA for three days and no one was paged. The metric's publisher crashed. What should have been configured?

A. A shorter alarm period
B. TreatMissingData set to breaching, plus an alarm on the publisher's own health
C. A higher threshold
D. A composite alarm with OR logic
Correct answer: B. INSUFFICIENT_DATA is a distinct state and does not page by default. Treating missing data as breaching, plus monitoring the publisher, closes the silent-gap failure mode.

Q4. A latency SLO requires that 99% of requests complete within 300ms. Which statistic should the alarm use?

A. Average
B. Maximum
C. p99
D. Sum
Correct answer: C. An SLO expressed as a percentile must be alarmed on with the matching percentile statistic. Average hides the tail; maximum is hypersensitive to a single outlier.

Q5. A metric's normal value varies significantly by time of day, and a static threshold either fires during peak hours or misses real problems overnight. What is the appropriate alarm configuration?

A. Two static alarms, one for peak and one for off-peak
B. An anomaly detection band alarm
C. A composite alarm with OR logic
D. A metric filter on the logs
Correct answer: B. Anomaly detection learns the metric's seasonal baseline and alarms on deviation from the predicted band, which a static threshold cannot track.

Q6. A team needs an alarm that fires only if a metric breaches its threshold on three consecutive one-minute periods. Which settings achieve this?

A. Period 60s, EvaluationPeriods 1, DatapointsToAlarm 1
B. Period 60s, EvaluationPeriods 3, DatapointsToAlarm 3
C. Period 180s, EvaluationPeriods 1, DatapointsToAlarm 1
D. Period 60s, EvaluationPeriods 3, DatapointsToAlarm 1
Correct answer: B. EvaluationPeriods sets the window and DatapointsToAlarm sets how many of those periods must breach. Requiring all three gives the consecutive-breach behavior.

Q7. A service must be detected as degraded within ten seconds of the condition occurring. What is required?

A. A standard-resolution metric with a 10-second period
B. A high-resolution custom metric with a 10-second alarm period
C. A composite alarm over two standard alarms
D. A CloudWatch Logs metric filter
Correct answer: B. Standard metrics are published at one-minute granularity, so sub-minute detection requires a high-resolution custom metric and a matching short alarm period.

Q8. An alarm fires correctly but the on-call engineer finds the fleet at 30% CPU with every request failing. What does this indicate about the alarm design?

A. The threshold is too low
B. The alarm watches an infrastructure metric that is a poor proxy for the user-visible symptom
C. The period is too short
D. The alarm should be deleted
Correct answer: B. CPU utilization is not a proxy for request success. The fix is to add a service-level alarm on the user-visible symptom, not to remove the infrastructure alarm.

Q9. A team wants to react automatically when an EC2 instance transitions to the stopped state. Which mechanism is appropriate?

A. A CloudWatch alarm on the instance's CPU metric
B. An EventBridge rule matching the instance state-change event
C. A composite alarm
D. A CloudWatch Logs metric filter
Correct answer: B. A state change is a discrete occurrence, which is EventBridge's domain. Alarms react to continuous measurements crossing a boundary.

Q10. A team wants to alarm on occurrences of a specific error string that appears only in application log output. What should they configure?

A. A composite alarm over the application's existing alarms
B. A CloudWatch Logs metric filter that produces a metric, then an alarm on that metric
C. An anomaly detection band on the Lambda error metric
D. An EventBridge rule on the log group
Correct answer: B. A metric filter derives a metric from log content; the alarm then watches that derived metric normally. The filter is a metric source, not an alarm type.

Q11. A composite alarm references two child alarms, but it never fires even when both children are in ALARM. What is the most likely cause?

A. Composite alarms cannot reference more than one child
B. The rule expression uses AND but the scenario requires OR, or the children are not actually both in ALARM at the same evaluation
C. Composite alarms require high-resolution metrics
D. The composite alarm needs its own SNS topic
Correct answer: B. Composite alarms evaluate the rule expression over child states at evaluation time. If the children never overlap in ALARM, an AND expression will never be true.

Q12. An alarm drives an automated remediation that restarts a service. What should accompany it?

A. Nothing — automated remediation should be silent
B. A notification action, plus an alarm on the remediation's own success
C. A second identical remediation action
D. A longer evaluation period
Correct answer: B. A silent automated action is indistinguishable from a bug. Pair remediation with notification and monitor whether the remediation actually worked.

Q13. A team wants to compare this month's p99 latency against the same month last year. What data will they actually be looking at?

A. One-minute datapoints, retained for fifteen months
B. One-hour aggregated datapoints, since one-minute data is retained for only fifteen days
C. Raw high-resolution datapoints
D. The data is not retained that long
Correct answer: B. Retention is tiered: one-minute data for fifteen days, five-minute for sixty-three days, one-hour for four hundred and fifty-five days. Year-over-year comparison uses hourly aggregates.

Q14. A Fault Injection Service experiment must abort automatically if production stability degrades. What does the experiment template require?

A. A manual abort button
B. A stop condition referencing a CloudWatch alarm with a real threshold
C. A composite alarm with OR logic
D. A Route 53 health check
Correct answer: B. FIS stop conditions are CloudWatch alarms. The alarm must be a real alarm with a real threshold, and it aborts the experiment when it enters ALARM.

Q15. A team wants to page only when the error budget burn rate is high enough to exhaust a monthly budget within hours. Which pattern fits?

A. A static alarm on raw error count
B. A fast-burn alert using a short window and a high burn rate, implemented with metric math over request and error counts
C. A composite alarm over CPU and memory alarms
D. An anomaly detection band on the error count
Correct answer: B. Burn-rate alerting derives a rate from request and error counts via metric math, then alarms on a short window with a high rate to catch acute outages.

Peek into Tomorrow

Everything in today's material shares one assumption: the metric you are alarming on is a faithful proxy for what the user actually experiences. That assumption is doing a lot of work, and it is worth interrogating before moving on. A fleet can report healthy CPU, healthy memory, healthy error rates at the load balancer, and still be serving a broken login page — because the failure is in a JavaScript bundle, a DNS record, a CDN configuration, or a third-party dependency that no internal metric observes. The alarm fires on the signal it was given, and the signal was never the user's experience in the first place.

Tomorrow's topic attacks that gap directly, with scripted probes that run from outside your infrastructure and exercise the actual user flow rather than the infrastructure underneath it. The open question today leaves unresolved is where the boundary sits: if synthetic probes can catch what internal metrics miss, why keep the internal metrics at all, and how do the two layers divide responsibility without duplicating each other's alerts? The answer turns out to be about what each layer can and cannot see, and it reframes the alarm design decisions from today as one half of a two-layer monitoring model.

Sources