Day 30 of 70 · Week 5
Day 30 / 70 Week 5 of 14 Phase 3: SRE Observability, Resilience & DR

CloudWatch Synthetics & Canaries for SLO Monitoring

🕑 ~58 min read · 2 services covered
CloudWatch Synthetics Canaries

Recap: From Alarm Logic to User-Visible Truth

Day 29 established that a CloudWatch alarm is only as good as the metric feeding it, and that composite alarms exist because single-metric thresholds generate noise. The AND/OR logic of a composite alarm — page only when high latency AND high error rate are both true — is a real advance over paging on either signal alone, and anomaly detection bands let you alarm on a metric whose healthy shape you cannot express as a static number. But every one of those mechanisms shares an assumption that deserves scrutiny: the metric being evaluated is produced by something inside your infrastructure, measuring something your infrastructure can see.

Today extends that thread by attacking the assumption directly. A composite alarm built on ALB target response time and Lambda error rate can be perfectly tuned and still miss a broken login page, because the failure lives in a layer no internal metric observes — a DNS record, a CDN edge rule, a third-party identity provider, a JavaScript bundle that fails to parse in one browser. Synthetics canaries are the outside-in complement to the inside-out alarm stack: scripted probes that exercise the real user path from outside your VPC and publish their own metrics into the same CloudWatch namespace your alarms already read. The alarm machinery from Day 29 does not change; what changes is that you now have a signal that can actually see the failure.

Foundations You'll Need Today

Today's topic sits on top of a handful of ideas that the rest of this curriculum treats as background knowledge. If you are coming from a foundational understanding of AWS rather than hands-on architecture work, these are the pieces worth having straight before the rest of the day makes sense.

CloudWatch Metrics, Alarms, and Composite Alarms

CloudWatch is AWS's monitoring service, and its core primitive is the metric: a named number that some AWS service publishes over time, such as the number of requests an application received in the last minute or the average response time. Metrics are grouped into namespaces so that metrics from different services do not collide, and each metric can carry dimensions — labels that narrow it down, such as which specific load balancer or function the number describes. An alarm is a rule that watches one metric and changes state when the metric crosses a threshold you set. It has a period (how wide each measurement window is), an evaluation (how many consecutive windows must breach before the alarm fires), and a missing-data treatment (what to do when no data point arrives at all — treat it as healthy, as breaching, or as an unknown state called INSUFFICIENT_DATA). A composite alarm is an alarm whose rule is built from the states of other alarms, using AND/OR logic, so that you can page someone only when two conditions are true at once rather than on either one alone. Today's material assumes you can read an alarm definition and reason about its period and missing-data behavior, because that is exactly where canary alarms go wrong.

SLOs and Error Budgets

An SLO, or service level objective, is a target you set for how reliable a service should be, expressed as a percentage over a time window — for example, "99.9% of checkout attempts succeed in a calendar month." The complement of that target is the error budget: the amount of failure you have implicitly decided is acceptable. At 99.9%, the budget is 0.1% of attempts, which works out to roughly 43 minutes of full outage per month. The error budget is useful because it turns reliability into a number you can spend and track rather than a vague aspiration, and it gives you a way to judge whether a monitoring gap matters: a failure that consumes a large fraction of the budget before anyone notices is a monitoring problem, while one that consumes a rounding error is not. Today's discussion of canary frequency is really a discussion about how much of the error budget a single undetected incident is allowed to burn.

Route 53 Health Checks

Route 53 is AWS's DNS service — the system that translates a human-readable name like example.com into the numeric address of a server. A health check is a feature of Route 53 that periodically requests a specific endpoint from several locations around the world and records whether it responded with a healthy status code. DNS records can be configured to point traffic only at endpoints whose health checks are passing, which is how automatic failover between regions is usually built. A health check is deliberately shallow: it asks "did this URL answer correctly?" and nothing more. It cannot log in, follow a multi-step flow, or check what a page actually rendered. Today's topic exists largely because that shallowness leaves a real gap.

IAM Roles and Trust Policies

IAM is AWS's identity and permissions system. A role is an identity that is not tied to a person — instead, it is assumed temporarily by something that needs to act, such as an AWS service running on your behalf. Every role has two distinct pieces of configuration. The trust policy answers "who is allowed to assume this role?" and is what lets, say, the CloudWatch Synthetics service take on the role when it runs your canary. The permissions policy answers "once assumed, what is this role allowed to do?" — for example, write an object into a specific S3 bucket or publish a metric into CloudWatch. Both must be correct for anything to work, and a mistake in either produces a failure that looks like a service problem rather than a permissions problem. Today's lab walks through creating exactly such a role, and one of the failure modes discussed later is what happens when its permissions are wrong.

VPCs and Subnets

A VPC, or virtual private cloud, is a private, isolated network that you define inside AWS. It is the boundary within which your resources live and the thing that determines what can talk to what. A VPC is divided into subnets, each of which is a slice of the VPC's address range placed in a specific availability zone; subnets are typically labeled public or private depending on whether they have a route to the internet. Resources launched inside a VPC reach the outside world only through the routes and gateways you configure, which is the whole point — it lets you keep databases and internal services unreachable from the internet. The relevance today is that a canary can be configured to run either outside any VPC, where it sees the same public path a real user takes, or inside one, where it sees only your private network. Choosing the wrong one means the canary is testing a path no real user ever uses.

With that grounding, here is why CloudWatch Synthetics exists and what problem it actually solves.

1. Why This Is on the Exam

SAP-C02 scenario questions are written around a specific kind of gap: the architecture looks healthy on every dashboard the team has built, yet the business is losing money. The exam's Domain 2 (Design for New Solutions) and Domain 4 (Continuous Improvement for Existing Solutions) both probe whether you can identify what a monitoring strategy is blind to, and Synthetics is the canonical answer whenever the blind spot is the user's actual experience rather than the server's internal state. A question that describes a multi-region active-active application with comprehensive CloudWatch dashboards, X-Ray tracing, and composite alarms, and then asks what additional monitoring would detect a regional edge failure before customers call the support line, is pointing at canaries.

The reason this is a durable exam topic rather than a trivia item is that it sits at the intersection of two things the exam cares about deeply: SLO definition and blast-radius reasoning. An availability SLO expressed as "99.9% of checkout attempts succeed" cannot be measured by any metric your application emits, because your application only knows about the requests that reached it. Requests that died at the DNS resolver, at the CDN, or in the client's browser never appear in your logs at all. Synthetics closes that measurement gap by generating known-good traffic on a schedule and asserting on the result, which turns an unmeasurable SLO into an observable one. Expect questions that give you an SLO target and ask which monitoring approach can actually attest to it.

There is also a cost-and-complexity dimension the exam likes to test. Canaries are not free, they are not instantaneous, and they cannot cover every path. A scenario that asks you to monitor a 40-step internal admin workflow with a canary is testing whether you understand that canaries are best reserved for high-value, stable, externally reachable user journeys. The exam will offer you a canary as a distractor for problems that are better solved with distributed tracing, log-based metrics, or synthetic load testing, and it will offer you those alternatives as distractors for problems that only a canary can solve. Knowing the boundary is the point.

2. How a Canary Actually Runs

A canary is not a ping. It is a small Node.js or Python script that CloudWatch Synthetics executes on a schedule inside an AWS-managed, serverless execution environment. The script receives a Synthetics runtime library that wraps Puppeteer (for browser-based canaries) or raw HTTP clients (for API canaries), and it is expected to do three things: drive the target, assert on what it observes, and emit structured results. The runtime handles the rest — launching the headless browser, capturing screenshots and HAR files on failure, uploading artifacts to S3, and publishing metrics into the CloudWatch namespace CloudWatchSynthetics.

The execution model matters for reasoning about cost and reliability. Each canary run is an isolated invocation with its own ephemeral environment; there is no persistent browser session between runs, which means every run pays the cold-start cost of launching Chromium. That cost is why the minimum useful interval is generally measured in minutes rather than seconds, and why a canary that takes 45 seconds to execute at a 1-minute frequency is effectively running continuously. The runtime also enforces a timeout, and a canary that exceeds it is recorded as a failure with a distinct status code rather than a generic error — a distinction that matters when you are writing alarms, because a timeout usually means the target is slow rather than down.

Results flow out on two channels. The first is metrics: every run publishes SuccessPercent, Duration, Failed, and 2xx/4xx/5xx counts, plus any custom metrics your script emits via the Synthetics library. The second is artifacts: screenshots, logs, and HAR files written to the S3 bucket you nominate when you create the canary. The metrics are what your alarms consume; the artifacts are what a human consumes at 3 a.m. when the alarm fires and they need to know whether the page rendered a login form or a CloudFront error page. Designing a canary without thinking about the artifact side is a common mistake — the metric tells you something broke, the screenshot tells you what.

Finally, canaries can run either on a schedule or on demand, and they can be triggered by other systems. The scheduled mode is the default and the one exam scenarios describe. On-demand invocation is useful for validating a deployment immediately after it completes, and for the "smoke test" step in a pipeline. Both modes publish to the same metrics, which means an alarm cannot distinguish a scheduled failure from a pipeline-triggered one unless you tag the runs — worth knowing when you are designing a deployment gate that should not page the on-call engineer.

3. The Core Decision Boundary: What Deserves a Canary

The single fork every Synthetics scenario question hinges on is whether the thing you need to monitor is observable from inside your infrastructure. If it is, a canary is redundant and expensive. If it is not, a canary is the only tool that can see it. The practical test is to ask where the failure would occur: if the failure would occur in code you deploy and instrument, internal metrics and traces will catch it; if the failure would occur in the path between the user and your code — DNS, CDN, load balancer configuration, TLS certificates, third-party dependencies, client-side rendering — only an outside-in probe will catch it.

This boundary has a second dimension that is easy to miss: stability. A canary is a scripted assertion, which means it encodes an expectation about what the target should look like. If the target's UI changes weekly, the canary becomes a maintenance burden that fails for reasons unrelated to availability, and the team learns to ignore its alarms. The exam will sometimes present a fast-moving internal tool as a canary candidate and expect you to recognize that the maintenance cost outweighs the benefit. Canaries earn their keep on stable, high-value, externally reachable journeys: login, search, checkout, the primary API endpoint a partner integrates against.

The table below is the decision boundary in its most compressed form. Read it as a set of rules for eliminating distractors rather than as a taxonomy to memorize.

Signal you needBest toolWhy not a canary
Is the service process alive and responding to health checks?Route 53 health checks, ALB target healthCanary adds cost and latency for a question already answered
Which internal service is slow on a specific request?X-Ray distributed tracingCanary sees the total duration, not the per-hop breakdown
Is the login page reachable and functional for a real user?Synthetics canaryNo internal metric observes the client-side path
Are error rates elevated in application logs?CloudWatch Logs metric filtersCanary samples one path; logs cover all traffic
Will the system hold up under 10x load?Load testing (Distributed Load Testing on AWS)Canary is a single synthetic user, not a load generator
Is a third-party payment provider degraded?Synthetics canary against the integration pathProvider's own status page is not an alarm source

4. Configuration Modes and Their Tradeoffs

Once you have decided a canary is warranted, the configuration knobs determine what it costs and how quickly it detects a problem. The first and most consequential is the blueprint. Synthetics ships pre-built blueprints for the common cases — Heartbeat Monitoring (a simple URL fetch and status assertion), API Canary (a scripted HTTP request sequence with response validation), Broken Link Checker (crawls a page and verifies every link), Visual Monitoring (compares a screenshot against a baseline and flags pixel differences), and GUI Workflow Builder (records a browser session and replays it). Choosing a blueprint is choosing how much of the assertion logic you write yourself. Heartbeat is nearly free to configure and catches availability failures; GUI Workflow catches functional regressions but requires you to maintain a recorded script that breaks whenever the UI changes.

The second knob is frequency, and it is where the cost conversation actually happens. A canary running every minute costs roughly sixty times what the same canary running every hour costs, and the detection latency for a failure is bounded below by the frequency. The right frequency is a function of your SLO: if your availability SLO is 99.9% measured monthly, a failure that lasts five minutes consumes about 0.7% of your monthly error budget, so a one-minute canary that detects it in under two minutes is proportionate. A five-minute canary on the same SLO would let a single incident consume several percent of the budget before anyone noticed. The exam rarely asks you to compute this, but it does ask you to reason about detection latency versus cost.

The third knob is the runtime version and the script itself. Synthetics runtimes are versioned (for example syn-nodejs-puppeteer-7.0 and later), and AWS deprecates older runtimes on a published schedule. A canary pinned to a deprecated runtime will eventually stop running, which is a silent monitoring failure — the worst kind, because the absence of alarms looks like health. The fourth knob is the S3 artifact location and retention, which determines how much forensic evidence you have after an incident and how much you pay for storing screenshots. The fifth is the VPC configuration: a canary can run inside a VPC to reach private endpoints, but doing so requires the canary's execution role to have ENI permissions and the subnets to have a route to the target, and it removes the canary's ability to test the public path. Most user-facing canaries should run outside the VPC precisely because the public path is what you are trying to verify.

5. Sizing, Limits and Quotas

The numbers that matter for Synthetics fall into three groups: execution limits, cost drivers, and integration limits. On execution, each canary run has a configurable timeout with a default of 30 seconds and a maximum of 15 minutes, and the runtime itself imposes a hard ceiling on how long a single run can occupy its environment. A canary whose script routinely approaches the timeout is a canary that will produce false failures under load, so the practical guidance is to keep the scripted path short and assert early rather than driving a long workflow.

On cost, the two drivers are the number of runs and the duration of each run. Synthetics pricing is per canary run, with a separate charge for the browser-based runs that use more memory. This is why the frequency decision from the previous section is the dominant cost lever: doubling the frequency doubles the bill, and a canary that takes 90 seconds to execute costs more per run than one that takes 15 seconds. The S3 artifacts add a smaller, usually negligible storage cost, but a canary that captures a full HAR file on every run rather than only on failure can accumulate meaningful storage over months.

On integration, the limits that bite in practice are the CloudWatch metrics retention (standard resolution metrics are retained for 15 months, but the canary's own success metrics are published at one-minute resolution and are best alarmed on with a short evaluation period) and the alarm evaluation window. A canary running every five minutes publishes a data point every five minutes; an alarm with a one-minute period and a single evaluation will see missing data for four out of five minutes and may enter INSUFFICIENT_DATA. The correct pattern is to set the alarm period to match the canary frequency and to treat missing data as breaching, because a canary that stops running is itself an incident. The table below collects the numbers worth remembering.

ParameterTypical value / limitPractical implication
Default run timeout30 secondsShort scripts; assert early
Maximum run timeout15 minutesLong workflows are possible but costly and fragile
Minimum useful frequency1 minuteDetection latency floor for availability failures
Metric namespaceCloudWatchSyntheticsAlarms read SuccessPercent and Duration from here
Artifact destinationS3 bucket you nominateScreenshots and HAR files for post-incident forensics
Runtime lifecycleVersioned, deprecated on a published schedulePinned old runtimes silently stop running

6. Failure Modes and What They Look Like in Production

The most dangerous canary failure is not a red alarm; it is a canary that has quietly stopped running. This happens when the runtime version is deprecated, when the execution role's permissions are revoked, when the S3 artifact bucket is deleted, or when the canary's schedule is disabled during a maintenance window and never re-enabled. In every case the symptom is identical: the SuccessPercent metric stops receiving data points, and an alarm configured with the default missing-data treatment may sit in INSUFFICIENT_DATA indefinitely without paging anyone. The first diagnostic move when you suspect this is to check the canary's last run timestamp in the Synthetics console, not the alarm state.

The second failure mode is the false positive caused by target drift. A GUI Workflow canary that asserts on a specific button label will fail the moment a designer renames the button, and the failure will look exactly like an availability incident. The distinguishing signal is that the canary's Duration metric stays normal while SuccessPercent drops — a real outage usually shows both degrading, because a slow or unreachable target takes longer to fail. When you see success drop with duration flat, the first thing to check is the screenshot artifact, which will show the new UI rather than an error page.

The third failure mode is the canary that is too aggressive and becomes a load source. A canary that drives a full checkout flow every minute from three regions is generating real orders and real payment authorizations, and if the target is not idempotent you have created a business problem. The mitigation is to design the canary against a dedicated test account or a sandbox path, and to make the scripted actions idempotent where possible. The fourth failure mode is credential expiry: a canary that logs in with a stored credential will start failing the day that credential rotates, and the failure will be indistinguishable from an authentication outage unless you monitor the canary's own error messages. Storing canary credentials in Secrets Manager and having the script fetch them at runtime is the pattern that avoids this.

7. The SRE Angle: Canaries as SLO Instruments

From an SRE perspective, a canary is not primarily a monitoring tool; it is a measurement instrument for an SLO that would otherwise be unmeasurable. The distinction matters because it changes how you configure it. If you are using a canary to page an on-call engineer, you want it to be sensitive and you accept some false positives. If you are using it to compute an availability SLO, you want it to be representative of real user traffic, which means the canary's success rate should track the real user success rate closely enough that you can defend the number in a post-incident review. Those two goals pull in different directions, and mature teams often run two canaries: a fast, sensitive one for alerting and a slower, more representative one for SLO reporting.

The alarm design follows from the SLO. A canary publishing SuccessPercent every minute should have an alarm on SuccessPercent < 100 with a period of one minute and an evaluation of one or two datapoints, treating missing data as breaching. That alarm should feed into the composite alarm structure from Day 29 rather than paging directly, so that a single canary failure in one region does not wake someone up when the other regions are healthy. The composite pattern — canary failure AND elevated real-user error rate — is the one that distinguishes a genuine outage from a canary-specific problem, and it is the pattern the exam is most likely to describe.

The runbook shape for a canary alarm is short and specific. First, open the screenshot artifact from the failing run; it answers the question "what did the user see" faster than any log query. Second, check whether the failure is isolated to one region or global, which distinguishes a regional edge problem from an application problem. Third, check the canary's own last successful run timestamp to rule out the silent-stop failure mode. Fourth, if the screenshot shows a valid page, check whether the assertion itself is stale. That four-step sequence resolves the large majority of canary alarms, and writing it down before the first incident is what turns a canary from a noise source into a useful signal.

8. Edge Cases and Exam Gotchas

The gotchas cluster around three themes: what a canary cannot see, what it costs to run, and how it interacts with the rest of the observability stack. On the first theme, a canary cannot see the experience of a real user on a slow connection, cannot see a failure that only affects one browser, and cannot see a failure that only affects authenticated users unless you script the authentication. It also cannot see a failure that occurs between the canary's region and the user's region if the canary runs from a single region — which is why multi-region canaries are the pattern for multi-region applications. A single-region canary on a multi-region application will report success while an entire region's users are down.

On the second theme, the cost trap is frequency. A team that sets a canary to run every minute "to be safe" and then adds a second canary for each of ten critical paths has created a monitoring bill that can rival the application's compute bill. The exam will sometimes present this as a scenario where the correct answer is to reduce frequency or consolidate paths rather than to add more canaries. The related trap is artifact retention: capturing a full HAR file on every successful run rather than only on failure is a common misconfiguration that quietly accumulates storage cost.

On the third theme, the interaction gotchas are about alarm wiring. A canary alarm with a one-minute period on a five-minute canary will flap between OK and INSUFFICIENT_DATA, and if missing data is treated as not-breaching the alarm will never fire. A canary alarm that pages directly rather than through a composite will generate noise during regional degradations. A canary that runs inside a VPC cannot test the public path, so a VPC-configured canary on a public-facing application is testing the wrong thing. And a canary whose execution role lacks s3:PutObject on the artifact bucket will fail to publish artifacts, which means the alarm fires but the screenshot is missing — a failure mode that is invisible until you need the evidence.

9. Synthetics vs. the Services It Gets Confused With

Synthetics is most often confused with three things: Route 53 health checks, X-Ray tracing, and load testing. The confusion with health checks is the most common, because both are outside-in probes that run on a schedule. The difference is depth. A health check asks "does this endpoint return a healthy status code" and nothing more; it cannot log in, cannot follow a redirect chain, cannot assert on page content, and cannot measure the time to interactive. A canary can do all of those things. The exam will offer a health check as the answer to a question about detecting a broken login flow, and the correct answer is a canary, because the health check would return 200 from the load balancer while the login page renders an error.

The confusion with X-Ray is about direction. X-Ray traces a request that already reached your application, breaking it into per-hop latency and identifying which internal service is slow. A canary generates a request from outside and measures the total time to a successful assertion. They are complementary: a canary tells you the user-visible path is broken, and X-Ray tells you which internal hop is responsible. The exam will sometimes describe a scenario where the canary is failing and ask what to do next, and the correct answer is to use X-Ray to find the slow hop — not to add more canaries.

The confusion with load testing is about volume. A canary is a single synthetic user; a load test is thousands. A canary detects availability and functional regressions; a load test detects capacity limits. The exam will offer a canary as the answer to a question about validating that a system can handle a traffic spike, and the correct answer is a load testing tool. The table below is the disambiguation in its most compressed form.

ServiceWhat it answersPick it when…
Route 53 health checkIs the endpoint returning a healthy status?You need fast, cheap failover routing
CloudWatch SyntheticsCan a real user complete the journey?The failure would occur outside your infrastructure
AWS X-RayWhich internal hop is slow or failing?The request reached your app and you need the breakdown
Distributed Load TestingDoes the system hold under N concurrent users?You are validating capacity, not availability
CloudWatch Logs metric filtersWhat is the error rate across all real traffic?You need aggregate behavior, not a single path

Hands-on Lab: A Checkout Canary with an SLO Alarm (45 min)

This lab builds a canary that exercises a login-and-checkout flow every five minutes, publishes its results to CloudWatch, and wires an alarm that pages only when the canary fails and real-user error rate is also elevated. You will need an AWS account with permissions to create Synthetics canaries, IAM roles, CloudWatch alarms, and an S3 bucket. If you do not have a real checkout flow, substitute a public test site or a simple static page you control; the mechanics are identical.

  1. Create the artifact bucket. In S3, create a bucket named for your canary artifacts, for example synthetics-artifacts-<account-id>-<region>. Enable default encryption and a lifecycle rule that expires objects after 30 days. This bucket holds the screenshots and HAR files that make the canary useful during an incident.
  2. Create the execution role. Create an IAM role that Synthetics can assume, with a trust policy for lambda.amazonaws.com and permissions for s3:PutObject on the artifact bucket, cloudwatch:PutMetricData, and logs:CreateLogStream/logs:PutLogEvents. If your canary needs to read a credential from Secrets Manager, add secretsmanager:GetSecretValue scoped to that specific secret.
  3. Create the canary. In the CloudWatch console, choose Synthetics, then Create canary. Select the GUI Workflow Builder blueprint if you want to record the flow, or the API Canary blueprint if you prefer to script the HTTP calls. Name it checkout-canary, point it at your target URL, and select the artifact bucket and execution role from the previous steps.
  4. Script the flow. For a GUI canary, record the login and checkout steps, then edit the generated script to replace any hard-coded credentials with a Secrets Manager fetch. For an API canary, write the request sequence explicitly and assert on the response body, not just the status code — a 200 response containing an error message is a failure the status check would miss.
  5. Set the schedule. Configure the canary to run every five minutes. Note the tradeoff you are accepting: a failure that begins immediately after a run will not be detected for up to five minutes, which on a 99.9% monthly SLO consumes roughly 0.7% of the error budget per incident. If that is too much, reduce the interval to one minute and accept the higher cost.
  6. Run it once manually. Use the Run now action and confirm the canary succeeds. Open the artifact in S3 and verify the screenshot shows the expected page. This step catches permission and networking problems before they become alarm noise.
  7. Create the alarm. Create a CloudWatch alarm on the SuccessPercent metric in the CloudWatchSynthetics namespace, dimensioned by your canary name. Set the period to five minutes to match the canary frequency, the statistic to Average, the threshold to less than 100, and the missing-data treatment to breaching. The last setting is the one that catches the silent-stop failure mode.
  8. Wire the composite alarm. Create a second alarm on your application's real-user error rate (for example, ALB HTTPCode_Target_5XX_Count or a Lambda error metric). Then create a composite alarm with the rule ALARM(canaryAlarm) AND ALARM(errorRateAlarm). Route the composite to your paging channel and route the individual canary alarm to a low-severity channel.
  9. Break it deliberately. Temporarily point the canary at a URL that returns a 500, or block the canary's egress with a security group change. Confirm the canary alarm fires, the composite does not (because real-user error rate is still healthy), and the screenshot artifact shows the failure. Then restore the target and confirm the alarm returns to OK.
  10. Write the runbook. Document the four-step response: open the screenshot, check regional scope, check the last successful run timestamp, and verify the assertion is not stale. Store it where the on-call engineer will find it at 3 a.m.

Clean up by deleting the canary, the alarms, the execution role, and the artifact bucket. The lab's real output is the runbook and the alarm wiring pattern, both of which transfer directly to production.

Scenario Question Drills (20 min)

Q1. Internal CloudWatch metrics show healthy servers, but customers report the login page is broken. What monitoring gap does this reveal?

A. Missing X-Ray tracing
B. No outside-in synthetic monitoring of the actual user flow
C. Missing VPC Flow Logs
D. Insufficient EC2 instance count
Correct answer: B. Server-side health metrics don't verify end-to-end user experience; Synthetics canaries probe from outside the infrastructure, catching failures (DNS, CDN, front-end bugs) invisible to internal metrics.

Q2. A multi-region active-active application has comprehensive CloudWatch dashboards and X-Ray tracing, but an edge failure in one region went undetected for 20 minutes. What should be added?

A. More X-Ray sampling rules
B. Synthetics canaries running from multiple regions against the user-facing endpoints
C. A larger ALB target group
D. CloudTrail data events
Correct answer: B. A single-region canary would have reported success while one region's users were down. Multi-region canaries are required to observe per-region user-visible availability.

Q3. A canary alarm is configured with a one-minute period, but the canary runs every five minutes. The alarm flaps between OK and INSUFFICIENT_DATA and never pages. What is the fix?

A. Increase the canary frequency to one minute
B. Set the alarm period to five minutes to match the canary frequency, and treat missing data as breaching
C. Remove the alarm and rely on the dashboard
D. Switch the statistic from Average to Sum
Correct answer: B. Alarm period should match the metric's publication cadence. Treating missing data as breaching also catches the silent-stop failure mode where the canary stops running entirely.

Q4. A team wants to detect when a third-party payment provider's integration path is degraded, before customers complain. What is the most direct approach?

A. Subscribe to the provider's status page RSS feed
B. A Synthetics canary that exercises the integration path and asserts on the response
C. Increase the Lambda timeout on the payment handler
D. Enable VPC Flow Logs on the payment subnet
Correct answer: B. A canary against the integration path measures the actual dependency from your perspective, which is what matters. A status page is not an alarm source and may lag or miss your specific path.

Q5. A canary's SuccessPercent drops to 0 while its Duration metric stays normal. What is the most likely cause?

A. A genuine availability outage
B. Target drift — the UI changed and the canary's assertion is now stale
C. The canary's runtime was deprecated
D. The artifact bucket was deleted
Correct answer: B. A real outage usually degrades both success and duration, because a slow or unreachable target takes longer to fail. Success dropping with duration flat points at an assertion mismatch, and the screenshot artifact will confirm it.

Q6. A canary that logs in with a stored credential starts failing every run on the same day each quarter. What is the most likely cause and the correct fix?

A. The canary runtime is deprecated; upgrade it
B. The credential rotated; store it in Secrets Manager and have the canary fetch it at runtime
C. The S3 artifact bucket expired; recreate it
D. The canary frequency is too high; reduce it
Correct answer: B. Hard-coded credentials in a canary script fail silently when the credential rotates. Fetching from Secrets Manager at runtime decouples the canary from the rotation schedule.

Q7. A team wants to validate that a new deployment did not break the checkout flow, immediately after the pipeline completes. What is the appropriate use of Synthetics?

A. Increase the canary frequency to one minute permanently
B. Invoke the canary on demand as a post-deployment smoke test, and gate the pipeline on its result
C. Replace the canary with a Route 53 health check
D. Disable the canary during deployments
Correct answer: B. On-demand invocation is the deployment-gate pattern. Tag the runs so the pipeline-triggered failures do not page the on-call engineer.

Q8. A canary is configured to run inside a VPC to reach a private endpoint, but the application it monitors is public-facing. What is wrong with this configuration?

A. Nothing — VPC configuration is always preferred
B. The canary is testing the private path, not the public path users actually take
C. VPC canaries cannot publish metrics
D. VPC canaries cannot use Puppeteer
Correct answer: B. A VPC-configured canary bypasses DNS, CDN, and edge layers — exactly the layers most likely to fail for a public-facing application. Run user-facing canaries outside the VPC.

Q9. A canary alarm fires, but the S3 artifact for the failing run is missing. What is the most likely cause?

A. The canary's execution role lacks s3:PutObject on the artifact bucket
B. The canary frequency is too low
C. The alarm period is misconfigured
D. The runtime version is too new
Correct answer: A. Missing artifacts almost always trace to IAM permissions on the artifact bucket. This failure mode is invisible until you need the evidence, which is why the lab verifies artifacts on the first manual run.

Q10. A team wants to detect a broken link on a marketing site with hundreds of pages. Which Synthetics blueprint fits best?

A. Heartbeat Monitoring
B. Broken Link Checker
C. Visual Monitoring
D. API Canary
Correct answer: B. The Broken Link Checker blueprint crawls a page and verifies every link, which is exactly the job. Heartbeat only checks one URL; Visual Monitoring compares screenshots; API Canary scripts HTTP requests.

Q11. A canary that drives a full checkout flow every minute from three regions is generating real orders. What is the correct mitigation?

A. Increase the frequency to reduce the number of orders
B. Point the canary at a dedicated test account or sandbox path, and make the scripted actions idempotent where possible
C. Disable the canary entirely
D. Switch to a Route 53 health check
Correct answer: B. A canary that mutates production state is a business problem, not a monitoring solution. Sandbox paths and idempotent actions keep the canary useful without side effects.

Q12. A canary has stopped running entirely — no failures, no successes, no data points. Which diagnostic step should come first?

A. Check the alarm state in CloudWatch
B. Check the canary's last run timestamp in the Synthetics console
C. Increase the canary frequency
D. Recreate the artifact bucket
Correct answer: B. The alarm may sit in INSUFFICIENT_DATA without paging. The canary's last run timestamp is the authoritative signal that it has stopped, and it points at the cause (deprecated runtime, revoked role, disabled schedule).

Q13. A team wants to detect a visual regression — a CSS change that breaks the layout of the homepage — before users notice. Which Synthetics blueprint fits?

A. Heartbeat Monitoring
B. Visual Monitoring
C. Broken Link Checker
D. API Canary
Correct answer: B. Visual Monitoring compares a screenshot against a baseline and flags pixel differences, which is the only blueprint that can detect a layout regression that leaves the page functional.

Q14. A canary is failing and the team wants to know which internal service is responsible for the slowdown. What is the correct next step?

A. Add more canaries
B. Use X-Ray distributed tracing to break the request into per-hop latency
C. Increase the canary timeout
D. Switch the canary to a health check
Correct answer: B. A canary measures total time to a successful assertion; it cannot attribute the delay to a specific internal hop. X-Ray is the tool that answers the "which service" question.

Q15. A team wants to validate that the system can handle a 10x traffic spike before a major launch. What is the appropriate tool?

A. A Synthetics canary running every minute
B. A distributed load testing tool
C. A Route 53 health check
D. A composite alarm
Correct answer: B. A canary is a single synthetic user; it detects availability and functional regressions, not capacity limits. Load testing is the tool for validating that the system holds under concurrent load.

Peek into Tomorrow

Today's canary can tell you that the checkout flow is broken and roughly how long it took to fail, but it cannot tell you why. The canary measures the total time from the first byte of the request to the final assertion, and when that number jumps from 800 milliseconds to 4 seconds, the canary's job is done — it has raised the alarm. The question it leaves entirely unanswered is which of the dozen services the request touched is responsible for the extra 3.2 seconds. That question is the one the on-call engineer actually has to answer at 3 a.m., and it is the question a canary is structurally incapable of answering, because it observes the request from outside the system rather than from within it.

Tomorrow's topic, X-Ray distributed tracing, is the inside-out complement. It instruments the request as it flows through each service, producing a service map that shows per-hop latency and error hotspots, and it lets you drill into a specific slow downstream call — the classic example being a slow DynamoDB query inside a Lambda invocation that the canary only sees as a slow checkout. The open question worth carrying forward is how you control the cost of that instrumentation: tracing every request at full fidelity is expensive, and the sampling rules that make it affordable are themselves a design decision with real consequences for what you can and cannot see after an incident.

Sources