Day 31 of 70 · Week 5
Day 31 / 70 Week 5 of 14 Phase 3: SRE Observability, Resilience & DR

AWS X-Ray Distributed Tracing

🕑 ~58 min read · 3 services covered
X-Ray Service Map Sampling Rules

Recap: From Outside-In Probes to Inside-Out Traces

Day 30 put monitoring outside the infrastructure. Synthetics canaries are scripted Puppeteer/Selenium probes that drive real user flows against your public endpoints on a schedule, which means they answer a question your internal dashboards structurally cannot: is the thing a customer actually does still working? That is the value of synthetic probes from outside — they catch SLO breach detection at the edge, where DNS, CDN, TLS, and front-end regressions live, and they do it before a human files a ticket. The limitation is equally structural. A canary tells you the checkout flow failed and roughly how long it took; it cannot tell you which of the twelve services behind that flow was responsible.

X-Ray is the same service family — CloudWatch-adjacent observability tooling, same console, same alarms — pointed at the opposite failure mode. Where canaries observe the system from outside and report a symptom, X-Ray instruments the system from inside and reports the causal chain. The two are complementary rather than competing: a canary alarm is often the trigger that sends you into the X-Ray service map to find the hop that actually degraded. Today is about building that inside-out view, and about the sampling economics that decide how much of it you can afford to keep.

Foundations You'll Need Today

Today's material assumes a handful of ideas that are second nature to anyone who has run a production system but are never spelled out in the AWS documentation, because the documentation is written for people who already have them. Here they are, built from the ground up.

A request is a journey, not a single event

When a user clicks a button in a web application, the browser sends one message to a server and waits for a reply. That much is intuitive. What is less intuitive is what happens on the other side. In a modern application, that one incoming message is rarely handled by one program. It is more like a relay race: a front-end program receives it, then calls a second program to look up the user's account, which calls a third program to fetch the user's order history, which calls a database to read the actual rows. Each of those programs is called a service, and each handoff from one service to the next is a hop. The user experiences one request that took 800 milliseconds. The system experiences five or six separate pieces of work, any one of which could be the slow one. This is the entire reason distributed tracing exists: without a way to see the individual hops, you only know the total time, not where it went.

A trace is a receipt for one request

If you wanted to reconstruct that relay race after the fact, you would need each runner to write down when they started, when they finished, and who handed them the baton. A trace is exactly that: a record of one request's journey, assembled from notes that each service wrote down as it did its part. Each service's note is called a segment, and if a service does several distinct things — say, a database read and then an HTTP call to a partner API — it can break its note into smaller pieces called subsegments. The critical detail is that the services have to agree on a shared identifier so their notes can be stitched together into one story. That identifier travels with the request from service to service, which is why the day keeps talking about "propagating the header" — it just means passing that identifier along on every outbound call. If one service forgets to pass it, the story breaks in half and you get two disconnected traces instead of one.

A daemon is a background helper process

A daemon is a small program that runs continuously in the background on a machine, doing a job nobody watches directly. You have almost certainly used one without knowing it — the thing that keeps your laptop's clock accurate is a daemon. In the X-Ray world, the daemon's job is to sit on each server, collect the trace notes that the application produces, hold them briefly in memory, and forward them in batches to the AWS service that stores them. The application does not talk to AWS directly; it drops its notes off with the local daemon and moves on. This matters because of how it communicates: the application sends to the daemon over UDP, a network protocol that fires a message and does not wait for or expect a confirmation. That is fast and it never slows the application down, but it also means that if the daemon is not running, the notes are simply thrown into the void with no error message anywhere. This is why "the traces are missing but the app is fine" is such a common symptom.

An IAM role is a permission slip a service wears

AWS is built on the principle that nothing gets to do anything unless it has been explicitly granted permission. An IAM role is a named bundle of permissions that a program can temporarily adopt, the way a contractor is issued a badge that opens specific doors for the duration of a job. When the day says "the execution role needs the appropriate X-Ray write permissions," it means the program has to be wearing a badge that includes the door labeled "send trace data to X-Ray." If it is not, the program still runs perfectly — it just quietly fails to deliver its trace notes, because tracing is deliberately designed never to break the application it is observing. The result is a service that works fine and is invisible in the tracing console, which is a confusing combination until you know to look at permissions.

Sampling means keeping some and discarding the rest

Recording a detailed receipt for every single request is expensive when you are serving millions of them. Sampling is the practice of recording only a fraction — say, five out of every hundred requests — and assuming the sample is representative of the whole. It is the same logic as a poll: you do not need to ask every voter to know roughly how the electorate feels. The tradeoff is that a rare event can fall entirely outside the sample. If one request in a thousand fails, and you are only recording one in twenty, you may have no record of the failure at all. The decision about whether to record a given request is made once, at the very beginning of the journey, and then carried along with the request — which is why you cannot decide after the fact to keep a trace you already discarded.

With that grounding, here is why X-Ray exists, what problem it actually solves, and where the exam expects you to reach for it.

1. Why X-Ray Is on the Exam

The architectural problem X-Ray exists to solve is the observability gap created by decomposition. When a monolith handles a request, a single stack trace or a single log line with a duration tells you where the time went. Once that monolith is split into a dozen services — an API Gateway fronting Lambda functions that call DynamoDB, publish to SNS, and invoke each other — no single component has a complete picture of the request. Each service logs its own work, each service emits its own metrics, and the aggregate view you get from CloudWatch is a set of independent time series that correlate only loosely. A p99 latency increase on the front-end service and a p99 increase on a downstream service look like two separate incidents until you can prove they are the same requests.

Distributed tracing closes that gap by propagating a correlation identifier across every hop and recording the timing of each hop as a structured segment. The result is a per-request timeline rather than a per-service aggregate, and the aggregate view built from those timelines — the service map — shows you the topology and the latency contribution of each edge. This is why the topic shows up in SAP-C02 Domain 2, the resilient architectures domain, and specifically in the observability and operational-excellence sub-areas. Exam scenarios rarely ask you to configure X-Ray in detail. They ask you to choose the right diagnostic tool for a described symptom, and the discriminator is almost always whether the question describes a request that crosses service boundaries.

The pattern to internalize is this: if the scenario says "a single service is slow," CloudWatch metrics or Logs Insights may be sufficient. If the scenario says "a request traverses multiple services and we cannot tell which one is responsible," or names a specific downstream dependency like a DynamoDB table inside a Lambda invocation, the answer is X-Ray. The exam also tests the cost-control dimension, because tracing every request at scale is expensive, which is why sampling rules are a first-class part of the service rather than an afterthought.

2. How a Trace Is Actually Built

X-Ray's data model has three levels, and understanding them is what lets you reason about the service map instead of just reading it. A trace is the end-to-end record of a single request, identified by a trace ID that is generated at the entry point and propagated to every downstream call. Within a trace, each unit of work is a segment, and segments are emitted by the service that performed the work — one segment per service per request. Segments can contain subsegments, which break a single service's work into finer pieces: an outbound HTTP call, a SQL query, a DynamoDB operation, or a block of custom instrumented code. The service map is built by aggregating segments across many traces and drawing an edge wherever one service's segment references another's.

Propagation is the part that makes or breaks an implementation. X-Ray uses a tracing header — historically X-Amzn-Trace-Id — that carries the trace ID, the parent segment ID, and a sampling decision. When a service receives a request with that header, it joins the existing trace rather than starting a new one. When it makes an outbound call, it must forward the header, which is why the X-Ray SDK matters: it instruments common clients (HTTP libraries, AWS SDK calls, database drivers) so the header is injected automatically. A single uninstrumented hop breaks the chain, and the service map will show the trace terminating at that boundary with no downstream edges. This is the most common reason a service map looks incomplete in a real account.

Instrumentation comes in three flavors, and the exam expects you to know which applies where. The X-Ray SDK is the code-level option for applications you control, available for Java, Node.js, Python, Ruby, Go, and .NET. The X-Ray daemon is a small process that listens on UDP port 2000 and buffers segments before forwarding them to the X-Ray API in batches — it is what EC2 instances, ECS tasks, and self-managed EKS nodes run. On Lambda, the daemon is managed for you: enabling active tracing on the function is sufficient, and the execution role needs the appropriate X-Ray write permissions. AWS-managed services like API Gateway, Elastic Load Balancing, and EventBridge can emit segments natively without any code change, which is often how you get the entry-point segment for free.

3. The Core Decision Boundary: Trace or Don't Trace

The fork that most scenario questions hinge on is not "which tracing tool" but "is tracing the right instrument at all." Tracing is a per-request, high-cardinality signal. It is excellent at answering "which hop in this specific request path is slow or failing" and poor at answering "how many requests did we serve last hour" or "what is our total error budget consumption." Metrics are cheap, aggregated, and always-on; traces are expensive, detailed, and sampled. A well-designed observability stack uses metrics to detect that something is wrong and traces to explain why, which is exactly the division of labor the exam rewards.

The second half of the boundary is where the trace boundary itself should be drawn. X-Ray traces requests, not background processes, and it traces across services you instrument. If a scenario describes a batch job that runs on a schedule with no inbound request, tracing it requires manual segment creation rather than automatic instrumentation. If a scenario describes a third-party SaaS dependency you cannot instrument, the trace will show the outbound call as a subsegment with a duration but no downstream detail — which is still useful, because it proves the time was spent outside your boundary.

Symptom in the scenarioRight instrumentWhy
One service is slow, no cross-service callCloudWatch metrics, Logs InsightsAggregate signal is sufficient; tracing adds cost without new information
Request crosses 3+ services, bottleneck unknownX-Ray service mapPer-hop latency attribution is the only way to isolate the responsible service
Specific downstream call is suspected (DynamoDB, HTTP)X-Ray subsegmentsSubsegments time individual operations inside a service
User flow broken but servers look healthySynthetics canaries (Day 30)Outside-in probe catches edge failures internal metrics miss
Need to search raw log text across accountsLogs Insights, subscription filters (Day 32)Logs are the right substrate for text search, not traces

4. Sampling Rules and Their Tradeoffs

Sampling is the knob that decides how much of your traffic becomes trace data, and it is the single most consequential configuration choice in X-Ray. The default rule is a reservoir of one trace per second plus five percent of additional requests, which is designed to give you a continuous baseline without proportional cost growth. The reservoir guarantees that even a low-traffic service produces some traces every second, which matters because a purely percentage-based sampler would produce nothing at all for a service handling one request per minute. The percentage above the reservoir is what scales with volume, and it is where the cost lives.

Custom sampling rules let you override the default per service, per URL path, per HTTP method, or per resource ARN, and they are evaluated in priority order. This is how you solve the classic tension between cost and coverage: sample aggressively for high-volume, low-value endpoints and sample at a much higher rate for the low-volume, high-value paths where a single failure is expensive. A checkout endpoint handling ten requests per second can be sampled at a high percentage cheaply; a health-check endpoint handling ten thousand requests per second should be sampled at a fraction of a percent or excluded entirely.

The tradeoff to understand is that sampling is a decision made at the entry point and propagated downstream, not re-evaluated at each hop. Once a request is marked as not sampled, no downstream service will record it, which means you cannot retroactively decide to keep a trace after the fact. This has a real operational consequence: if you sample at one percent and an incident affects only a small subset of requests, you may have very few traces of the failing path. The mitigation is to raise the sampling rate temporarily during an investigation, or to use a rule that samples error paths at a higher rate — but the latter requires the sampling decision to be made where the error is known, which is usually not the entry point.

Sampling configurationWhat it costsWhen to use it
Default (1/sec reservoir + 5%)Low, predictableBaseline coverage across all services with no tuning
High-rate custom rule on a specific pathProportional to that path's volumeLow-volume, high-value endpoints (checkout, auth)
Low-rate or excluded high-volume pathMinimalHealth checks, static asset fetches, noisy pollers
100% samplingHighest; can throttle segment deliveryShort-lived debugging windows only, never steady state

5. Sizing, Limits and Quotas

X-Ray's limits fall into three buckets: trace data size, API throughput, and retention. The trace document size limit is 64 KB per segment, and exceeding it causes the segment to be dropped rather than truncated — which is why dumping large payloads into annotations or metadata is a common self-inflicted failure. Annotations are indexed and searchable but limited in number and value size; metadata is not indexed and is intended for larger, non-searchable payloads. If you need to attach a large object to a trace, the correct pattern is to store it elsewhere and put a reference in the trace.

On the API side, the daemon buffers and batches segments before sending them, which smooths bursts, but there are still per-account limits on the rate at which segments can be sent and retrieved. The practical implication is that 100% sampling on a high-throughput service can cause the daemon to drop segments under load, producing a service map that looks like it has gaps. The daemon's own logs are the place to look for this, and the fix is almost always to lower the sampling rate rather than to raise a quota.

Retention is the limit most people are surprised by. Trace data is retained for 30 days, and there is no configuration to extend it. If a scenario requires long-term retention of request-level data for compliance or trend analysis, X-Ray is the wrong tool — the answer is to export to S3 or a data warehouse, or to rely on aggregated metrics for the long view. The 30-day window is generous for incident response and adequate for weekly trend comparison, but it is not an audit log.

LimitValueConsequence of exceeding
Trace document size64 KB per segmentSegment dropped, trace incomplete
Trace data retention30 daysOlder traces unavailable; export required for longer retention
Daemon listenerUDP port 2000Segments never reach the API if the daemon is unreachable
Default sampling1 trace/sec reservoir + 5%Higher rates increase cost and risk of dropped segments

6. Failure Modes and What They Look Like

The most common X-Ray failure is not an outage — it is a silently incomplete service map. The symptom is a map where traces terminate at a service boundary with no downstream edges, or where a service appears in the map but never as a caller of anything. The first diagnostic move is to check whether the tracing header is being propagated across that boundary. If the caller is instrumented but the callee is not, the callee's work is invisible. If the callee is instrumented but the caller is not forwarding the header, the callee starts a new trace instead of joining the existing one, and you get two disconnected traces rather than one connected one.

The second failure mode is missing permissions. X-Ray requires IAM permissions to send segments, and the failure is quiet: the application works, the trace simply never appears. On Lambda, the execution role needs the X-Ray write actions; on EC2 and ECS, the instance role or task role does. The diagnostic move is to check the service's own logs for X-Ray SDK errors, because the SDK logs delivery failures rather than throwing them into the application path. This is deliberate — tracing must never break the application — but it means a misconfigured role can go unnoticed for a long time.

The third failure mode is the daemon not running or not reachable. On EC2 and ECS, the SDK sends segments to the local daemon over UDP, and UDP is fire-and-forget: if nothing is listening on port 2000, the segments are simply lost with no error at the application layer. The symptom is a service that appears in the map intermittently or not at all, correlated with which instances have the daemon running. The diagnostic move is to check the daemon process and its logs on the affected instances, and to confirm the security group or network policy allows the local UDP traffic.

The fourth failure mode is sampling starvation. If the sampling rate is set very low and the traffic is bursty, an incident may occur entirely within unsampled requests, leaving you with no traces of the failure. The symptom is a service map that looks healthy during an incident that is clearly happening. The diagnostic move is to check the current sampling configuration and temporarily raise it, accepting the cost increase for the duration of the investigation.

7. The Operational and SRE Angle

X-Ray's role in an SRE workflow is diagnostic rather than alerting. You do not typically alarm on trace data directly — you alarm on metrics, and you use traces to explain the alarm. The exception is the X-Ray-specific metrics that CloudWatch exposes, such as the count of traces with errors or faults, which can be alarmed on if you want a signal that is closer to the request level than a service-level error rate. In practice, most teams alarm on the service's own error and latency metrics and treat the service map as the first place to look once an alarm fires.

The runbook shape that works is a short decision tree. An alarm fires on elevated latency or error rate for a service. The on-call engineer opens the service map filtered to the alarm window and looks for the edge with the largest latency contribution or the highest error rate. If the responsible edge is an outbound call to a dependency, the next step is to check that dependency's own health. If the responsible edge is internal to the service, the next step is to drill into subsegments to find the specific operation. This is a five-minute investigation rather than a thirty-minute one, and the value of tracing is entirely in that compression.

From an SLO perspective, traces are what let you attribute error budget consumption to a specific dependency rather than to the service that happens to be at the edge. If your checkout SLO is breached because a downstream inventory service is slow, the trace is the evidence that lets you have the right conversation with the right team. This is also why annotations matter operationally: adding a customer tier, a tenant ID, or a feature flag as an annotation lets you filter traces by dimensions that matter to the business, which turns the service map from a debugging tool into a tool for answering questions like "is this affecting enterprise customers specifically."

8. Edge Cases and Exam Gotchas

The first gotcha is the distinction between annotations and metadata, which the exam likes because it is a real design decision. Annotations are indexed and searchable, so they are what you filter traces by; they are limited in number and size. Metadata is not indexed and is intended for larger payloads you want attached to the trace for later inspection. If a question asks how to make traces filterable by a custom attribute, the answer is annotations. If it asks how to attach a large diagnostic payload, the answer is metadata.

The second gotcha is the propagation requirement. X-Ray does not magically connect services; it connects services that forward the tracing header. Any scenario that describes an incomplete service map should be diagnosed as a propagation or instrumentation gap, not as a service outage. This is also why managed services that emit segments natively are valuable — they remove one propagation hop from the equation.

The third gotcha is the 30-day retention limit, which is a hard boundary with no configuration. Any scenario requiring long-term request-level retention needs an export path, not an X-Ray configuration change.

The fourth gotcha is the sampling decision's one-way nature. Because the decision is made at the entry point and propagated, you cannot recover a trace that was not sampled. Scenarios that describe "we need to investigate an incident that happened last week and we have no traces" are usually describing a sampling rate that was too low, and the fix is forward-looking rather than retrospective.

The fifth gotcha is the daemon's UDP transport. Because UDP is fire-and-forget, a missing or unreachable daemon produces silent data loss rather than an error. Any scenario describing intermittent missing traces on EC2 or ECS should prompt a check of the daemon, not the application.

9. X-Ray vs. the Services It Gets Confused With

The services X-Ray is most often confused with are CloudWatch Logs, CloudWatch metrics, and CloudWatch Synthetics, and the confusion is understandable because they all live in the same console and all answer "what is happening in production." The distinction is the shape of the data. Metrics are aggregated numeric time series with no per-request identity. Logs are per-event text records with no cross-service correlation unless you build it. Traces are per-request structured timelines with explicit cross-service correlation. Synthetics are scheduled probes that generate requests from outside. Each answers a different question, and the exam scenarios are usually written so that exactly one of these shapes fits.

A second comparison worth holding is X-Ray against third-party APM tools. The exam is AWS-centric, so the expected answer for a native AWS scenario is X-Ray, but the architectural distinction is real: X-Ray is a tracing service, not a full APM platform, and it does not do continuous profiling or code-level flame graphs. If a scenario requires that depth, the answer is likely a partner solution rather than X-Ray, though such scenarios are rare on SAP-C02.

ServiceData shapePick it when…
X-RayPer-request trace with cross-service segmentsThe question is which hop in a multi-service request is responsible
CloudWatch metricsAggregated numeric time seriesYou need cheap, always-on detection and alarming
CloudWatch Logs / Logs InsightsPer-event text, queryableYou need to search log content or build ad-hoc queries
CloudWatch SyntheticsScheduled outside-in probesYou need to verify user flows work from the customer's perspective
VPC Flow LogsNetwork-level connection recordsThe question is about IP-level connectivity, not application latency

Hands-on Lab: Instrumenting a Lambda-to-DynamoDB Chain

The goal of this lab is to build a working trace across a real service boundary and then use the service map to attribute latency to a specific hop. You will need an AWS account with permission to create Lambda functions, DynamoDB tables, and IAM roles. Budget roughly 45 minutes, most of which is waiting for deployments.

Step 1 — Create the downstream dependency. Create a DynamoDB table named xray-lab-items with a partition key of pk (string). Populate it with a few hundred items so that a scan or query takes measurable time. A table with ten items will complete too fast to produce a useful latency signal, which is the most common reason this lab appears not to work.

Step 2 — Create the producer function. Create a Lambda function (Node.js or Python) that reads from the table and returns the result. Attach an execution role that includes the X-Ray write permissions in addition to the DynamoDB read permissions. Enable active tracing on the function in the Lambda console. This is the step that removes the need to run the daemon yourself — Lambda manages it.

Step 3 — Create the caller function. Create a second Lambda function that invokes the first one via the AWS SDK. Enable active tracing on this function as well. The AWS SDK is instrumented by the X-Ray SDK, so the invoke call will automatically propagate the tracing header and appear as a downstream edge in the service map. If you use a raw HTTP client instead of the SDK, you will need to inject the header manually — which is a useful exercise in understanding why propagation matters, but not necessary for the lab.

Step 4 — Generate traffic. Invoke the caller function a few dozen times, either from the console or with a small script. Because the default sampling rule includes a one-per-second reservoir, even a modest number of invocations will produce traces. Wait a minute or two for segments to be delivered and indexed.

Step 5 — Read the service map. Open the X-Ray console and view the service map for the last fifteen minutes. You should see the caller function, the callee function, and the DynamoDB table as three nodes with edges between them. Click the edge between the callee and DynamoDB and look at the latency distribution. This is the hop that the lab is designed to isolate.

Step 6 — Attribute the latency. Open a specific trace and examine the segment timeline. Identify how much of the total request duration was spent in the DynamoDB subsegment versus the Lambda execution overhead. Then deliberately slow the DynamoDB call — for example, by scanning a larger table or adding a small artificial delay — and repeat. The service map should show the callee-to-DynamoDB edge as the dominant contributor to p99 latency, which is exactly the conclusion the original lab prompt asks you to reach.

Step 7 — Break it on purpose. Remove the X-Ray write permissions from the callee's execution role and invoke the caller again. The application will continue to work, but the callee will disappear from the service map. This demonstrates the silent-failure property described in section 6 and is worth doing once so you recognize the symptom in a real account.

Scenario Question Drills (20 min)

Q1. A microservices application has growing p99 latency but it is unclear which of twelve services is the bottleneck. What AWS service pinpoints this?

A. CloudWatch Logs Insights alone
B. AWS X-Ray, using the service map to isolate the slow hop
C. VPC Flow Logs
D. AWS Config
Correct answer: B. X-Ray's distributed tracing and service map visualize per-hop latency across the full call chain, directly identifying the bottleneck service that aggregate metrics cannot isolate.

Q2. A service map shows traces terminating at an API Gateway boundary with no downstream edges, even though the backend services are healthy and serving traffic. What is the most likely cause?

A. The backend services are not running
B. The tracing header is not being propagated across that boundary, or the downstream services are not instrumented
C. X-Ray retention has expired
D. The service map requires a paid tier to show downstream edges
Correct answer: B. X-Ray connects services only when the tracing header is forwarded and the downstream service is instrumented. A map that stops at a boundary is a propagation or instrumentation gap, not an outage.

Q3. A team wants to filter traces by customer tier so they can investigate whether an incident affects enterprise customers specifically. What X-Ray feature should they use?

A. Metadata, because it holds arbitrary payloads
B. Annotations, because they are indexed and searchable
C. Subsegments, because they break work into finer pieces
D. Sampling rules, because they control which traces are kept
Correct answer: B. Annotations are indexed and searchable, which is what makes them the right choice for filtering traces by custom dimensions. Metadata is not indexed and is intended for larger, non-searchable payloads.

Q4. A compliance requirement mandates that request-level trace data be retained for two years. What is the correct approach?

A. Configure X-Ray retention to two years in the console
B. Export trace data to S3 or a data warehouse, because X-Ray retains traces for 30 days with no configuration to extend it
C. Enable X-Ray long-term storage mode
D. Use CloudWatch Logs instead, which retains data indefinitely by default
Correct answer: B. X-Ray trace retention is fixed at 30 days. Any longer retention requirement needs an export path to durable storage.

Q5. A high-throughput service is traced at 100% sampling and the service map shows intermittent gaps during peak traffic. What is the most likely cause and the correct fix?

A. The service is failing; investigate the application logs
B. The daemon is dropping segments under load; lower the sampling rate
C. X-Ray does not support high-throughput services
D. The trace document size limit is being exceeded
Correct answer: B. 100% sampling on a high-volume service can overwhelm the daemon's buffering and delivery capacity, producing gaps. The fix is to lower the sampling rate rather than to raise a quota.

Q6. An application on EC2 emits traces intermittently, and the pattern correlates with which instances are serving the request. The application itself is healthy. What should you check first?

A. The DynamoDB table's provisioned capacity
B. Whether the X-Ray daemon is running and reachable on UDP port 2000 on the affected instances
C. The CloudWatch Logs retention setting
D. The VPC route tables
Correct answer: B. On EC2 the SDK sends segments to the local daemon over UDP, which is fire-and-forget. A missing or unreachable daemon produces silent data loss with no application error.

Q7. A Lambda function's traces never appear in the X-Ray console, but the function works correctly and its logs show successful invocations. What is the most likely cause?

A. Lambda does not support X-Ray
B. Active tracing is not enabled on the function, or the execution role lacks X-Ray write permissions
C. The function's timeout is too short
D. The function needs a NAT gateway to reach X-Ray
Correct answer: B. Lambda manages the daemon for you, but tracing must be enabled on the function and the execution role must permit sending segments. Delivery failures are logged rather than thrown, so the application appears healthy.

Q8. A team needs to attach a large JSON diagnostic payload to a trace for later inspection, but does not need to search by its contents. What should they use?

A. Annotations
B. Metadata
C. A separate subsegment per field
D. A sampling rule
Correct answer: B. Metadata is not indexed and is intended for larger payloads attached to a trace. Annotations are indexed and searchable but limited in number and size.

Q9. An incident occurred last week and the team has no traces of the failing requests, even though tracing has been enabled for months. What is the most likely explanation?

A. X-Ray was down during the incident
B. The sampling rate was too low, and the failing requests were not sampled
C. The traces were deleted by a lifecycle policy
D. X-Ray only retains traces for 24 hours
Correct answer: B. The sampling decision is made at the entry point and propagated, so unsampled requests leave no trace. A low sampling rate can miss a failure that affects only a subset of traffic.

Q10. A team wants to reduce X-Ray cost without losing visibility into their low-volume checkout endpoint. What is the correct configuration?

A. Disable tracing entirely and rely on metrics
B. Use a custom sampling rule with a high rate for the checkout path and a low rate or exclusion for high-volume, low-value paths
C. Set the global sampling rate to 100% and rely on retention to control cost
D. Move tracing to a different region
Correct answer: B. Custom sampling rules are evaluated per path, which lets you spend the trace budget where it matters and exclude noisy, low-value traffic.

Q11. A trace segment is being dropped, and the application logs show no errors. The segment includes a large serialized request body. What is the most likely cause?

A. The segment exceeded the 64 KB trace document size limit
B. The daemon is not running
C. The IAM role lacks permissions
D. The trace ID was not propagated
Correct answer: A. Segments larger than 64 KB are dropped rather than truncated. Large payloads should be stored elsewhere with a reference placed in the trace.

Q12. An on-call engineer receives an alarm for elevated error rate on the front-end service. What is the most efficient first step using X-Ray?

A. Re-read the application source code
B. Open the service map filtered to the alarm window and identify the edge with the highest error rate or latency contribution
C. Restart all services
D. Increase the sampling rate to 100% permanently
Correct answer: B. The service map is the fastest way to attribute an alarm to a specific hop. Drilling into subsegments comes next if the responsible edge is internal to a service.

Q13. A scenario describes a scheduled batch job with no inbound request that the team wants to trace. What is the correct approach?

A. X-Ray cannot trace batch jobs under any circumstances
B. Create a segment manually at the start of the job, since there is no inbound request to generate one automatically
C. Convert the batch job to a Lambda function
D. Use VPC Flow Logs instead
Correct answer: B. X-Ray traces requests, so a job with no inbound request needs a manually created segment to establish the trace root.

Q14. A team wants to know whether an incident is affecting enterprise customers specifically, and their traces include a customer tier annotation. What does this enable?

A. Nothing; annotations are not searchable
B. Filtering traces by the annotation value, which turns the service map into a tool for answering business-relevant questions
C. Automatic alerting on enterprise customer errors
D. Longer trace retention for those requests
Correct answer: B. Indexed annotations make traces filterable by custom dimensions, which is what allows per-segment analysis of an incident.

Q15. A scenario requires code-level flame graphs and continuous profiling of a Java application, beyond per-request tracing. What is the correct conclusion?

A. X-Ray provides flame graphs natively
B. X-Ray is a tracing service, not a full APM platform; continuous profiling requires a different tool
C. Enable X-Ray Insights to get flame graphs
D. Use CloudWatch Synthetics instead
Correct answer: B. X-Ray traces requests across services; it does not perform continuous profiling or produce code-level flame graphs. Those requirements point to a different class of tool.

Peek into Tomorrow

X-Ray answers the question "which hop in this request was slow," but it deliberately does not answer the question "what did the application actually say when it failed." Traces carry timing and structure; they do not carry the stack trace, the validation error message, or the business-logic branch that was taken. When an incident is caused by a bad input, a misconfigured feature flag, or an exception that was caught and swallowed, the trace will show you the hop and the duration but leave you guessing at the cause. That gap is where logs live, and it is a gap that becomes acute at organizational scale, because the logs you need are often in an account you do not have open.

Tomorrow's material addresses both halves of that problem. The first half is querying: how to ask structured questions of log data without standing up a separate analytics stack, which is what the Logs Insights query language provides. The second half is aggregation: how to get logs from every account in an organization into one place in near real time, which is what subscription filters streaming to a centralized logging account via Kinesis Data Firehose or Streams accomplish. The open question worth carrying forward is how you would correlate a trace you have already identified with the log lines from the specific invocation it represents — and what has to be true about your logging setup for that correlation to be possible at all.

Sources