Amazon CloudWatch Logs Insights & Centralized Logging
Recap: From Traces to Logs
X-Ray gave us the shape of a request: a service map that shows which hop is slow, sampling rules that keep trace volume affordable, and annotations and subsegments that let us drill into a specific slow DynamoDB query inside a Lambda invocation. That is a powerful lens, but it is a sampled, structural lens. It tells you that the DynamoDB call inside the checkout function accounts for 340ms of a 900ms p99, and it tells you that the error rate on that hop doubled at 14:05. What it does not tell you is why. The trace carries no stack trace, no exception message, no request payload, no downstream error code. For that you need the log line the application wrote immediately before it gave up.
Today extends the observability signal thread by moving from per-hop latency hotspots to the raw event stream underneath them. Logs Insights is the query layer over CloudWatch Logs, and subscription filters are the transport that moves those events out of the account that produced them and into a place where the whole organization can search them. The two halves solve different problems — one is ad-hoc analysis in the account you are already in, the other is durable, org-wide aggregation — and the exam tests whether you can tell which one a scenario is asking for.
Foundations You'll Need Today
Today's material sits on top of a handful of AWS building blocks that the rest of this curriculum assumes you already know. If you have been working at the Cloud Practitioner level — comfortable with the idea of cloud services but without hands-on experience wiring them together — the concepts below are the ones this specific day leans on hardest. None of them are complicated on their own; the difficulty is that the day's prose uses them as shorthand, and shorthand only works if the underlying idea is already in your head.
CloudWatch Logs: What It Is and Why It Exists
When an application runs, it produces a stream of text messages about what it is doing: a request came in, a database call took 340 milliseconds, an error occurred, a background job finished. Those messages are called logs, and they are the most detailed record of what a system actually did. The problem is that in a cloud environment there is no single machine to log into and no single file to open — a workload might run across dozens of containers that start and stop on their own, and the moment a container is replaced, anything it wrote to its local disk is gone. CloudWatch Logs is AWS's answer to that: a managed service that collects log messages from wherever they are produced, stores them durably, and makes them searchable later. You do not run a server for it and you do not manage storage for it; you send it messages and it keeps them. That is the whole value proposition, and it is why almost every AWS service can be configured to write its logs there.
Log Groups, Log Streams, and Log Events
CloudWatch Logs organizes what it stores in a three-level hierarchy, and today's material uses all three terms without stopping to define them. A log event is the smallest unit: one timestamped message, one line of text. A log stream is an ordered sequence of events from a single source — in practice, one stream per running process, per container, or per Lambda function instance. A log group is the container that holds streams, and it is the level you actually configure things on: retention, permissions, and the subscription filters discussed today are all set on a log group, not on an individual stream. The mental model that helps is a filing cabinet: the log group is the drawer, each log stream is a folder inside it, and each log event is a single page. When you hear "attach a filter to the log group," picture putting a rule on the whole drawer rather than on one folder.
Subscription Filters: Moving Logs Somewhere Else
By default, logs live in the account and region where they were produced, and only people with permission in that account can read them. A subscription filter is a rule you attach to a log group that says "every log event matching this pattern should also be sent to this destination." It is a one-way pipe: the original events stay where they are, and a copy flows out to wherever you pointed it. This is the mechanism that makes centralized logging possible, because it is how logs escape the account that created them. The important thing to hold onto is that a subscription filter is a pipeline, not a search tool — it does not answer questions, it just moves data. Today's material contrasts it constantly with Logs Insights, which is the search tool, and the two are easy to conflate because they both attach to a log group and both take a filter pattern.
Kinesis Data Firehose and Kinesis Data Streams
When a subscription filter forwards logs, it has to forward them to something, and today's material names two AWS services as the usual destinations. Kinesis Data Firehose is a delivery service: you tell it where you want data to end up — an S3 bucket, an OpenSearch cluster, a third-party analytics platform — and it buffers incoming records, optionally transforms them, and writes them out reliably, retrying if the destination is temporarily unavailable. It is a pipe with a destination attached, and it has no memory of what it delivered. Kinesis Data Streams is a different shape: it is a durable, replayable buffer that holds records for a configurable window (24 hours by default, up to a year) and lets multiple independent consumers read the same records at their own pace. The distinction that matters today is replay: if you need to re-read events from yesterday, Firehose cannot help you because it already delivered them and forgot, while Data Streams still has them. The capacity unit for Data Streams is a shard, which is simply a slice of throughput — more shards means more records per second, and running out of shard capacity is what produces the throttling failure mode discussed later in this day.
AWS Accounts, Organizations, and SCPs
Today's material talks constantly about "40 accounts," "a central logging account," and "an SCP," and those phrases assume you know that an AWS account is not just a billing construct but a hard security and isolation boundary. Every resource lives inside exactly one account, permissions are granted within an account, and by default one account cannot see another's resources at all. Large organizations deliberately spread workloads across many accounts so that a mistake or compromise in one does not touch the others. AWS Organizations is the service that groups those accounts under one management umbrella, and a Service Control Policy (SCP) is a guardrail you attach at the organization level that sets the maximum permissions any account underneath is allowed to use — it cannot grant access, only restrict it. When today's material says an SCP prevents a workload account from deleting its own subscription filter, it means a rule set at the organization level that overrides whatever the account's own administrators might try to do. That is the mechanism that makes a centralized logging design tamper-resistant rather than merely convenient.
With that grounding — a service that stores logs, a hierarchy of groups and streams, a pipe that moves them between accounts, two services that can receive them, and an organizational structure that makes the whole thing enforceable — here is why CloudWatch Logs Insights and centralized logging exist and what problem they actually solve.
1. Why This Is on the Exam
Centralized logging sits at the intersection of two SAP-C02 domains that are otherwise easy to study separately: Domain 1 (Design Solutions for Organizational Complexity) and Domain 2 (Design for New Solutions, specifically the observability and operational-excellence sub-areas). The organizational-complexity angle is the one that catches people. A multi-account landing zone built with Control Tower produces dozens or hundreds of accounts, each with its own CloudWatch Logs log groups, each with its own retention setting, each with its own IAM boundary around who can read them. An incident that spans three accounts — an API Gateway in the workload account, a Lambda in a shared-services account, and an RDS instance in a data account — produces evidence in three separate places that no single engineer has permission to read. The exam wants to know whether you recognize that as an architecture problem rather than an IAM problem.
The second reason this topic appears is that it is a natural place to test whether you understand the difference between a query tool and a pipeline. Logs Insights is a query tool: it runs against log groups that already exist, in the account and region where they exist, and it returns results to whoever is looking at the console. A subscription filter is a pipeline: it is a persistent, always-on configuration attached to a log group that forwards every matching event to a destination. Candidates who have only used Logs Insights in a single account tend to answer "use Logs Insights" for questions that are actually about cross-account aggregation, and that is precisely the distractor the exam is built around.
There is also a compliance dimension that shows up in Domain 4 (Continuous Improvement for Existing Solutions) and in the security-adjacent scenarios. Regulated workloads frequently require that logs be retained for a fixed period, that they be immutable once written, and that the team who operates the workload cannot delete the evidence of what they did. A centralized logging account owned by a security team, with the workload accounts holding only write permission to it, is the standard answer. That pattern is built out of the same two primitives — log groups and subscription filters — plus an IAM and Organizations structure around them.
Finally, this is one of the few observability topics where cost is a first-class design constraint rather than an afterthought. CloudWatch Logs charges for ingestion, for storage, and for the Insights queries you run against stored data. A naive "send everything everywhere" design can cost more than the compute it is monitoring. The exam does not usually ask you to compute a bill, but it does ask you to choose between designs with obviously different cost profiles, and knowing which knob controls which cost is what makes that choice defensible.
2. How CloudWatch Logs Actually Works
Everything in this day rests on one data model, so it is worth being precise about it. A log group is a named container that holds log streams. A log stream is an ordered sequence of log events from a single source — in practice, one stream per process, per container, per Lambda execution environment, or per file being tailed. A log event is a timestamp plus a UTF-8 payload, and that is the entire schema. CloudWatch Logs does not parse your payload, does not index individual fields, and does not know that a line beginning with {"level":"ERROR" is different from a line beginning with {"level":"INFO" until you tell it so. This is the single most important fact for understanding why Logs Insights queries look the way they do.
When an agent or SDK writes to a log group, the events are durably stored and become queryable. The CloudWatch Logs agent (or the unified CloudWatch agent, or the Lambda service itself, or the ECS/EKS log drivers) batches events and calls the PutLogEvents API. There is a per-stream ordering guarantee and a per-event timestamp, but there is no cross-stream ordering guarantee at all. Two containers writing to the same log group will interleave in whatever order their batches arrived, which is why a Logs Insights query that sorts by @timestamp is reconstructing an approximation of the timeline rather than reading a canonical one. For most debugging this is fine; for reconstructing a distributed transaction across services, it is exactly why you also want X-Ray.
Logs Insights is a query engine that runs over one or more log groups within a single region and a single account. You give it a time range and a query written in a purpose-built language, and it scans the matching events, applies your filter and parse stages, and returns aggregated or projected results. The language is pipeline-shaped: a fields clause selects and computes, filter narrows, parse extracts fields out of unstructured text using glob patterns, stats aggregates with functions like count, avg, pct, and sum, and sort/limit shape the output. Because the underlying data is unindexed text, a query that filters on a parsed field still has to scan the events in the time range to extract that field — which is the direct link between query design and query cost.
Subscription filters are a completely separate mechanism that happens to live on the same resource. A subscription filter is attached to a log group and declares a filter pattern plus a destination. Every log event that matches the pattern is delivered, in near-real time, to that destination. The destination is one of a small set: a Kinesis Data Streams stream, a Kinesis Data Firehose delivery stream, or a Lambda function. The filter pattern language is not the Insights query language — it is a simpler matching syntax with terms, quoted phrases, and ? wildcards, and it is evaluated on the raw event text. This distinction matters because a pattern that works in a subscription filter will not paste into an Insights query and vice versa.
The delivery path is asynchronous and best-effort in the sense that it is not transactional with the log write. CloudWatch Logs buffers matching events and forwards them in batches; if the destination is unavailable or throttled, events can be retried and, in pathological cases, dropped. This is why the centralized logging pattern is usually described as near-real-time rather than real-time, and why a design that treats the central account as the system of record for compliance needs to think about the failure window rather than assuming perfect delivery.
3. The Core Decision Boundary: Query In Place vs. Aggregate Centrally
Almost every scenario question on this topic reduces to one fork. Either the logs you need are already in the account and region where you are working, in which case Logs Insights is the right tool and the design question is how to write an efficient query; or the logs you need are spread across accounts, or need to outlive the account that produced them, or need to be readable by a team that does not have access to the producing account, in which case you need a subscription filter and a destination in a central account. The two are not alternatives to each other — a mature design uses both — but a scenario will usually be testing whether you recognize which one the stated constraint demands.
The tell is almost always in the constraint language. "An engineer needs to find the slowest requests in the last hour" is a query problem. "The security team needs to search application logs from all 40 accounts from one place" is an aggregation problem. "Logs must be retained for seven years even if the workload account is deleted" is an aggregation problem with a retention twist. "We need to alert when a specific error string appears" is neither, strictly — that is a metric filter, which converts matching log events into a CloudWatch metric you can alarm on, and it is worth knowing as a third option because it is a common distractor.
| Requirement | Right primitive | Why the others fail |
|---|---|---|
| Ad-hoc search of logs in the account you are in | Logs Insights query | Subscription filters move data out; they do not answer questions |
| Org-wide search across many accounts | Subscription filter to a central destination | Insights cannot query across accounts or regions |
| Alert when an error pattern appears | Metric filter + CloudWatch alarm | Insights is pull-based; it does not evaluate continuously |
| Retain logs beyond the producing account's lifetime | Subscription filter to S3 via Firehose | Log group retention is bounded and dies with the account |
| Feed a SIEM or third-party analytics platform | Subscription filter to Firehose or Kinesis Data Streams | Insights has no export API for streaming consumers |
| Correlate a request across services | X-Ray (with logs as supporting detail) | Logs have no cross-stream ordering guarantee |
One nuance worth internalizing: Logs Insights can query multiple log groups in a single query, and it can query log groups in a different account if you have been granted cross-account access to them. That is a real capability and it is occasionally the right answer for a small number of accounts. But it does not scale to an organization, it does not survive account deletion, and it requires the querying principal to hold permissions in every producing account — which is exactly the access sprawl a centralized logging account exists to prevent. When a scenario says "40 accounts" or "the security team should not have write access to workload accounts," cross-account Insights queries are the wrong answer even though they are technically possible.
4. Configuration Modes and Their Tradeoffs
The first knob is retention, set per log group. The default is "Never expire," which sounds generous and is actually the most expensive setting in the service, because you pay for stored bytes indefinitely. The practical pattern is to set a short retention on high-volume, low-value log groups (application debug logs, load balancer access logs) and a long or infinite retention only on the groups that carry compliance weight. Retention is a per-log-group setting, so a landing-zone design usually enforces it through a Config rule or a Control Tower guardrail rather than trusting each team to pick sensibly.
The second knob is the subscription filter's filter pattern, and this is where cost control actually happens. A subscription filter with an empty pattern matches every event and forwards all of it, which means you pay CloudWatch Logs ingestion, plus the destination's ingestion, plus storage at the destination, for every debug line your application emits. A pattern that matches only ERROR and WARN reduces that volume by whatever fraction of your logs are noise — often the large majority. The tradeoff is that you cannot retroactively decide you wanted the INFO lines; if they were filtered out at the subscription, they exist only in the source account's log group, subject to that group's retention. Teams that get this wrong usually discover it during an incident, when the detail they need was filtered away three weeks earlier.
The third knob is the destination type, and the choice between Kinesis Data Streams, Kinesis Data Firehose, and Lambda is a genuine architectural fork. Firehose is the default for archival and SIEM ingestion: it buffers, optionally transforms with a Lambda processor, and delivers to S3, OpenSearch, Splunk, Datadog, or a handful of other destinations, with built-in retry and a configurable buffer interval. Kinesis Data Streams is the right choice when you need multiple independent consumers reading the same stream, or when you need to replay a window of events — Firehose is a delivery pipe, not a replayable log. Lambda as a destination is for custom routing or enrichment logic, and it is the most operationally expensive option because you now own the function's concurrency, error handling, and dead-letter behavior.
| Destination | Best for | What you give up |
|---|---|---|
| Kinesis Data Firehose | S3 archival, SIEM delivery, managed retry | No replay; single logical consumer path |
| Kinesis Data Streams | Multiple consumers, replay, custom processing | You manage shards, retention, and consumers |
| Lambda | Custom routing, enrichment, conditional fan-out | You own concurrency, errors, and DLQ design |
The fourth knob is the cross-account permission model, and it is the one most often gotten wrong in practice. The producing account's log group needs an IAM role or resource policy that allows logs:PutSubscriptionFilter against the destination, and the destination account needs a resource policy that accepts the incoming stream. In an Organizations setup this is usually expressed as a destination in the central account with a policy that allows any account in the organization to write to it, plus an SCP that prevents workload accounts from deleting their own subscription filters. Without that SCP, the pattern has a hole: a compromised workload account can simply remove the filter and stop shipping evidence.
5. Sizing, Limits, and Quotas
The numbers that matter here fall into three groups: limits on the log data itself, limits on queries, and limits on delivery. On the data side, a single log event payload is capped at 256 KB, and a PutLogEvents batch is capped at 1 MB or 10,000 events, whichever comes first. Events larger than the payload cap are truncated, which is a real failure mode for applications that log large JSON documents — the truncation is silent from the application's perspective, and the missing tail is only discovered when someone tries to parse the log line and finds it malformed. The practical mitigation is to log a reference (a request ID, an S3 key) rather than the full payload.
On the query side, Logs Insights has a default limit of 50 concurrent queries per account per region, and a single query can scan up to 50 log groups. Query results are capped at 10,000 rows returned to the console, though the stats aggregation runs over the full matched set before that cap applies — which is why an aggregation query is often the right way to answer "how many" questions even when the raw result set is huge. There is also a timeout on query execution; a query that scans an enormous time range across many high-volume log groups can time out rather than return, and the fix is to narrow the time range or the log group set rather than to retry.
On the delivery side, the numbers to remember are the Firehose buffer settings and the subscription filter's own throughput behavior. Firehose buffers by size (default 5 MB) and by interval (default 300 seconds), whichever is reached first, so a low-volume log group can take up to five minutes to appear at the destination. Lowering the buffer interval to 60 seconds reduces that latency at the cost of more, smaller S3 objects or more frequent downstream writes. This is the number that determines whether your "near-real-time" claim is honest for a given log group, and it is a common source of confusion when someone expects a log line to appear in the central account within seconds.
| Limit | Value | Practical consequence |
|---|---|---|
| Max log event size | 256 KB | Larger events are truncated silently |
| Max PutLogEvents batch | 1 MB / 10,000 events | Agents batch automatically; custom writers must chunk |
| Concurrent Insights queries | 50 per account per region | Shared dashboards can contend with ad-hoc queries |
| Log groups per query | 50 | Org-wide search needs aggregation, not a wide query |
| Insights result rows returned | 10,000 | Aggregate with stats rather than returning raw rows |
| Firehose buffer interval | 60-900 s (default 300) | Sets the floor on end-to-end delivery latency |
One more number worth carrying: log group retention is configurable from one day up to ten years, or "never expire." There is no retention setting that is both infinite and free, and there is no way to change retention retroactively for events already deleted. If a compliance requirement says seven years, the retention setting has to be set to seven years before the events age out, not after someone notices.
6. Failure Modes and What They Look Like in Production
The most common failure is silent: logs stop arriving in the central account and nobody notices, because the workload is healthy and the source log group still has everything. The symptom is a gap in the central destination that only becomes visible during an incident, when someone goes looking for evidence that should be there. The first diagnostic move is to check the subscription filter still exists on the source log group — a filter deleted by a well-meaning cleanup script, or never created because a new log group was added after the Terraform module was written, produces exactly this symptom. The second move is to check the destination's resource policy, because a policy that was correct when written can break when the producing account is moved to a different OU or when an SCP is tightened.
The second failure mode is throttling at the destination. Kinesis Data Streams has per-shard write limits, and a log group that suddenly produces a burst — a retry storm, a debug flag left on in production, a runaway loop — can exceed the stream's capacity. The symptom is ThrottlingException in the CloudWatch Logs service metrics and a growing gap between events written and events delivered. Firehose is more forgiving because it buffers and retries, but it has its own limits and a sustained overload will eventually surface as delivery failures to S3. The diagnostic is to compare the source log group's incoming bytes metric against the destination's incoming records metric; a divergence that grows over time is the signature.
The third failure mode is the query that never returns. A Logs Insights query over a wide time range against a high-volume log group can run long enough to time out, and the console gives you a generic failure rather than a partial result. This is not a bug; it is the direct consequence of scanning unindexed text. The fix is to narrow the time range first, then the log group set, then to push filtering into the query as early as possible so fewer events reach the expensive parse and stats stages. A query that filters on a parsed field is doing the parse work on every event in range before it can filter, which is the opposite of what you want.
The fourth failure mode is the one that hurts most in a regulated environment: the log group's retention expires before the compliance window does. This happens when retention is left at the default and someone later discovers the requirement, or when a log group is created by an application's infrastructure-as-code template that hardcodes a 30-day retention while the compliance requirement is seven years. The symptom is not an error — it is the quiet absence of data from a period that is now unrecoverable. The only real mitigation is enforcement at creation time, via a Config rule or a Control Tower guardrail that rejects log groups without an approved retention setting.
7. The Operational and SRE Angle
From an SRE perspective, the logging pipeline is itself a service with an SLO, and it is usually an unmonitored one. The two metrics that matter are delivery latency (how long between an event being written and it being queryable at the destination) and delivery completeness (what fraction of written events arrive). Neither is exposed as a first-class CloudWatch metric, so both have to be constructed. Delivery latency can be approximated by comparing the @timestamp of the newest event in the destination against wall-clock time; completeness is harder and is usually approximated by comparing source incoming-bytes against destination incoming-records on a rolling window, accepting that the two are not directly convertible.
The alarms worth having are unglamorous. An alarm on the source log group's IncomingBytes dropping to zero for a sustained period catches the "application stopped logging" case, which is often the first symptom of a deployment that broke the logging configuration. An alarm on the destination's delivery failures catches the throttling case. An alarm on the subscription filter's existence is not directly expressible in CloudWatch, which is why the standard approach is a Config rule that asserts the filter is present on every log group matching a naming convention — a configuration-drift check rather than a metric alarm.
The runbook shape follows from the failure modes above. Step one is always "confirm the source is producing," because a surprising fraction of "logs are missing" incidents are actually "the application stopped writing logs." Step two is "confirm the filter exists and matches," which means checking both the filter's presence and its pattern against a known-good sample event. Step three is "confirm the destination is accepting," which means checking the destination's resource policy and its own throttling metrics. Only after those three are exhausted is it worth looking at the query layer, because a query problem and a delivery problem look identical from the console — both produce "I searched and found nothing."
There is also a cost-observability angle that belongs in the SRE rotation. CloudWatch Logs spend is dominated by ingestion volume, and ingestion volume is dominated by whatever the noisiest application is doing. A monthly review of per-log-group IncomingBytes against the value that log group actually provided during incidents is a cheap way to find the debug flag that has been on since a release three months ago. This is the same discipline as right-sizing compute, applied to telemetry: the goal is not to log less, it is to stop paying for logs nobody reads.
8. Edge Cases and Exam Gotchas
The first gotcha is the filter pattern language. It is not regex, it is not the Insights query language, and it does not support the same operators. A pattern of ERROR matches any event containing that term as a word; a pattern of ?"ERROR" ?"WARN" matches events containing either; a pattern of [timestamp, request_id, level = "ERROR", ...] matches space-delimited events whose third field is exactly ERROR. Candidates who assume regex semantics write patterns that silently match nothing, and a filter that matches nothing looks exactly like a filter that is working correctly on a quiet log group.
The second gotcha is that subscription filters are per-log-group, not per-account. There is no account-level "send all logs to the central account" switch. Every log group that needs to ship has to have its own filter, which means the pattern has to be applied at creation time — usually by the same infrastructure-as-code module that creates the log group, or by a Config-driven remediation. This is the single most common reason a centralized logging design works for the accounts that were onboarded carefully and silently fails for the ones that were not.
The third gotcha is the distinction between a metric filter and a subscription filter. Both attach to a log group, both take a filter pattern, and both are commonly confused. A metric filter converts matching events into a CloudWatch metric and does not move the log data anywhere; a subscription filter moves the matching events to a destination and does not create a metric. If a scenario asks for an alarm on an error pattern, the answer is a metric filter. If it asks for the events to be searchable elsewhere, the answer is a subscription filter. Some designs need both, attached to the same log group with the same pattern.
The fourth gotcha is region. Logs Insights queries run in one region. A subscription filter delivers to a destination in the same region as the log group, unless you are using a cross-region destination, which is supported for some destination types but adds latency and a second failure domain. A multi-region workload therefore has a multi-region logging problem, and the central logging account needs a destination per region or a cross-region delivery design. Scenarios that mention "all regions" are testing whether you noticed this.
The fifth gotcha is the retention default. "Never expire" is the default, and it is the wrong answer for almost every high-volume log group. Conversely, setting a short retention on a log group that carries compliance weight is a data-loss bug that will not surface until an audit. The exam tends to present this as a cost question when it is really a governance question, and the correct answer usually involves enforcing retention through policy rather than choosing a number per log group.
9. This vs. the Services It Gets Confused With
The nearest neighbor is CloudTrail, and the confusion is understandable because both are "logs" and both are commonly centralized. CloudTrail records API activity — who called what, when, from where, and whether it succeeded. CloudWatch Logs records whatever your application chose to write. They answer different questions, they have different retention and integrity models, and a complete audit story needs both. CloudTrail's own organization trail is the mechanism for centralizing API activity across accounts, and it is a separate design from the CloudWatch Logs subscription filter pattern described here. A scenario that says "we need to know who deleted the bucket" is a CloudTrail question; "we need to know why the deletion request failed" is a CloudWatch Logs question.
The second neighbor is X-Ray, which we covered yesterday. X-Ray is sampled and structural; CloudWatch Logs is complete and unstructured. X-Ray tells you which hop is slow; logs tell you what the code was doing when it was slow. The two are complementary, and the standard pattern is to correlate them with a shared trace ID that the application writes into its log lines — which is exactly the kind of thing that makes a Logs Insights query useful during an incident, because you can pull every log line for a single failing request across every service that handled it.
The third neighbor is OpenSearch Service, which is where a lot of centralized logging designs eventually land. The relationship is that CloudWatch Logs is the collection and transport layer, and OpenSearch is the search and analytics layer. Firehose delivers from one to the other. The decision boundary is scale and query sophistication: Logs Insights is sufficient for ad-hoc investigation and simple aggregation, while OpenSearch is the right answer when you need full-text search across very large volumes, complex dashboards, or long-term retention with fast query. Choosing OpenSearch when Logs Insights would do adds real operational cost, and choosing Logs Insights when the requirement is a searchable multi-year corpus is a design that will not scale.
| Need | Pick | Because |
|---|---|---|
| Who did what in the AWS API | CloudTrail (org trail) | Logs only contain what the app wrote |
| Which hop in a request is slow | X-Ray | Logs have no cross-stream ordering |
| Ad-hoc search of recent app logs | Logs Insights | No pipeline to build or pay for |
| Org-wide searchable log corpus | Subscription filter + Firehose + OpenSearch/S3 | Insights cannot span accounts at scale |
| Alarm on an error string | Metric filter + alarm | Subscription filters do not alarm |
| Immutable long-term archive | Subscription filter + Firehose to S3 with Object Lock | Log group retention is mutable and bounded |
The rule of thumb: reach for Logs Insights when the question is "what happened in this account recently," and reach for a subscription filter when the question is "how do we make this evidence available to people and systems that are not in this account." The exam rarely asks you to choose one and exclude the other; it asks you to recognize which constraint in the scenario is the binding one.
Hands-On Lab: Query, Then Ship (45 min)
This lab has two halves that mirror the decision boundary above. First you will write Logs Insights queries against a log group you control, then you will build the aggregation path that moves a filtered subset of those events into a second account. Do it in a sandbox organization with at least two accounts; if you only have one, create a second account in the same organization and treat it as the central logging account.
Step 1 — Produce structured logs. Create a Lambda function that writes a JSON log line per invocation containing at minimum a request ID, a level, a duration in milliseconds, and a path. Invoke it a few hundred times with a mix of fast and slow paths so you have something to query. Confirm the log group exists and that the events are structured JSON rather than free text; the rest of the lab depends on this.
Step 2 — Query the slowest requests. Open Logs Insights against that log group and write a query that parses the JSON, filters to the last hour, sorts by duration descending, and returns the top 10. The shape is a fields clause that pulls @timestamp and the parsed fields, a parse stage that extracts them from the message, a filter on the time window, and a sort plus limit. Note how long the query takes and how many records it scanned — the console reports both, and that number is your cost signal.
Step 3 — Aggregate instead of returning rows. Rewrite the query to use stats to compute the count, average, and p99 duration grouped by path. Compare the two queries' scanned-record counts. The aggregation returns a handful of rows but scans the same events, which is the point: aggregation controls the size of the answer, not the cost of the scan.
Step 4 — Create the central destination. In the second account, create a Kinesis Data Firehose delivery stream with an S3 bucket as its destination. Set the buffer interval to 60 seconds so you are not waiting five minutes for results. Note the delivery stream ARN; you will need it in the next step.
Step 5 — Attach the subscription filter. Back in the producing account, create a subscription filter on the Lambda's log group with a pattern that matches only error-level events, pointing at the Firehose stream in the other account. You will need to grant CloudWatch Logs permission to write to the stream, which is done with an IAM role that the log group assumes. Verify the filter is listed on the log group.
Step 6 — Prove the path end to end. Invoke the Lambda with a payload that produces an error-level log line, wait for the buffer interval to elapse, and confirm the object appears in the S3 bucket in the central account. Then invoke it with an info-level line and confirm that one does not appear — that negative test is the one that proves your filter pattern is doing what you think.
Step 7 — Break it deliberately. Delete the subscription filter, generate more error events, and confirm nothing new arrives centrally while the source log group still has everything. This is the silent-failure mode from section 6, and seeing it once makes it recognizable in a scenario question. Recreate the filter and confirm delivery resumes.
Step 8 — Write the drift check. Sketch (do not necessarily deploy) a Config rule or a scheduled Lambda that lists log groups in the producing account and asserts each one has a subscription filter attached. This is the control that turns the pattern from "works for the accounts we onboarded carefully" into "works for every account."
Scenario Question Drills (20 min)
Q1. An organization wants every account's application logs centrally searchable in a dedicated logging account in near-real time. What's the mechanism?
Q2. A security team needs to search application logs from 40 accounts from a single console, and must not hold write permissions in the workload accounts. What is the correct design?
Q3. An application logs a 400 KB JSON document per request. Engineers report that the log lines are unparseable in Logs Insights. What is happening?
Q4. A team wants to be paged when the string OutOfMemoryError appears in application logs. What should they configure?
Q5. A compliance requirement mandates that application logs be retained for seven years and be immutable once written. Which design satisfies this?
Q6. A Logs Insights query over the last 30 days against a high-volume log group times out. What is the correct first move?
Q7. Logs stopped appearing in the central logging account three weeks ago, but the workload is healthy and the source log group contains everything. What is the most likely cause?
Q8. A team needs multiple independent consumers to read the same log stream, including the ability to replay the last 24 hours of events. Which destination type fits?
Q9. Which statement about the subscription filter pattern language is correct?
Q10. A workload runs in three regions and the organization wants all logs in one central account. What must the design account for?
Q11. A team wants to reduce CloudWatch Logs spend without losing the ability to debug production incidents. What is the most effective first step?
Q12. An engineer needs to correlate every log line from a single failing request across four microservices. What makes this practical?
Q13. A new log group is created by an application's Terraform module, but its logs never reach the central account. The module creates the log group and nothing else. What is missing?
Q14. A team needs full-text search across a multi-year log corpus with complex dashboards and fast query at high volume. What is the appropriate architecture?
Q15. A workload account is compromised and the attacker wants to stop evidence from reaching the central logging account. Which control prevents this?
Peek into Tomorrow
Everything in this day assumes the system is behaving badly in ways you did not choose. You write a query because something already went wrong, you build a pipeline because an auditor already asked, you alarm on an error string because a customer already complained. The logs are a record of failures that happened on their own schedule, which means the failure modes you are best prepared for are the ones that have already occurred at least once. That is a strange position for an architect to be in: the evidence you have is biased toward the failures you have already survived, and the ones you have not survived are precisely the ones you have no data on.
Tomorrow's topic inverts that. Fault Injection Service lets you schedule the failure instead of waiting for it — terminating EC2 instances, stressing CPU and network, failing over an AZ — and the interesting design question is not whether you can inject a fault but how you bound the blast radius while you do it. The mechanism for that bounding is a stop condition tied to a CloudWatch alarm, which raises a question this day's material cannot answer: if the alarm that halts your chaos experiment is built from the same metrics and logs you have been discussing, what happens when the experiment itself degrades the telemetry you are relying on to stop it?
Sources
- Amazon CloudWatch Logs Insights query syntax
- Real-time processing of log data with subscriptions
- Filter and pattern syntax for metric filters and subscription filters
- CloudWatch Logs concepts: log groups, streams, and events
- CloudWatch Logs quotas
- Kinesis Data Firehose buffering hints and intervals
- Centralized logging across accounts with CloudWatch Logs
- S3 Object Lock for WORM retention
- CloudTrail concepts (for the CloudTrail vs. CloudWatch Logs boundary)