Deep Review — Domain 4 Weak Areas (Cost Control)
Recap — From Migration Mechanics to Cost Mechanics
Yesterday's pass over Domain 3 was about movement: the 7 Rs as a classification discipline, MGN's continuous block-level replication, the SCT-then-DMS split for heterogeneous engines, DataSync for file and object transfer, and the Snow Family as the answer whenever the network itself is the bottleneck. The through-line was that migration questions are almost always decided by a constraint you cannot negotiate — a deadline, a licensing blocker, a link that cannot carry the volume in the window. Cost questions are the same shape but a different concern: the constraint is a commitment you are about to sign, and the wrong answer is usually wrong because it locks flexibility you still need.
That is the same lifecycle stage with a different concern. Migration decides how a workload arrives; cost optimization decides what you promise about it once it is running. Both domains punish the candidate who reaches for the most aggressive-sounding option. In migration that was Snowmobile for a 300TB dataset; here it is a three-year EC2 Instance Savings Plan for a fleet whose instance families are still moving. The distractor patterns below are the cost-domain versions of that same reflex.
Foundations You'll Need Today
Today's review is about cost, and cost questions on this exam are almost never about arithmetic. They are about commitments, visibility, and accountability — three ideas that only make sense once you know what the underlying billing mechanisms actually are. If you have never bought a discount on AWS, the vocabulary in the patterns below will read as a wall of acronyms. This section builds the five concepts the rest of the day assumes you already have.
Commitment-Based Discounts: Savings Plans and Reserved Instances
By default, AWS charges you for compute by the hour at what is called On-Demand pricing — you pay for exactly what you use, with no promise about the future, and you pay the highest rate for that freedom. If you are willing to promise that you will spend a certain amount on compute for the next one or three years, AWS will charge you a lower rate for the same machines. That promise is the whole product. A Savings Plan is a commitment to spend a fixed dollar amount per hour on compute; in exchange, the usage you run is billed at a discounted rate until the commitment is used up. A Reserved Instance is a similar promise, but it is attached to a specific instance type rather than to a dollar amount.
The important thing to understand is that these are billing arrangements, not machines. Buying a Savings Plan does not create a server, reserve a server, or guarantee that a server will be available when you ask for one. It only changes the price on the bill for usage you were going to run anyway. That distinction is the entire point of Patterns 1, 2, and 13 below, and it is the single most common thing candidates get wrong.
Spot Capacity and the Two-Minute Interruption Notice
AWS runs far more servers than any single customer needs at any moment, and the spare capacity is sold at a steep discount under the name Spot. The trade-off is that AWS can take the capacity back whenever it needs it for full-price customers. When that happens, your instance receives a two-minute warning and is then terminated. Two minutes is enough time to stop accepting new work and finish or hand off what is in flight; it is not enough time to save a long-running computation that has no checkpoint.
This is why Spot is described as suitable for "fault-tolerant" work. A stateless web server behind a load balancer is fault-tolerant because another server can absorb its traffic. A database holding the only copy of your data is not, because losing it loses the data. The exam uses Spot as a distractor by attaching it to workloads that sound cost-sensitive but are not actually restartable.
Right-Sizing and Utilization Data
Every EC2 instance comes in a size — a number of virtual CPUs and an amount of memory — and you choose that size when you launch it. If you pick a size larger than the workload needs, you pay for capacity you never use. Right-sizing is the practice of measuring how much of the machine the workload actually consumes and moving it to a smaller size that still fits. The measurement is called utilization, and AWS records it automatically as CloudWatch metrics: CPU percentage, memory percentage, network throughput, and so on.
Right-sizing matters here because it competes with discounting as a way to lower the bill, and the two are not interchangeable. Discounting reduces the rate you pay for a given size; right-sizing reduces the size. If a fleet is running at ten percent CPU, the larger saving is almost always to shrink the machine first, because a discount on an oversized machine is a discount on waste. The exam tests the ordering explicitly.
Cost Allocation Tags
A tag is a label you attach to an AWS resource — a key and a value, like CostCenter: 4471 or Environment: production. Tags are free-form and you can attach as many as you like. A cost allocation tag is a tag you have specifically activated in the billing console so that AWS will break your bill down along that dimension. Once activated, you can ask "how much did the 4471 cost center spend last month" and get a real answer.
The catch is that the answer is only as good as the tagging. If half your resources have no CostCenter tag, half your spend lands in an "untagged" bucket and the report is misleading in a way that is easy to miss. This is why the exam insists that tag enforcement comes before reporting: no reporting tool can attribute spend to a label that was never applied. Patterns 8 and 9 are both versions of this.
S3 Storage Classes and Retrieval Time
Amazon S3 stores objects — files, essentially — and offers several storage classes that differ in two ways: how much they cost per gigabyte per month, and how quickly you can get an object back when you ask for it. S3 Standard is the most expensive and returns objects immediately. The Glacier classes are far cheaper and take anywhere from minutes to many hours to produce an object once you request it. There is also a class called One Zone-IA that is cheaper because it stores your data in a single Availability Zone rather than replicating it across several.
The trap in cost questions is choosing the cheapest class without checking the retrieval requirement. If a scenario says records must be produced within a working day, the cheapest class may be too slow; if it says the data must survive the loss of an Availability Zone, the single-zone class is disqualified no matter how cheap it is. Price is only half the decision, and Patterns 10 and 11 are built on candidates who forget the other half.
With that grounding — commitments are billing constructs, Spot is interruptible, right-sizing precedes discounting, tags gate reporting, and storage classes trade price against retrieval — here is why Domain 4 questions are written the way they are, and what each distractor is actually testing.
Distractor Patterns — Domain 4
Domain 4 questions rarely ask you to compute a number. They ask you to pick a commitment, a visibility tool, or an accountability mechanism, and the wrong options are engineered to sound either more aggressive or more precise than the right one. The fifteen patterns below are the ones that recur across practice banks. Each entry names the tempting answer, explains why it is tempting, and gives the tell that separates it from the correct choice.
Pattern 1 — The Most Flexible Commitment Is Not Always the Right One
The tempting answer is always Compute Savings Plans, because they apply across any instance family, region, and operating system, and they even cover Fargate and Lambda usage. That flexibility is real and it is the reason the option appears so often. The trap is that a question can describe a fleet that is genuinely frozen — a fixed instance family, a fixed region, a workload that has not changed shape in two years and will not change shape during the commitment term — and still offer Compute Savings Plans as the "safe" answer.
The tell is whether the scenario gives you any reason to believe the fleet will change. If the question says the workload is steady-state and predictable and the only stated goal is maximum discount, the correct answer is the EC2 Instance Savings Plan, which trades flexibility for a deeper rate. If the question mentions any of the things that force change — a pending migration, a re-platform, a move to containers, a region expansion — Compute Savings Plans win because the alternative would strand the commitment. Read the scenario for verbs of change, not for adjectives of stability.
Pattern 2 — Reserved Instances When You Actually Need a Capacity Reservation
Savings Plans and Reserved Instances are both commitment-based discounts, so a question that only mentions cost will usually accept either. The distractor is the reverse case: a scenario that quietly requires guaranteed capacity — a regulated workload that must be able to launch a specific instance type in a specific AZ during a regional event, or a capacity-constrained instance family — where a Savings Plan is offered as the answer because it is cheaper and more flexible.
The tell is the word "guarantee," or any phrasing about being able to launch during a shortage. Savings Plans are a billing construct; they discount usage but reserve nothing. Only zonal Reserved Instances carry a capacity reservation, and only for that specific instance type in that specific AZ. If the scenario's requirement is availability rather than price, the discount mechanism is the wrong axis entirely and the answer is the RI with the reservation, or an On-Demand Capacity Reservation paired with whatever discount instrument you like.
Pattern 3 — Spot for Anything With a Stateful or Long-Running Component
Spot pricing is the most aggressive discount on the exam and therefore the most attractive distractor. The option shows up attached to workloads that are described as "cost-sensitive" without any statement about fault tolerance, and candidates who have internalized "Spot is up to 90% off" reach for it. The interruption notice is only two minutes, which is enough to drain a stateless web tier and nowhere near enough to checkpoint a long-running simulation or fail over a primary database.
The tell is whether the scenario describes work that can be restarted from scratch without consequence. Batch rendering, CI runners, stateless API workers behind a load balancer, and big-data jobs with checkpointing all qualify. Anything holding a lock, a leader election, a session, or an in-memory state that would be lost does not. A second tell is whether the scenario mentions a mixed-instance policy or Spot Fleet — if the question is testing Spot correctly, it usually gives you the diversification mechanism as part of the setup rather than asking you to invent it.
Pattern 4 — Confusing Right-Sizing With Discounting
Compute Optimizer and Savings Plans both reduce your bill, so a question about an over-provisioned fleet can be answered with either if you are reading only the headline. The distractor is a commitment purchase offered for a fleet whose utilization data shows the instances are simply too large. Committing to three years of an oversized instance locks in the waste at a discount, which is worse than fixing the size and paying On-Demand for a month while you measure.
The tell is whether the scenario gives you utilization numbers. If it says instances are running at low CPU or memory utilization, the question is about right-sizing and the answer is Compute Optimizer's recommendations. If it says utilization is high and stable and the goal is to reduce the rate, the question is about commitment. A scenario that gives you both — high utilization on an oversized family — is testing whether you right-size first and commit second, in that order.
Pattern 5 — Cost Explorer as a Forecasting Tool vs. a Reporting Tool
Cost Explorer appears in almost every cost question and is correct in most of them, which makes it a weak distractor and a dangerous one. The trap is a scenario that needs granular, line-item billing data exported to a data warehouse for custom analysis, where Cost Explorer is offered because it is the tool everyone knows. Cost Explorer is a visualization and forecasting console; it is not the source of the most granular billing data.
The tell is the word "granular," or any mention of hourly or resource-level line items, or of feeding a BI tool. That is the Cost and Usage Report, which lands in S3 and can be queried with Athena. Cost Explorer is the right answer when the question is about trends, forecasts, or the RI and Savings Plan purchase recommendations it generates. When the question is about the raw data underneath those views, the answer moves to CUR.
Pattern 6 — Trusted Advisor as a Right-Sizing Engine
Trusted Advisor is a broad best-practice scanner covering cost, security, performance, fault tolerance, and service quotas, and its cost category does flag idle resources. That breadth makes it a plausible answer to almost any cost question, including ones that need a specific instance-type recommendation. The distractor is a scenario asking which service will tell you the exact smaller instance type to move to.
The tell is whether the question wants a recommendation or a flag. Trusted Advisor tells you that you have low-utilization instances; Compute Optimizer tells you which instance type to move to and what the projected savings are, based on historical utilization metrics. If the scenario asks for a list of problems, Trusted Advisor is fine. If it asks for a specific target configuration, it is Compute Optimizer.
Pattern 7 — Budgets as an Enforcement Mechanism
AWS Budgets alerts on actual and forecasted spend, and the distractor is a scenario that needs spending to actually stop when a threshold is crossed. Budgets can trigger actions, but the reflex answer — "create a budget" — does not by itself prevent anything. The question is usually testing whether you understand that alerting and enforcement are separate concerns.
The tell is the verb. If the scenario says "notify," "alert," or "warn," a budget with a notification threshold is correct. If it says "prevent," "block," or "stop," you need something that acts on the account — a budget action that applies an SCP or stops instances, or an SCP that denies expensive resource creation outright. A budget alone will happily let you spend past the threshold while it emails you about it.
Pattern 8 — Tagging as a Reporting Fix Rather Than a Prerequisite
Cost allocation by team, project, or environment is a common scenario, and the tempting answer is to enable Cost Explorer or build a report. The trap is that the scenario has already told you tagging is inconsistent. No reporting tool can attribute spend to a dimension that is not recorded on the resource, so the report will simply show a large "untagged" bucket.
The tell is any sentence about tags being missing, inconsistent, or applied after the fact. The correct answer is to enforce the tags at creation time — an SCP that denies resource creation without the required cost-allocation tags, or a Config rule that flags and remediates untagged resources — and only then activate those tags for cost allocation. The ordering matters and the exam tests it: enforcement first, reporting second.
Pattern 9 — The Single-Account Answer to a Multi-Account Cost Problem
When a scenario describes an organization with many accounts and asks how to see total spend, one distractor is to consolidate everything into one account so the bill is simpler. That solves the reporting problem by destroying the isolation the organization was built for, and it is almost never the intended answer.
The tell is whether the scenario mentions AWS Organizations, OUs, or separate accounts for security or blast-radius reasons. If it does, the answer is consolidated billing at the management account plus cost allocation tags and CUR for the breakdown — you get one bill and per-account attribution without merging accounts. The distractor is attractive because it is simple, and the exam is specifically checking whether you will trade governance for simplicity.
Pattern 10 — Storage Class Selection Driven by Price Alone
Cost questions that touch storage often offer the cheapest class as the answer, and the cheapest class is usually Glacier Deep Archive. The trap is a scenario with a retrieval requirement — a compliance hold that legal may need to produce within hours, or an analytics job that reads the data monthly — where the price is right and the retrieval time is not.
The tell is any stated retrieval latency or access frequency. Deep Archive is correct when the data is genuinely cold and a twelve-hour retrieval is acceptable. If the scenario says the data is accessed occasionally, or that retrieval must complete within a working day, the answer moves up the ladder to Glacier Flexible Retrieval or Standard-IA. A second tell is durability scope: if the scenario requires surviving the loss of an entire Availability Zone, One Zone-IA is disqualified regardless of price.
Pattern 11 — Intelligent-Tiering as a Universal Answer
S3 Intelligent-Tiering is the answer that requires no analysis, which makes it the default distractor for any storage-cost question. It monitors access patterns and moves objects between tiers automatically, so it is genuinely correct when access patterns are unknown or changing. It is the wrong answer when the scenario tells you the access pattern explicitly.
The tell is whether the scenario gives you a known lifecycle. If it says logs are written once and read only during incident response, or that objects are hot for thirty days and then never touched, a lifecycle policy with explicit transitions is cheaper because Intelligent-Tiering charges a monitoring and automation fee per object. If the scenario says access patterns are unpredictable or the team does not know how the data will be used, Intelligent-Tiering is the right call and the fee is the price of not guessing.
Pattern 12 — Data Transfer Costs Attributed to the Wrong Hop
Cost scenarios involving multi-AZ or multi-region architectures often include a distractor that blames the wrong component. The classic is a scenario where a chatty application crosses Availability Zones on every request, and the offered answers focus on instance sizing or storage class rather than the cross-AZ traffic itself.
The tell is whether the architecture diagram or description shows traffic crossing a boundary — AZ, region, or the internet — on a high-frequency path. Cross-AZ traffic is billed in both directions, and a service that makes many small calls between AZs can cost more in transfer than in compute. The correct answer is usually to co-locate the chatty components in a single AZ with the understanding that you are trading some resilience, or to reduce the call volume. If the scenario mentions NAT Gateway data processing charges, the same logic applies: the fix is a gateway endpoint for S3 and DynamoDB traffic, not a bigger NAT.
Pattern 13 — Savings Plans Applied to the Wrong Scope
Both Savings Plan types are commitment instruments, and a scenario can be constructed where the candidate picks the right type but the wrong scope. The distractor is an EC2 Instance Savings Plan offered for a workload that includes Fargate tasks or Lambda invocations, where the commitment would not apply to that portion of the spend at all.
The tell is the composition of the bill. If the scenario mentions containers on Fargate, Lambda functions, or a mix of compute types, only Compute Savings Plans cover that breadth. EC2 Instance Savings Plans apply to EC2 usage within a specific instance family in a specific region and nothing else. Read the scenario for the words "Fargate" and "Lambda" specifically — they are the signal that the narrower plan is disqualified.
Pattern 14 — Spot Fleet Without Interruption Handling
Spot Fleet and mixed-instance Auto Scaling Groups are the correct mechanism for running fault-tolerant work on Spot capacity, and the distractor is a scenario that names Spot Fleet but never addresses what happens when a Spot instance is reclaimed. The exam expects you to know that the two-minute notice has to be consumed by something.
The tell is whether the scenario mentions draining, checkpointing, or a handler. A correct Spot answer includes a mechanism that reacts to the interruption notice — a lifecycle hook that drains connections, a checkpoint written to durable storage, or a queue that lets the work be picked up elsewhere. If the scenario describes a workload that cannot tolerate any interruption and the offered answer is Spot with a diversification policy, the diversification reduces the frequency of interruptions but does not eliminate them, and the answer is wrong.
Pattern 15 — The Commitment That Outlives the Workload
The last pattern is the one that ties the domain together. A scenario describes a project with a known end date — a six-month migration, a seasonal campaign, a proof of concept — and offers a one- or three-year commitment because the discount is larger. The math in the scenario is a trap: the discount only pays if the usage continues for the full term.
The tell is any explicit duration on the workload itself. If the scenario says the environment will be decommissioned, or that the workload is temporary, the correct answer is On-Demand or Spot for the duration, or a shorter commitment if one is offered. This is the same reflex that produced the wrong answers in yesterday's migration review — reaching for the most powerful tool without checking whether the constraint permits it. In Domain 4, the constraint is almost always time.
Hands-On Lab — Building a Cost Review From Your Own Bill (45 min)
The most useful version of this review is not abstract. Take a real or sample account and work the patterns in order, because the order matters: you cannot commit to anything until you know what you are running, and you cannot attribute anything until the tags exist.
Start in Cost Explorer and group by service for the last three months. Identify the top three services by spend and write down, for each, whether the cost is driven by rate, by volume, or by waste. Rate problems are commitment candidates; volume problems are architectural; waste problems are right-sizing. Most accounts have one of each and the temptation is to buy a Savings Plan to fix all three, which only addresses the first.
Next, open Compute Optimizer and look at the recommendations for the same period. For every over-provisioned instance it flags, note the recommended type and the projected savings. Compare that number against what a one-year Compute Savings Plan would save on the same instance at its current size. In most fleets the right-sizing number is larger, which is the empirical version of Pattern 4.
Then check tag coverage. Pick the cost allocation tags you would want — CostCenter, Environment, Owner — and count how many resources in the top three services actually carry them. If coverage is below roughly ninety percent, stop and fix that before doing anything else, because every report you build afterward will be wrong in a way that is hard to notice. Write the SCP or Config rule that would enforce the tags at creation time, even if you do not apply it.
Finally, build one budget per cost center with a forecasted-spend alert at eighty percent and an actual-spend alert at one hundred percent, and decide explicitly whether either should trigger an action rather than a notification. That decision is Pattern 7 in practice, and writing it down forces you to be honest about whether you want enforcement or just visibility.
Finish by writing a one-page summary in the same shape as the pattern entries above: for each of the three services, the tempting fix, the correct fix, and the tell that distinguishes them. That page is what you review on Day 68, and it is more valuable than any of the individual numbers you collected.
Drill — 15 Distractor-Pattern Questions
Each question below is built around one of the patterns above, and each includes the tempting wrong answer as a live option. Answer before revealing.
Q1. A company runs a steady-state EC2 fleet in a single region on a single instance family. The fleet has not changed in two years and there are no plans to change it. The only stated goal is the largest possible discount. Which commitment should they purchase?
Q2. A regulated workload must be able to launch a specific instance type in a specific Availability Zone during a regional capacity shortage. Cost is a secondary concern. What should be provisioned?
Q3. A team runs a nightly batch rendering job that processes independent frames and can restart any frame from scratch. They want the lowest possible cost. What is the appropriate capacity strategy?
Q4. Compute Optimizer reports that a fleet of m5.2xlarge instances is running at under ten percent CPU utilization. The team wants to reduce cost. What should they do first?
Q5. Finance needs hourly, resource-level billing line items loaded into a data warehouse for custom analysis. Which source provides this?
Q6. A team wants to know which specific smaller instance type to move their over-provisioned instances to, with projected savings. Which service answers this?
Q7. A company wants spending to actually stop when a project exceeds its monthly budget, not just be reported. What should be configured?
Q8. Finance wants spend broken out per business unit, but an audit shows that fewer than half of the resources carry any cost allocation tags. What should be done first?
Q9. An organization with forty accounts under AWS Organizations wants a single view of total spend with per-account attribution, without weakening account isolation. What is the correct approach?
Q10. A compliance team must be able to produce archived records within a working day if requested, and the records are otherwise never accessed. The data must survive the loss of an entire Availability Zone. Which storage class fits?
Q11. A team has a large S3 bucket whose access patterns are genuinely unpredictable — some objects are read daily, others never. They want to minimize cost without analyzing access patterns themselves. What should they use?
Q12. A three-tier application makes a high volume of small calls between services in different Availability Zones, and the bill shows unexpectedly high data transfer charges. What is the most likely cause and fix?
Q13. A company's bill includes significant EC2 usage, a large Fargate footprint, and steady Lambda invocations. They want a single commitment instrument covering all three. What should they buy?
Q14. A team runs a fault-tolerant workload on Spot Fleet and wants to minimize the impact of interruptions. Which addition is required for the design to be complete?
Q15. A six-month migration project needs a temporary environment that will be decommissioned when the project ends. The team wants to minimize cost. What is the appropriate commitment?
Preview — Mock Exam 3
Everything in this review assumed you already know which answer is right and are working on why the wrong ones are tempting. That assumption is about to be tested under conditions where you will not have time to run the pattern list in your head. Mock Exam 3 is the third full 75-question, 180-minute sitting, and by this point the score itself matters less than the shape of the misses. If the cost patterns above have landed, Domain 4 questions should stop producing the "I narrowed it to two and picked wrong" outcome and start producing clean, fast answers.
The open question is whether the deep reviews on Days 59 through 64 have actually moved the needle, or whether they have only made the material feel more familiar. Familiarity is the failure mode here: recognizing a pattern on review and applying it under a 2.4-minute-per-question clock are different skills, and only the timed sitting distinguishes them. Take Mock Exam 3 without pausing, without notes, and without the pattern list open, and let the score tell you which of those two things you actually have.