Cost Optimization — Savings Plans, RIs & Spot Fleet Strategy
Recap: Day 52
Day 52 was about provisioning capacity ahead of a migration window — Link Aggregation Groups (LAG) to bundle Direct Connect connections, Snowball for the bulk historical dataset, and DX/DMS CDC for the ongoing delta. The lesson there was that a migration plan which only accounts for the steady-state network path will fail on the day the bulk copy actually runs, because the transfer itself is a temporary, enormous, and entirely predictable load that nobody sized for.
Today sets up the same trap in a different service. The commitment instruments — Savings Plans, Reserved Instances, and Spot — are all capacity and pricing decisions made in advance of the workload that consumes them, and each one has a failure mode that only appears when the workload's actual shape diverges from the shape you assumed at purchase time. A LAG sized for average throughput and a Compute Savings Plan sized for average compute both look correct on the day you buy them. The exam tests whether you can tell which commitment instrument tolerates divergence and which one quietly converts your discount into a stranded liability.
Foundations You'll Need Today
Today's material is written in the vocabulary of AWS billing, and that vocabulary only makes sense if you already know what is being billed. Before the commitment instruments make any sense, here are the five pieces of background this day quietly assumes.
EC2 Instances, Families, and Sizes
An EC2 instance is a virtual server you rent by the hour (or by the second, for some types). When you launch one, you pick an instance type, which is a combination of a family and a size. The family describes what the machine is optimized for — general purpose, compute-heavy, memory-heavy, storage-heavy, or accelerated with GPUs — and the size describes how much of that resource you get, from a couple of virtual CPUs up to machines with hundreds. The family name is the part before the dot and the size is the part after it, so a m5.large and an m5.2xlarge are the same family at different sizes, while an m5.large and an r5.large are different families at the same size. This matters today because a Reserved Instance is a commitment to a specific family and size, and a Savings Plan is a commitment that ignores family and size entirely. That single distinction is the source of most of the day's trade-offs.
On-Demand Pricing and Why It Is the Baseline
On-Demand is the default way to pay for compute: you launch an instance, you pay the published hourly rate for as long as it runs, and you stop paying when you stop it. There is no contract, no minimum, and no discount. Every other pricing instrument in this day is described as a percentage off On-Demand, which is why On-Demand is the reference point rather than just another option. When a scenario says a workload is "steady-state and predictable," it is really saying that the workload runs enough hours that paying the undiscounted rate is leaving money on the table — and that is the situation the commitment instruments exist to fix.
Auto Scaling Groups and Mixed-Instance Fleets
An Auto Scaling Group (ASG) is a controller that keeps a target number of EC2 instances running on your behalf. You tell it a desired capacity — say, six instances — and it launches instances when you have fewer and terminates them when you have more. It launches them from a launch template, which is a saved recipe describing the AMI (the disk image), the instance type, the security groups, and the startup script. A mixed-instance ASG is one where the launch template offers several acceptable instance types instead of pinning one, and the ASG is allowed to pick whichever is available. That flexibility is what makes Spot capacity practical, because a fleet that will accept any of six comparable instance types can almost always find room somewhere, while a fleet that demands one exact type often cannot.
Regions and Availability Zones
A region is a geographic area — us-east-1, eu-west-2 — and each region contains several Availability Zones, which are physically separate data centers with independent power and networking. An AZ is identified by appending a letter to the region, so us-east-1a and us-east-1b are two different AZs in the same region. The reason this matters today is that "regional" and "zonal" are not synonyms in AWS pricing. A regional Reserved Instance gives you a discount that floats across every AZ in the region, while a zonal Reserved Instance pins the discount — and, crucially, a capacity reservation — to one specific AZ. Spreading a fleet across multiple AZs is also the standard way to survive the failure of a single data center, which is why the Spot diversification advice in this day tells you to spread across AZs as well as families.
Load Balancers, Target Groups, and Connection Draining
A load balancer sits in front of your instances and distributes incoming requests across them. It does not send traffic to instances directly; it sends traffic to a target group, which is a list of instances that have been registered as healthy and eligible. When an instance is about to go away — because it is being replaced, or because Spot is about to reclaim it — the correct sequence is to deregister it from the target group first, which tells the load balancer to stop sending it new requests, and then wait for the in-flight requests it is already handling to finish. That waiting period is called connection draining. Today's lab builds an interruption handler around exactly this sequence, and the two-minute Spot notice is the budget it has to complete in.
With that grounding, here's why the commitment instruments exist and what problem each one actually solves.
1. Why This Is on the Exam
Cost optimization is Domain 4 of SAP-C02, and it is the domain candidates most often underestimate because it feels like arithmetic rather than architecture. It is not arithmetic. Every cost question on the exam is really a question about which constraint the business is willing to accept: a commitment term, a capacity guarantee, an interruption risk, or a migration effort. The pricing instruments are just the vocabulary the question is written in.
The reason this cluster carries real weight is that it is the one place where a technically correct architecture can still be the wrong answer. A design that satisfies every availability and performance requirement can be marked incorrect because it leaves a steady-state fleet on On-Demand pricing, or because it puts a stateful, non-idempotent workload on Spot, or because it buys a three-year commitment for a workload the business has already told you will be re-platformed. The exam expects you to read the business constraints in the scenario — predictability, tolerance for interruption, willingness to commit, and how long the workload will exist in its current form — and select the instrument that matches.
There is also a structural reason this day sits where it does. The migration phase just covered the 7 Rs, and the 7 Rs are themselves a cost decision: Rehost is cheap and fast, Refactor is expensive and slow, Retain is free today and expensive later. Cost optimization is the discipline that quantifies those trade-offs after the migration lands. The exam will hand you a workload that has just been rehosted and ask what to do with its pricing, which is exactly the question this day answers.
2. How the Commitment Instruments Actually Work
All three instruments are mechanisms for trading flexibility for price, but they trade different things, and the difference matters more than the discount percentage. A Savings Plan is a commitment to spend a fixed dollar amount per hour on compute for a term. It is not a commitment to a specific instance; it is a commitment to a billing rate. AWS applies that commitment against whatever eligible usage you generate, and anything above the commitment is billed at On-Demand rates. This is why a Savings Plan can be partially utilized without being wasted — you simply stop receiving the discount on the portion of usage that exceeds the commitment.
A Reserved Instance is a commitment to a specific instance configuration in a specific scope. You are reserving a shape, not a spend rate. That specificity is what buys the deeper discount, and it is also what creates the stranded-capacity problem: if the workload moves to a different family, a different region, or a different tenancy, the reservation no longer applies to it and the discount evaporates while the commitment continues to bill. Convertible RIs soften this by allowing an exchange for a different configuration of equal or greater value, but the exchange is a manual operation with its own constraints, not an automatic reallocation.
Spot is a different mechanism entirely. Spot capacity is spare EC2 capacity that AWS sells at a discount with the explicit understanding that it can be reclaimed. The reclaim is announced through a two-minute interruption notice delivered to the instance metadata service and to EventBridge, after which the instance is terminated or stopped depending on the request configuration. Spot is not a pricing tier you graduate into; it is a capacity class with a different availability contract. The exam consistently rewards candidates who treat the two-minute notice as an architectural requirement rather than a footnote — a workload that cannot drain, checkpoint, or re-route within two minutes is not a Spot workload, regardless of how much it would save.
3. The Core Decision Boundary
The single fork that most scenario questions hinge on is whether the workload's shape is predictable and stable enough to commit to, and if so, how much specificity the commitment can tolerate. Predictability is the gate. A workload with a stable baseline and modest variance can be covered by a commitment; a workload whose baseline moves every quarter cannot, because the commitment will be either under-utilized or exceeded, and both outcomes are worse than On-Demand for different reasons.
Once predictability is established, the second question is how much the workload is expected to change. If the instance families, regions, or operating systems are likely to shift — because the team is mid-migration, or because the workload is being re-platformed onto Graviton, or because a region strategy is still being decided — then the commitment must be flexible, and the flexibility costs discount depth. If the workload is genuinely frozen in shape for the term, the specific commitment is strictly better. The exam phrases this as "maximum flexibility" versus "maximum savings," and the correct answer is whichever the scenario's constraints actually demand.
| Instrument | What you commit to | Flexibility | Discount depth | Capacity guarantee |
|---|---|---|---|---|
| Compute Savings Plan | $/hour of compute spend | Any family, region, OS, tenancy; includes Fargate and Lambda | Moderate | No |
| EC2 Instance Savings Plan | $/hour within one family in one region | Any size, OS, tenancy within that family/region | Deeper | No |
| Standard RI | A specific instance configuration | None (region or zonal scope only) | Deepest | Optional, zonal only |
| Convertible RI | A specific configuration, exchangeable | Exchange for equal-or-greater value | Moderate | Optional, zonal only |
| Spot | Nothing — capacity is reclaimable | N/A | Largest, variable | None, by definition |
4. Configuration Modes and Their Tradeoffs
Each instrument has configuration knobs that change what it covers, and the exam tests the knobs more often than the headline discount. For Savings Plans the primary knob is the hourly commitment amount, which is a spend rate rather than an instance count. Setting it too low leaves usage uncovered; setting it too high means you are paying for compute you did not consume, because the commitment bills whether or not you use it. The practical approach is to cover the stable baseline and let the variable portion ride at On-Demand, which is the same logic as the mixed-instance Auto Scaling Group covered in the lab.
For Reserved Instances the knobs are term length, payment option, scope, and offering class. Term length is one or three years, with the three-year term carrying the deeper discount and the longer exposure. Payment option — All Upfront, Partial Upfront, or No Upfront — trades cash flow against discount depth, with All Upfront cheapest and No Upfront most expensive. Scope is regional or zonal, and this is the knob candidates most often miss: a zonal RI can reserve capacity in a specific Availability Zone, which is the only way any of these instruments guarantees you capacity at all. A regional RI gives you the billing discount and the flexibility to move between AZs, but no capacity reservation.
Spot configuration is where the interruption contract becomes concrete. A Spot request can be configured to terminate, stop, or hibernate on interruption, and the choice determines what the workload can recover. Terminate is the default and loses everything on the instance. Stop preserves the EBS volumes and allows a later restart, which is useful for workloads that can tolerate a gap. Hibernate preserves the in-memory state, which is the closest thing to a checkpoint but requires the instance to support hibernation and the workload to tolerate the pause. The exam will describe a workload that needs to survive interruption and expect you to recognize that the answer is not a Spot configuration at all — it is a Spot configuration plus an application that can checkpoint, or a decision to move that workload off Spot entirely.
5. Sizing, Limits and Quotas
The numbers that matter here are the ones that constrain how much of a fleet you can actually cover with each instrument, and the ones that define the interruption window. Spot's two-minute interruption notice is the hard architectural constraint: it is the entire budget a workload has to drain connections, flush state, and deregister from a load balancer. Anything that cannot complete in that window needs a different capacity strategy, and the exam will describe workloads whose shutdown sequence obviously exceeds it.
Spot also has a service quota on the number of Spot Instance requests and on vCPUs, and the vCPU quota is the one that bites during a scale-out event. A mixed-instance group that tries to burst into Spot capacity beyond the account's Spot vCPU quota will simply fail to launch the Spot portion, which means the burst never happens and the workload degrades to its On-Demand baseline. This is a quota problem, not a pricing problem, and it is exactly the kind of thing the Reliability pillar's "service quotas as foundations" guidance is pointing at.
| Constraint | Value / behavior | Why it matters |
|---|---|---|
| Spot interruption notice | Two minutes before reclaim | Defines the maximum graceful-shutdown budget |
| Spot request types | One-time or persistent | Persistent requests re-launch after interruption; one-time do not |
| Spot allocation strategy | lowest-price, capacity-optimized, price-capacity-optimized | capacity-optimized reduces interruption frequency at some cost |
| Savings Plan term | One or three years | Longer term, deeper discount, longer exposure |
| RI payment options | All Upfront, Partial Upfront, No Upfront | Trades cash flow against discount depth |
| RI scope | Regional or zonal | Only zonal RIs reserve capacity in a specific AZ |
One more number worth internalizing: the discount on Spot is described as up to roughly ninety percent off On-Demand, but it is variable and moves with supply and demand in the pool. Any scenario that quotes a fixed Spot savings figure is describing an assumption, not a guarantee, and the exam is more interested in whether you understand the variability than in the specific percentage.
6. Failure Modes and What They Look Like in Production
The first failure mode is the stranded commitment. It presents as a cost anomaly rather than an availability incident: the monthly bill shows a Savings Plan or RI charge that is no longer offset by matching usage, because the workload moved to a different family, region, or account. The symptom is a utilization metric on the commitment itself — AWS reports Savings Plan and RI utilization and coverage, and a utilization figure that has dropped below the baseline is the signal. The first diagnostic move is to compare the commitment's coverage against the current fleet's actual instance mix, which usually reveals that a migration or a Graviton adoption quietly invalidated the reservation.
The second failure mode is Spot interruption during a non-idempotent operation. This presents as intermittent, hard-to-reproduce data corruption or duplicate processing, and it is the classic "same trap different service" pattern: the workload was designed for a capacity class that never disappears, and it was moved to one that does. The symptom is a spike in retries, duplicate records, or partial writes correlated with Spot interruption events in EventBridge. The first diagnostic move is to check whether the workload's write path is idempotent and whether the interruption handler actually drains in-flight work, rather than assuming the interruption itself is the problem.
The third failure mode is Spot capacity unavailability during a scale-out. This presents as a fleet that fails to reach its desired capacity exactly when demand is highest, because the Spot pools it targets are thin. The symptom is a scaling activity that reports insufficient capacity, and the fix is diversification — spreading the request across multiple instance families and Availability Zones, or using the capacity-optimized allocation strategy, which selects pools with the most available capacity rather than the lowest price. A single-family Spot request is the most fragile configuration available, and it is the one candidates most often propose.
7. The Operational and SRE Angle
From an operations standpoint, commitment instruments are a monitoring problem before they are a purchasing problem. Savings Plan utilization and coverage, RI utilization and coverage, and Spot interruption frequency are all first-class metrics that belong on a dashboard alongside the workload's own SLOs, because a commitment that has drifted out of alignment is a slow, silent cost leak that no availability alarm will ever catch. The runbook shape is a periodic review: compare coverage against the current fleet, identify commitments whose utilization has fallen below the expected baseline, and decide whether to exchange, resell on the RI Marketplace where eligible, or simply let the term expire.
Spot changes the reliability posture of the workload itself, and that has to be reflected in the SLO rather than hidden from it. A service running a mixed fleet has a baseline capacity that is guaranteed and a burst capacity that is not, so its effective availability during a demand spike depends on Spot pool depth as much as on its own code. The honest way to model this is to treat the On-Demand baseline as the capacity floor and the Spot portion as elastic headroom, then set the SLO against the floor and monitor how often the headroom is actually available. A service that quietly depends on Spot headroom to meet its SLO has an availability risk that is invisible in its own metrics.
The interruption handler is the operational artifact that makes Spot safe, and it deserves the same treatment as any other failure-handling code. It should be tested, not assumed — the same discipline the chaos engineering days applied to AZ failure applies here, and an interruption handler that has never been exercised under load is an untested assumption. The handler's job is narrow and specific: deregister from the load balancer, stop accepting new work, finish or checkpoint in-flight work, and exit within the two-minute window. Anything the handler cannot complete in that window is work that will be lost, and that should be a deliberate decision rather than a discovery.
8. Edge Cases and Exam Gotchas
The most common trap is conflating a billing discount with a capacity guarantee. A regional Reserved Instance and a Savings Plan both reduce the bill and neither reserves capacity; only a zonal RI does. A scenario that says the workload must be guaranteed to launch in a specific Availability Zone during a regional capacity crunch is asking for a zonal RI, and a scenario that only asks for cost reduction is not. Read the requirement for the word "guarantee" and treat it as decisive.
The second trap is assuming a Savings Plan is wasted when it is under-utilized. It is not — the unused portion simply stops discounting, and the commitment continues to bill at the agreed rate. This is a real cost, but it is a bounded one, and it is materially different from a stranded RI, which cannot be reallocated to a different family at all. The exam will sometimes describe a workload that is expected to change shape and expect you to choose the flexible instrument precisely because partial utilization is recoverable while a mismatched reservation is not.
The third trap is treating Spot as suitable for any stateless workload. Statelessness is necessary but not sufficient; the workload also has to tolerate a two-minute termination, which rules out long-running batch jobs without checkpointing, in-memory session stores, and anything holding a lock. The fourth trap is forgetting that Spot and Savings Plans interact: a Savings Plan commitment applies to On-Demand usage, and Spot usage is billed separately, so a fleet that shifts heavily toward Spot can leave a Savings Plan under-utilized. The fifth is the RI Marketplace, which allows you to sell Standard RIs you no longer need — a detail the exam occasionally uses to distinguish a recoverable mistake from an unrecoverable one.
9. This vs. the Instruments It Gets Confused With
The cleanest way to hold all of this is to separate the instruments by what they actually promise. Savings Plans promise a discount on a spend rate. Reserved Instances promise a discount on a configuration, and optionally a capacity reservation in one AZ. Spot promises a discount in exchange for accepting reclaim. On-Demand promises nothing except availability at list price. Every scenario question is asking you to pick the promise that matches the workload's constraints, and the constraints are almost always stated in the scenario's first two sentences.
| Pick this | When the scenario says… |
|---|---|
| Compute Savings Plan | Steady usage, but families/regions/OS may change; or the fleet includes Fargate or Lambda |
| EC2 Instance Savings Plan | Steady usage locked to one family in one region, and you want a deeper discount than Compute |
| Standard RI | Frozen configuration for the term, and you want the deepest discount |
| Zonal RI | You need a capacity reservation in a specific Availability Zone |
| Convertible RI | You want RI depth but expect to exchange the configuration mid-term |
| Spot with mixed-instance ASG | Fault-tolerant, interruptible, checkpointable work; burst capacity above a stable baseline |
| On-Demand | Unpredictable or short-lived usage; no commitment is justified |
The pairing that shows up most often is a mixed-instance Auto Scaling Group: an On-Demand baseline sized to the guaranteed floor, Spot capacity above it for burst, and a diversification strategy across families and AZs so that no single thin pool can starve the fleet. That combination is the standard answer for "cost-optimize this fault-tolerant workload without risking availability," and it is the design the lab below builds.
Hands-on Lab: A Mixed-Instance Fleet with a Tested Interruption Handler
Objective. Build an Auto Scaling Group that runs a guaranteed On-Demand baseline with Spot burst capacity above it, and prove that the interruption handler actually drains in-flight work inside the two-minute window.
Step 1 — Establish the baseline. Pick a small stateless web workload and measure its steady-state instance count over a week. That number, not the peak, is the On-Demand baseline. Write it down; the entire design depends on not confusing the floor with the average.
Step 2 — Build the launch template with a diversification-friendly instance type list. Create a launch template that does not pin a single instance type. Instead, supply a list of several comparable families and sizes so the ASG can substitute when a pool is thin. This is the single highest-leverage change for Spot reliability.
Step 3 — Configure the mixed-instances policy. Set the On-Demand base capacity to the baseline from Step 1, set On-Demand percentage to zero above the base, and set the Spot allocation strategy to capacity-optimized. Capacity-optimized selects the pools with the most available capacity rather than the lowest price, which trades a little discount for materially fewer interruptions.
Step 4 — Write the interruption handler. Subscribe a handler to the EC2 Spot Instance Interruption Warning event in EventBridge. The handler should immediately deregister the instance from the target group, wait for connection draining to complete, finish or checkpoint any in-flight work, and then allow termination. Log the elapsed time from event receipt to drain completion — that number is the thing you are testing.
Step 5 — Test it under load. Drive real traffic at the fleet, then trigger an interruption deliberately. Confirm that no request fails, that the instance leaves the target group before termination, and that the drain completes well inside two minutes. If it does not, the workload is not a Spot workload yet, and the fix is in the application, not the ASG configuration.
Step 6 — Add the cost guardrails. Create a Savings Plan sized to the On-Demand baseline only, not the peak. Then build a CloudWatch dashboard showing Spot interruption frequency, ASG capacity versus desired capacity, and Savings Plan utilization side by side, so a drift in any of the three is visible before it shows up on the bill.
Step 7 — Break it on purpose. Reduce the instance type list back to a single family and re-run the scale-out test during a busy period. Watch the fleet fail to reach desired capacity. That failure is the lesson: diversification is not a nice-to-have, it is the mechanism that makes Spot capacity reliable enough to depend on.
Scenario Question Drills
Q1. A company runs a steady-state, predictable production fleet but wants maximum flexibility to change instance families and regions over the commitment period.
Q2. A regulated workload must be guaranteed to launch in a specific Availability Zone during a regional capacity crunch, and the team also wants a discount. What should they purchase?
Q3. A batch processing job runs for 40 minutes per task and cannot checkpoint. The team wants to cut its cost by running it on Spot. What is the correct assessment?
Q4. A company's monthly bill shows a Savings Plan charge that is no longer offset by matching usage after the team migrated to Graviton instances. What happened, and what is the first diagnostic step?
Q5. A mixed-instance Auto Scaling Group repeatedly fails to reach desired capacity during peak demand, reporting insufficient capacity. What is the most effective fix?
Q6. A team wants the deepest possible discount on a fleet whose instance family and region are frozen for the next three years, and they do not need a capacity reservation. What should they buy?
Q7. A workload's write path is not idempotent, and it has been running on Spot. The team sees intermittent duplicate records correlated with interruption events. What is the correct remediation?
Q8. A company has a Savings Plan that is consistently under-utilized. What is the accurate characterization of the cost impact?
Q9. A team expects to change its instance family mid-term but still wants a Reserved Instance discount. Which offering class fits?
Q10. A workload must survive a Spot interruption without losing in-memory state, and the application cannot be modified. Which Spot request configuration is the closest fit?
Q11. Which metric pair should be on a cost dashboard to detect a commitment that has drifted out of alignment with the fleet?
Q12. A company has a Standard RI it no longer needs after re-platforming. What is the recovery option?
Q13. A service's SLO is met only when Spot burst capacity is available. How should this be modeled honestly?
Q14. A fleet shifts heavily toward Spot over a quarter. What second-order effect should the team check on their existing Savings Plan?
Q15. A team is designing a cost-optimized fleet for a fault-tolerant, checkpointable workload with a stable baseline and unpredictable peaks. Which design is correct?
Peek into Tomorrow
Everything here assumed you already knew which instances were oversized. The commitment instruments optimize the price of the capacity you have; none of them tells you whether that capacity is the right shape in the first place. A fleet of consistently under-utilized instances covered by a perfectly sized Savings Plan is still paying for compute it does not need, and the discount simply makes the waste slightly cheaper. The open question is how you find the over-provisioned instances in the first place, and how you decide what to do with them.
Tomorrow's material answers that with three different lenses on the same problem. Compute Optimizer turns utilization metrics into concrete right-sizing recommendations, Cost Explorer visualizes and forecasts spend and drives RI and Savings Plan purchase advice, and Trusted Advisor flags cost, security, performance, and fault-tolerance issues against AWS best practices. The interesting part is that they disagree in useful ways — a right-sizing recommendation and a commitment purchase recommendation can point in opposite directions, and knowing which to act on first is the judgment the exam is testing.
Sources
- AWS Savings Plans User Guide — What Are Savings Plans?
- Amazon EC2 User Guide — Reserved Instances
- Amazon EC2 User Guide — Spot Instances
- Amazon EC2 User Guide — Spot Instance Interruptions
- Amazon EC2 Auto Scaling — Mixed Instances Policies
- AWS Billing — Reserved Instance Marketplace
- AWS Well-Architected Framework — Cost Optimization Pillar
- AWS Well-Architected Framework — Reliability Pillar (service quotas)