Week Synthesis — Compute Decision Matrix & Scenario Drills
Week 3 Recap
Week 3 opened with Fargate's task definition and its awsvpc network mode, where every task gets its own ENI and security group instead of sharing a host's port space, and then moved to EKS, where the same serverless idea reappears as Fargate profiles that lose DaemonSets and cost more per pod than managed node groups. From there the week turned to capacity mechanics: lifecycle hooks that pause an instance in Pending:Wait or Terminating:Wait so bootstrap and drain scripts can run, and warm pools that keep pre-initialized stopped instances ready to promote instead of booting cold. Day 18 added ALB weighted target groups for blue/green and canary shifting, Day 19 added reserved and provisioned concurrency as the two independent levers on Lambda scaling, and Day 20 closed with Step Functions' Standard versus Express split and API Gateway integrations that bypass Lambda entirely.
What ties these together is that none of them is a compute platform decision in isolation. Each one is a trade of operational overhead against cost against control, and the exam almost never asks which service is "best" — it asks which service satisfies a stated constraint. This synthesis day is about holding all four services in one frame at once.
Foundations You'll Need Today
Containers and Images
A container is a way of packaging an application together with everything it needs to run — its code, its libraries, its configuration — into a single unit that behaves the same way wherever it is started. The packaged artifact is called an image, and when you actually run an image you get a container. The problem this solves is the classic "it worked on my laptop" failure: instead of installing dependencies by hand on each server and hoping the versions match, you build the image once and run that exact image everywhere. Containers are lighter than virtual machines because they share the host's operating system kernel rather than each booting a full operating system of their own, which is why you can pack many more of them onto the same hardware.
Orchestration
Running one container is easy; running hundreds of them across many machines, restarting the ones that crash, replacing them when you deploy a new version, and connecting them to a network and a load balancer is not. An orchestrator is the software that does that bookkeeping for you. It takes a description of what you want — "keep five copies of this image running, give each one a network address, restart any that fail" — and continuously works to make reality match that description. ECS and EKS are AWS's two orchestrators, and the difference between them is mostly about which description format and ecosystem they speak, not about whether they can run containers.
Serverless
Serverless does not mean there are no servers; it means you never see or manage them. With a traditional server you rent a machine by the hour whether or not it is doing work, and you are responsible for patching its operating system and keeping it running. With a serverless service like Lambda you hand over a piece of code, and the provider runs it only when something triggers it, bills you for the milliseconds it actually executes, and handles all the underlying machines. The trade is control: you give up the ability to install things at the operating system level or keep a process running indefinitely, and in exchange you stop paying for idle time and stop doing capacity planning.
The Three Pricing Models You'll See Named
Almost every cost question in this curriculum comes down to choosing among three ways to pay for compute. On-Demand means paying the full hourly rate with no commitment, which is the most flexible and the most expensive per hour. Commitments — Reserved Instances and Savings Plans — mean promising to spend a certain amount for one or three years in exchange for a meaningful discount, which only makes sense for a workload you know will be running the whole time. Spot means bidding on spare AWS capacity at a steep discount, with the catch that AWS can take the capacity back on short notice when it needs it, so it only suits work that can be interrupted and restarted. The skill being tested is not arithmetic but matching the model to the shape of the traffic.
With that grounding, here's why this day exists: the four services covered this week are not four interchangeable ways to run code, and the exam is really asking you to eliminate the ones that violate a stated constraint before you ever compare prices.
The Frame: Three Constraints, Not Four Services
Every compute question on SAP-C02 resolves to a ranking of three constraints: how much operational overhead the team is willing to absorb, how much the workload's cost profile matters relative to its traffic shape, and how much control over the runtime the workload actually requires. The four services in this week are not four answers to one question — they are four positions on those three axes, and the exam's job is to give you a scenario where exactly one position is defensible.
The reason this framing matters is that candidates tend to memorize service features and then pattern-match on keywords. A scenario mentions Kubernetes and the answer becomes EKS; a scenario mentions unpredictable traffic and the answer becomes Lambda. That works until the scenario mentions Kubernetes and a hard cold-start budget, or unpredictable traffic and a 15-minute execution ceiling, at which point the keyword match produces the wrong answer. The reliable approach is to extract the constraints first, then eliminate services that violate one of them, then choose among the survivors on cost.
Operational overhead is usually the first eliminator because it is stated most explicitly. A scenario that says the team has no container or Kubernetes experience, or that the company wants to minimize the number of services it operates, is telling you to eliminate EKS and probably ECS-on-EC2 as well. Cost is the second eliminator, and it is almost always stated as a shape rather than a number: steady and predictable favors commitments, spiky and unpredictable favors per-request billing, and fault-tolerant batch favors Spot. Control is the last eliminator and the one candidates most often miss, because it is usually implied by a requirement rather than named — a need for a specific kernel module, a custom AMI, a GPU driver version, or a persistent local scratch disk all eliminate the serverless options without ever using the word "control."
| Constraint stated in the scenario | Eliminates | Survivors |
|---|---|---|
| No container or Kubernetes operational experience | EKS, self-managed ECS on EC2 | Lambda, Fargate, managed node groups with a platform team |
| Steady, predictable, 24/7 baseline load | Pure on-demand pricing | Savings Plans, RIs, provisioned capacity |
| Fault-tolerant, interruptible batch work | On-Demand-only fleets | Spot, Spot Fleet, mixed-instance ASGs |
| Custom kernel module, GPU driver, or persistent local disk | Lambda, Fargate | EC2, ECS on EC2, EKS managed node groups |
| Sub-second cold-start requirement on spiky traffic | Lambda without provisioned concurrency | Lambda with provisioned concurrency, always-on containers |
| Run longer than 15 minutes per invocation | Lambda | ECS, EKS, EC2, Step Functions orchestrating shorter steps |
Orchestration: ECS, EKS, and When Neither Fits
The ECS-versus-EKS question is the one candidates most often get wrong for the wrong reason. Both run containers, both can run on Fargate or on EC2 capacity, and both integrate with the same load balancers, IAM roles, and CloudWatch surfaces. The difference that actually decides exam scenarios is not capability — it is what the organization already knows and what it needs to be portable to. EKS runs upstream Kubernetes, which means existing manifests, Helm charts, operators, and admission controllers transfer unchanged, and a workload built for EKS can move to another conformant Kubernetes distribution with limited rework. ECS is AWS's own orchestrator: simpler to operate, fewer moving parts, no control-plane upgrade cycle to manage, but the task definitions and service definitions are AWS-specific and do not port anywhere.
The second deciding factor is who owns the control plane. EKS charges a per-cluster hourly fee for the managed control plane regardless of how much work the cluster does, which makes many small clusters expensive and pushes teams toward fewer, larger clusters with namespace-based tenancy. ECS has no equivalent control-plane charge, so a fleet of small, isolated services costs nothing extra to orchestrate. A scenario describing dozens of small independent services with no Kubernetes requirement is usually pointing at ECS, while a scenario describing a platform team standardizing on Kubernetes across clouds is pointing at EKS even if ECS would be cheaper.
There is also a third position that candidates forget: sometimes the right answer is neither orchestrator. A single long-running process that needs a specific AMI, a fixed IP, or a persistent local disk is an EC2 instance with an Auto Scaling Group, not a container platform. The exam will occasionally present a scenario with all the trappings of a container question — a Dockerfile, a registry, a service that needs to scale — where the actual constraint (a licensed kernel module, a 500 GB local scratch volume) makes plain EC2 the only correct answer.
| Dimension | ECS | EKS | Plain EC2 + ASG |
|---|---|---|---|
| Orchestration model | AWS-proprietary task/service definitions | Upstream Kubernetes API | None — you manage processes |
| Control-plane cost | None | Per-cluster hourly charge | None |
| Portability | AWS-only | Any conformant Kubernetes | Anywhere, but no orchestration |
| Operational burden | Low | Moderate to high (upgrades, add-ons, RBAC) | Highest (patching, AMIs, scaling logic) |
| Best fit | AWS-native microservices, small teams | Existing Kubernetes investment, multi-cloud | Custom runtime, licensed software, local disk |
Capacity: Fargate, Managed Nodes, and the Scaling Levers
Once the orchestrator is chosen, the next fork is where the compute comes from. Fargate removes the node layer entirely: there is no AMI to patch, no cluster autoscaler to tune, no bin-packing to reason about, and each task or pod gets its own isolated compute allocation. That isolation is also the limitation. Because there is no shared node, DaemonSets cannot run, host-level agents cannot be installed, and per-pod pricing is higher than the equivalent share of an EC2 instance. Managed node groups sit in the middle: AWS owns the Auto Scaling Group and the node lifecycle, you still choose instance types and can run DaemonSets, and you pay EC2 rates. Self-managed nodes give the most control and the most work.
The scaling levers differ by service in ways the exam tests directly. EC2 Auto Scaling scales on metrics, schedules, or predictive policies, and its two most-tested features are lifecycle hooks and warm pools. A lifecycle hook pauses an instance in Pending:Wait before it enters service or Terminating:Wait before it is destroyed, which is where bootstrap scripts, configuration management runs, and load-balancer deregistration belong. A warm pool keeps pre-initialized instances — usually stopped — ready to be promoted into the group, which converts a multi-minute cold boot into a seconds-long promotion. Lambda scales differently: reserved concurrency caps and guarantees a function's maximum simultaneous executions, while provisioned concurrency pre-warms execution environments so that the first request after a scale-out does not pay the initialization cost. These are independent settings, and a scenario that needs both a hard ceiling on downstream connections and no cold starts requires both.
Container services add a third layer on top of the compute layer: task or pod count. An ECS service with a target-tracking policy on CPU or ALB request count adjusts desired count, and the underlying capacity provider decides whether that means launching more Fargate tasks or adding EC2 instances to a node group. The exam's favorite trap here is a scenario where the service-level scaling policy is correct but the capacity provider has no headroom, so tasks sit in a pending state — the fix is capacity provider managed scaling or a larger maximum size on the node group, not a change to the service policy.
| Scaling lever | Applies to | What it actually controls |
|---|---|---|
| Target-tracking / step scaling policy | EC2 ASG, ECS service, EKS HPA | Desired instance, task, or pod count |
| Lifecycle hook | EC2 ASG | Pauses launch or termination for scripts to run |
| Warm pool | EC2 ASG | Pre-initialized instances promoted instead of booted |
| Reserved concurrency | Lambda | Hard cap and guarantee on simultaneous executions |
| Provisioned concurrency | Lambda | Pre-warmed environments, removes cold starts |
| Capacity provider managed scaling | ECS on EC2 | Node count behind the task count |
Serverless Boundaries: Where Lambda Stops Being the Answer
Lambda is the default answer for event-driven, short-lived, spiky work, and the exam rewards recognizing when that default breaks. The hard boundaries are execution duration, payload size, and runtime control. A function cannot run indefinitely, so any workload with a long-running process — a video transcode that takes an hour, a long-lived WebSocket connection, a stateful stream processor — has to move to containers or EC2. A function also cannot host a persistent local filesystem or install arbitrary system packages at the kernel level, which eliminates anything needing a custom driver or a licensed daemon.
The softer boundary is cost shape. Lambda's per-request and per-GB-second billing is excellent for spiky traffic and poor for a steady, high-volume baseline, where a continuously running container or a committed EC2 instance is cheaper per unit of work. A scenario that describes a constant, predictable request rate with no idle periods is often pointing away from Lambda even though Lambda would technically work. The reverse is also tested: a scenario with long idle periods punctuated by bursts is almost always Lambda, because paying for idle capacity is the thing being avoided.
The third boundary is orchestration. Lambda functions are individually simple, and complexity moves into whatever coordinates them. Step Functions handles that with explicit state, retries, and catch blocks, and its Standard versus Express split matters here: Standard workflows are exactly-once, long-running, and auditable, while Express workflows are at-least-once, high-throughput, and cheaper per execution. A scenario describing millions of short executions per day that can tolerate at-least-once semantics is an Express scenario; a scenario describing a multi-step order process with a human approval step and an audit requirement is a Standard scenario. API Gateway sits in front of both and can integrate directly with AWS services, bypassing Lambda entirely for simple request transformations — a detail the exam uses to test whether you reach for a function when none is needed.
Cost Shape: Matching the Pricing Model to the Traffic
Cost questions in this week are almost never arithmetic. They are about matching a pricing model to a traffic shape, and the four shapes the exam uses are steady baseline, predictable peak, spiky and unpredictable, and interruptible batch. A steady baseline wants a commitment — a Compute Savings Plan for maximum flexibility across families and regions, or an EC2 Instance Savings Plan for a deeper discount locked to one family and region. A predictable peak wants a committed baseline plus on-demand or Spot burst capacity above it. Spiky and unpredictable traffic wants per-request billing, which means Lambda or Fargate rather than a fleet sized for the peak. Interruptible batch wants Spot, with a mixed-instance Auto Scaling Group or Spot Fleet to diversify capacity pools and a two-minute interruption notice handler to checkpoint work.
The trap is that these shapes are often described indirectly. "The workload runs continuously but utilization varies between 20% and 90% across the day" is a steady baseline with a peak, not a spiky workload — the instance is running either way, so a commitment on the baseline is correct and the peak is handled by scaling. "The workload runs for ten minutes every hour" is spiky, and paying for a full hour of instance time to serve ten minutes of work is the waste the scenario is describing. Reading the shape correctly is most of the answer.
| Traffic shape | Pricing model | Typical service pairing |
|---|---|---|
| Steady 24/7 baseline | Compute or EC2 Instance Savings Plan, RIs | EC2 ASG, ECS on EC2, EKS managed nodes |
| Predictable peak above a baseline | Commitment on baseline + on-demand burst | ASG with mixed instance policy |
| Spiky, unpredictable, idle between bursts | Per-request / per-GB-second | Lambda, Fargate, API Gateway |
| Interruptible batch | Spot with diversification | Spot Fleet, mixed-instance ASG, AWS Batch |
Hands-on Lab: Build the Decision Matrix Against Real Constraints
This lab is a paper exercise with a verification step, and it is designed to be done in about 45 minutes. The goal is to stop pattern-matching on service names and start eliminating on constraints, so the output is a written decision record rather than a deployed stack.
Start by writing out eight short scenarios of your own, each two or three sentences, drawn from systems you have actually operated. Give each one a stated constraint that eliminates at least one of the four services in this week. Good constraint types to use: a team with no Kubernetes experience, a workload that must run for 40 minutes per invocation, a service that needs a 200 GB local scratch volume, a request rate that is flat at 3,000 requests per second around the clock, a batch job that can be restarted from a checkpoint, a workload with a hard 200 ms p99 including cold start, a service that must run a licensed host-level agent, and a workload with long idle periods between bursts.
For each scenario, write three lines: the constraint you extracted, the services it eliminates and why, and the service you would choose with the pricing model that matches its traffic shape. Do not write more than three lines per scenario — the discipline of compression is the point, because the exam gives you roughly two and a half minutes per question and the elimination has to happen fast.
Then verify the mechanics you are reasoning about rather than trusting memory. In a sandbox account, create a Lambda function with reserved concurrency set to a low value and confirm that concurrent invocations beyond that value are throttled rather than queued indefinitely. Create an EC2 Auto Scaling Group with a lifecycle hook on termination and confirm the instance sits in Terminating:Wait until you complete the hook. If you have an ECS cluster, set a service's desired count above what the capacity provider can place and observe tasks sitting in a pending state — that is the failure mode the exam describes when it says scaling is configured but capacity is not.
Finish by writing a one-page cheat sheet with four columns: service, the constraint that eliminates it, the constraint that selects it, and the pricing model that pairs with it. Keep it to one page. If it does not fit on one page, you are still memorizing features instead of constraints, and the sheet will not be usable under exam time pressure.
Mixed Scenario Quiz — 25 Questions Across the Week
Q1. A platform team already runs Kubernetes on-premises and in another cloud, and wants to move a set of services to AWS with minimal manifest changes. Which compute platform fits best?
Q2. A workload requires a DaemonSet that runs a host-level log shipper on every node. Which EKS compute option cannot satisfy this?
Q3. An application takes four minutes to bootstrap before it can serve traffic, and scale-out currently lags demand spikes badly. What reduces the scale-out latency?
Q4. During scale-in, an application must finish in-flight requests and upload final logs before the instance is destroyed. What implements this?
Q5. A canary release needs 5% of production traffic sent to a new task set before full cutover. Which ALB feature does this?
Q6. A Lambda function reading from Kinesis is exhausting connections on a downstream RDS instance during traffic spikes. What bounds this without changing the database?
Q7. A latency-sensitive function must never pay a cold-start penalty, and the team also needs a hard ceiling on its concurrency. What configuration satisfies both?
Q8. An IoT pipeline runs millions of short workflow executions per day and can tolerate at-least-once semantics. Which Step Functions type fits?
Q9. An order process requires a human approval step, must run for up to three days, and must be auditable with exactly-once execution. Which Step Functions type fits?
Q10. A simple API endpoint only needs to write a request body directly into a DynamoDB table with no transformation logic. What avoids an unnecessary function?
Q11. A workload runs a licensed host-level security agent that cannot be containerized. Which compute option is viable?
Q12. A batch job runs for 40 minutes per invocation and can restart from a checkpoint. Which platform and pricing model fit?
Q13. A service runs continuously at a flat 3,000 requests per second around the clock with no idle periods. Which pricing approach is most cost-effective?
Q14. A workload runs for ten minutes every hour and is idle the rest of the time. What is the cost-optimal model?
Q15. An ECS service's target-tracking policy is raising desired count, but new tasks sit in a pending state and never start. What is the most likely cause?
Q16. A team wants each container task to have its own security group and IP address rather than sharing a host's port space. Which network mode provides this?
Q17. A company runs dozens of small independent containerized services and has no Kubernetes requirement or existing Kubernetes tooling. Which orchestrator minimizes operational overhead?
Q18. A workload needs a 500 GB local scratch volume for temporary processing files and runs for several hours at a time. Which platform fits?
Q19. A team wants to shift 10% of traffic to a new container version, watch error rates for an hour, then shift the rest. Which combination supports this?
Q20. A function processes SQS messages in batches, and one bad message in a batch causes the entire batch to be retried repeatedly. What addresses this?
Q21. A team needs a hard guarantee that a function can always run at least 50 concurrent executions even when other functions in the account are consuming the regional limit. What provides this?
Q22. A workload's traffic is unpredictable, bursts to thousands of requests per second, and drops to zero for hours. The team wants no idle cost. Which pairing fits?
Q23. A team wants to run a containerized web service without managing any EC2 instances, but the service needs a sidecar container that must run on every node. What is the constraint conflict?
Q24. A company wants maximum flexibility to change instance families and regions over a one-year commitment while still receiving a discount on a steady baseline. What should they purchase?
Q25. A batch workload is fault-tolerant and can be restarted, and the team wants the lowest possible compute cost with protection against a single capacity pool being unavailable. What should they configure?
Preview
Everything in this week assumed that the compute layer is the hard part — that once you have chosen a platform and a scaling model, the data layer is a solved problem you configure once. That assumption holds right up until the workload has to survive the loss of an Availability Zone, or serve reads from a second continent, or accept writes in two regions at the same time. The open question this week leaves unresolved is what happens to the database when the compute layer is already multi-AZ and the storage underneath it is not.
Day 22 answers that for the relational case with Aurora, whose storage layer is replicated six ways across three Availability Zones and self-heals without a failover event, and then extends it across regions with Aurora Global Database — replication that happens in the storage layer rather than through binlog shipping, which is why the cross-region lag is measured in sub-second terms and a secondary region can be promoted in under a minute. The interesting part is not the numbers but the mechanism: understanding why storage-layer replication behaves differently from a read replica is what separates the two answers on an exam question that mentions both.