Day 21 of 70 · Week 3
Day 21 / 70 Week 3 of 14 Phase 2: Compute, Containers & Global Databases

Week Synthesis — Compute Decision Matrix & Scenario Drills

🕑 ~32 min read · 4 services covered
ECS EKS Lambda EC2 Auto Scaling

Week 3 Recap

Week 3 opened with Fargate's task definition and its awsvpc network mode, where every task gets its own ENI and security group instead of sharing a host's port space, and then moved to EKS, where the same serverless idea reappears as Fargate profiles that lose DaemonSets and cost more per pod than managed node groups. From there the week turned to capacity mechanics: lifecycle hooks that pause an instance in Pending:Wait or Terminating:Wait so bootstrap and drain scripts can run, and warm pools that keep pre-initialized stopped instances ready to promote instead of booting cold. Day 18 added ALB weighted target groups for blue/green and canary shifting, Day 19 added reserved and provisioned concurrency as the two independent levers on Lambda scaling, and Day 20 closed with Step Functions' Standard versus Express split and API Gateway integrations that bypass Lambda entirely.

What ties these together is that none of them is a compute platform decision in isolation. Each one is a trade of operational overhead against cost against control, and the exam almost never asks which service is "best" — it asks which service satisfies a stated constraint. This synthesis day is about holding all four services in one frame at once.

Foundations You'll Need Today

Containers and Images

A container is a way of packaging an application together with everything it needs to run — its code, its libraries, its configuration — into a single unit that behaves the same way wherever it is started. The packaged artifact is called an image, and when you actually run an image you get a container. The problem this solves is the classic "it worked on my laptop" failure: instead of installing dependencies by hand on each server and hoping the versions match, you build the image once and run that exact image everywhere. Containers are lighter than virtual machines because they share the host's operating system kernel rather than each booting a full operating system of their own, which is why you can pack many more of them onto the same hardware.

Orchestration

Running one container is easy; running hundreds of them across many machines, restarting the ones that crash, replacing them when you deploy a new version, and connecting them to a network and a load balancer is not. An orchestrator is the software that does that bookkeeping for you. It takes a description of what you want — "keep five copies of this image running, give each one a network address, restart any that fail" — and continuously works to make reality match that description. ECS and EKS are AWS's two orchestrators, and the difference between them is mostly about which description format and ecosystem they speak, not about whether they can run containers.

Serverless

Serverless does not mean there are no servers; it means you never see or manage them. With a traditional server you rent a machine by the hour whether or not it is doing work, and you are responsible for patching its operating system and keeping it running. With a serverless service like Lambda you hand over a piece of code, and the provider runs it only when something triggers it, bills you for the milliseconds it actually executes, and handles all the underlying machines. The trade is control: you give up the ability to install things at the operating system level or keep a process running indefinitely, and in exchange you stop paying for idle time and stop doing capacity planning.

The Three Pricing Models You'll See Named

Almost every cost question in this curriculum comes down to choosing among three ways to pay for compute. On-Demand means paying the full hourly rate with no commitment, which is the most flexible and the most expensive per hour. Commitments — Reserved Instances and Savings Plans — mean promising to spend a certain amount for one or three years in exchange for a meaningful discount, which only makes sense for a workload you know will be running the whole time. Spot means bidding on spare AWS capacity at a steep discount, with the catch that AWS can take the capacity back on short notice when it needs it, so it only suits work that can be interrupted and restarted. The skill being tested is not arithmetic but matching the model to the shape of the traffic.

With that grounding, here's why this day exists: the four services covered this week are not four interchangeable ways to run code, and the exam is really asking you to eliminate the ones that violate a stated constraint before you ever compare prices.

The Frame: Three Constraints, Not Four Services

Every compute question on SAP-C02 resolves to a ranking of three constraints: how much operational overhead the team is willing to absorb, how much the workload's cost profile matters relative to its traffic shape, and how much control over the runtime the workload actually requires. The four services in this week are not four answers to one question — they are four positions on those three axes, and the exam's job is to give you a scenario where exactly one position is defensible.

The reason this framing matters is that candidates tend to memorize service features and then pattern-match on keywords. A scenario mentions Kubernetes and the answer becomes EKS; a scenario mentions unpredictable traffic and the answer becomes Lambda. That works until the scenario mentions Kubernetes and a hard cold-start budget, or unpredictable traffic and a 15-minute execution ceiling, at which point the keyword match produces the wrong answer. The reliable approach is to extract the constraints first, then eliminate services that violate one of them, then choose among the survivors on cost.

Operational overhead is usually the first eliminator because it is stated most explicitly. A scenario that says the team has no container or Kubernetes experience, or that the company wants to minimize the number of services it operates, is telling you to eliminate EKS and probably ECS-on-EC2 as well. Cost is the second eliminator, and it is almost always stated as a shape rather than a number: steady and predictable favors commitments, spiky and unpredictable favors per-request billing, and fault-tolerant batch favors Spot. Control is the last eliminator and the one candidates most often miss, because it is usually implied by a requirement rather than named — a need for a specific kernel module, a custom AMI, a GPU driver version, or a persistent local scratch disk all eliminate the serverless options without ever using the word "control."

Constraint stated in the scenarioEliminatesSurvivors
No container or Kubernetes operational experienceEKS, self-managed ECS on EC2Lambda, Fargate, managed node groups with a platform team
Steady, predictable, 24/7 baseline loadPure on-demand pricingSavings Plans, RIs, provisioned capacity
Fault-tolerant, interruptible batch workOn-Demand-only fleetsSpot, Spot Fleet, mixed-instance ASGs
Custom kernel module, GPU driver, or persistent local diskLambda, FargateEC2, ECS on EC2, EKS managed node groups
Sub-second cold-start requirement on spiky trafficLambda without provisioned concurrencyLambda with provisioned concurrency, always-on containers
Run longer than 15 minutes per invocationLambdaECS, EKS, EC2, Step Functions orchestrating shorter steps

Orchestration: ECS, EKS, and When Neither Fits

The ECS-versus-EKS question is the one candidates most often get wrong for the wrong reason. Both run containers, both can run on Fargate or on EC2 capacity, and both integrate with the same load balancers, IAM roles, and CloudWatch surfaces. The difference that actually decides exam scenarios is not capability — it is what the organization already knows and what it needs to be portable to. EKS runs upstream Kubernetes, which means existing manifests, Helm charts, operators, and admission controllers transfer unchanged, and a workload built for EKS can move to another conformant Kubernetes distribution with limited rework. ECS is AWS's own orchestrator: simpler to operate, fewer moving parts, no control-plane upgrade cycle to manage, but the task definitions and service definitions are AWS-specific and do not port anywhere.

The second deciding factor is who owns the control plane. EKS charges a per-cluster hourly fee for the managed control plane regardless of how much work the cluster does, which makes many small clusters expensive and pushes teams toward fewer, larger clusters with namespace-based tenancy. ECS has no equivalent control-plane charge, so a fleet of small, isolated services costs nothing extra to orchestrate. A scenario describing dozens of small independent services with no Kubernetes requirement is usually pointing at ECS, while a scenario describing a platform team standardizing on Kubernetes across clouds is pointing at EKS even if ECS would be cheaper.

There is also a third position that candidates forget: sometimes the right answer is neither orchestrator. A single long-running process that needs a specific AMI, a fixed IP, or a persistent local disk is an EC2 instance with an Auto Scaling Group, not a container platform. The exam will occasionally present a scenario with all the trappings of a container question — a Dockerfile, a registry, a service that needs to scale — where the actual constraint (a licensed kernel module, a 500 GB local scratch volume) makes plain EC2 the only correct answer.

DimensionECSEKSPlain EC2 + ASG
Orchestration modelAWS-proprietary task/service definitionsUpstream Kubernetes APINone — you manage processes
Control-plane costNonePer-cluster hourly chargeNone
PortabilityAWS-onlyAny conformant KubernetesAnywhere, but no orchestration
Operational burdenLowModerate to high (upgrades, add-ons, RBAC)Highest (patching, AMIs, scaling logic)
Best fitAWS-native microservices, small teamsExisting Kubernetes investment, multi-cloudCustom runtime, licensed software, local disk

Capacity: Fargate, Managed Nodes, and the Scaling Levers

Once the orchestrator is chosen, the next fork is where the compute comes from. Fargate removes the node layer entirely: there is no AMI to patch, no cluster autoscaler to tune, no bin-packing to reason about, and each task or pod gets its own isolated compute allocation. That isolation is also the limitation. Because there is no shared node, DaemonSets cannot run, host-level agents cannot be installed, and per-pod pricing is higher than the equivalent share of an EC2 instance. Managed node groups sit in the middle: AWS owns the Auto Scaling Group and the node lifecycle, you still choose instance types and can run DaemonSets, and you pay EC2 rates. Self-managed nodes give the most control and the most work.

The scaling levers differ by service in ways the exam tests directly. EC2 Auto Scaling scales on metrics, schedules, or predictive policies, and its two most-tested features are lifecycle hooks and warm pools. A lifecycle hook pauses an instance in Pending:Wait before it enters service or Terminating:Wait before it is destroyed, which is where bootstrap scripts, configuration management runs, and load-balancer deregistration belong. A warm pool keeps pre-initialized instances — usually stopped — ready to be promoted into the group, which converts a multi-minute cold boot into a seconds-long promotion. Lambda scales differently: reserved concurrency caps and guarantees a function's maximum simultaneous executions, while provisioned concurrency pre-warms execution environments so that the first request after a scale-out does not pay the initialization cost. These are independent settings, and a scenario that needs both a hard ceiling on downstream connections and no cold starts requires both.

Container services add a third layer on top of the compute layer: task or pod count. An ECS service with a target-tracking policy on CPU or ALB request count adjusts desired count, and the underlying capacity provider decides whether that means launching more Fargate tasks or adding EC2 instances to a node group. The exam's favorite trap here is a scenario where the service-level scaling policy is correct but the capacity provider has no headroom, so tasks sit in a pending state — the fix is capacity provider managed scaling or a larger maximum size on the node group, not a change to the service policy.

Scaling leverApplies toWhat it actually controls
Target-tracking / step scaling policyEC2 ASG, ECS service, EKS HPADesired instance, task, or pod count
Lifecycle hookEC2 ASGPauses launch or termination for scripts to run
Warm poolEC2 ASGPre-initialized instances promoted instead of booted
Reserved concurrencyLambdaHard cap and guarantee on simultaneous executions
Provisioned concurrencyLambdaPre-warmed environments, removes cold starts
Capacity provider managed scalingECS on EC2Node count behind the task count

Serverless Boundaries: Where Lambda Stops Being the Answer

Lambda is the default answer for event-driven, short-lived, spiky work, and the exam rewards recognizing when that default breaks. The hard boundaries are execution duration, payload size, and runtime control. A function cannot run indefinitely, so any workload with a long-running process — a video transcode that takes an hour, a long-lived WebSocket connection, a stateful stream processor — has to move to containers or EC2. A function also cannot host a persistent local filesystem or install arbitrary system packages at the kernel level, which eliminates anything needing a custom driver or a licensed daemon.

The softer boundary is cost shape. Lambda's per-request and per-GB-second billing is excellent for spiky traffic and poor for a steady, high-volume baseline, where a continuously running container or a committed EC2 instance is cheaper per unit of work. A scenario that describes a constant, predictable request rate with no idle periods is often pointing away from Lambda even though Lambda would technically work. The reverse is also tested: a scenario with long idle periods punctuated by bursts is almost always Lambda, because paying for idle capacity is the thing being avoided.

The third boundary is orchestration. Lambda functions are individually simple, and complexity moves into whatever coordinates them. Step Functions handles that with explicit state, retries, and catch blocks, and its Standard versus Express split matters here: Standard workflows are exactly-once, long-running, and auditable, while Express workflows are at-least-once, high-throughput, and cheaper per execution. A scenario describing millions of short executions per day that can tolerate at-least-once semantics is an Express scenario; a scenario describing a multi-step order process with a human approval step and an audit requirement is a Standard scenario. API Gateway sits in front of both and can integrate directly with AWS services, bypassing Lambda entirely for simple request transformations — a detail the exam uses to test whether you reach for a function when none is needed.

Cost Shape: Matching the Pricing Model to the Traffic

Cost questions in this week are almost never arithmetic. They are about matching a pricing model to a traffic shape, and the four shapes the exam uses are steady baseline, predictable peak, spiky and unpredictable, and interruptible batch. A steady baseline wants a commitment — a Compute Savings Plan for maximum flexibility across families and regions, or an EC2 Instance Savings Plan for a deeper discount locked to one family and region. A predictable peak wants a committed baseline plus on-demand or Spot burst capacity above it. Spiky and unpredictable traffic wants per-request billing, which means Lambda or Fargate rather than a fleet sized for the peak. Interruptible batch wants Spot, with a mixed-instance Auto Scaling Group or Spot Fleet to diversify capacity pools and a two-minute interruption notice handler to checkpoint work.

The trap is that these shapes are often described indirectly. "The workload runs continuously but utilization varies between 20% and 90% across the day" is a steady baseline with a peak, not a spiky workload — the instance is running either way, so a commitment on the baseline is correct and the peak is handled by scaling. "The workload runs for ten minutes every hour" is spiky, and paying for a full hour of instance time to serve ten minutes of work is the waste the scenario is describing. Reading the shape correctly is most of the answer.

Traffic shapePricing modelTypical service pairing
Steady 24/7 baselineCompute or EC2 Instance Savings Plan, RIsEC2 ASG, ECS on EC2, EKS managed nodes
Predictable peak above a baselineCommitment on baseline + on-demand burstASG with mixed instance policy
Spiky, unpredictable, idle between burstsPer-request / per-GB-secondLambda, Fargate, API Gateway
Interruptible batchSpot with diversificationSpot Fleet, mixed-instance ASG, AWS Batch

Hands-on Lab: Build the Decision Matrix Against Real Constraints

This lab is a paper exercise with a verification step, and it is designed to be done in about 45 minutes. The goal is to stop pattern-matching on service names and start eliminating on constraints, so the output is a written decision record rather than a deployed stack.

Start by writing out eight short scenarios of your own, each two or three sentences, drawn from systems you have actually operated. Give each one a stated constraint that eliminates at least one of the four services in this week. Good constraint types to use: a team with no Kubernetes experience, a workload that must run for 40 minutes per invocation, a service that needs a 200 GB local scratch volume, a request rate that is flat at 3,000 requests per second around the clock, a batch job that can be restarted from a checkpoint, a workload with a hard 200 ms p99 including cold start, a service that must run a licensed host-level agent, and a workload with long idle periods between bursts.

For each scenario, write three lines: the constraint you extracted, the services it eliminates and why, and the service you would choose with the pricing model that matches its traffic shape. Do not write more than three lines per scenario — the discipline of compression is the point, because the exam gives you roughly two and a half minutes per question and the elimination has to happen fast.

Then verify the mechanics you are reasoning about rather than trusting memory. In a sandbox account, create a Lambda function with reserved concurrency set to a low value and confirm that concurrent invocations beyond that value are throttled rather than queued indefinitely. Create an EC2 Auto Scaling Group with a lifecycle hook on termination and confirm the instance sits in Terminating:Wait until you complete the hook. If you have an ECS cluster, set a service's desired count above what the capacity provider can place and observe tasks sitting in a pending state — that is the failure mode the exam describes when it says scaling is configured but capacity is not.

Finish by writing a one-page cheat sheet with four columns: service, the constraint that eliminates it, the constraint that selects it, and the pricing model that pairs with it. Keep it to one page. If it does not fit on one page, you are still memorizing features instead of constraints, and the sheet will not be usable under exam time pressure.

Mixed Scenario Quiz — 25 Questions Across the Week

Q1. A platform team already runs Kubernetes on-premises and in another cloud, and wants to move a set of services to AWS with minimal manifest changes. Which compute platform fits best?

A. Amazon ECS with the EC2 launch type
B. Amazon EKS
C. AWS Lambda behind API Gateway
D. EC2 instances in an Auto Scaling Group
Correct answer: B. EKS runs upstream Kubernetes, so existing manifests, Helm charts, and operators transfer with minimal rework, and the workload stays portable to other conformant distributions.

Q2. A workload requires a DaemonSet that runs a host-level log shipper on every node. Which EKS compute option cannot satisfy this?

A. Managed node groups
B. Self-managed EC2 node groups
C. EKS Fargate profiles
D. A mixed node group with Spot capacity
Correct answer: C. Fargate pods run in isolated micro-VMs with no shared node, so DaemonSets — which require a persistent per-node agent — are not supported.

Q3. An application takes four minutes to bootstrap before it can serve traffic, and scale-out currently lags demand spikes badly. What reduces the scale-out latency?

A. Increase the Auto Scaling Group cooldown period
B. Use a warm pool of pre-initialized stopped instances
C. Switch the fleet to Spot Instances
D. Add a lifecycle hook on instance launch
Correct answer: B. A warm pool keeps instances pre-initialized so the group promotes them in seconds instead of re-running the full bootstrap sequence.

Q4. During scale-in, an application must finish in-flight requests and upload final logs before the instance is destroyed. What implements this?

A. A lifecycle hook on Terminating:Wait
B. A longer health check grace period
C. A larger instance type
D. Enabling connection draining on the target group only
Correct answer: A. A termination lifecycle hook pauses the instance in Terminating:Wait, giving drain and log-flush scripts time to complete before the instance is removed.

Q5. A canary release needs 5% of production traffic sent to a new task set before full cutover. Which ALB feature does this?

A. Sticky sessions
B. Weighted target groups on a single listener rule
C. Cross-zone load balancing
D. A second listener on a different port
Correct answer: B. Weighted target groups let one routing rule split traffic by percentage across two target groups, which is the ALB-native canary and blue/green mechanism.

Q6. A Lambda function reading from Kinesis is exhausting connections on a downstream RDS instance during traffic spikes. What bounds this without changing the database?

A. Increase the function timeout
B. Set reserved concurrency on the function
C. Increase the Kinesis shard count
D. Move RDS to Multi-AZ
Correct answer: B. Reserved concurrency caps how many instances of the function run simultaneously, which directly bounds the number of concurrent downstream connections.

Q7. A latency-sensitive function must never pay a cold-start penalty, and the team also needs a hard ceiling on its concurrency. What configuration satisfies both?

A. Provisioned concurrency only
B. Reserved concurrency only
C. Both provisioned concurrency and reserved concurrency
D. A larger memory allocation
Correct answer: C. Provisioned concurrency pre-warms environments to remove cold starts, while reserved concurrency caps and guarantees maximum simultaneous executions — they are independent settings.

Q8. An IoT pipeline runs millions of short workflow executions per day and can tolerate at-least-once semantics. Which Step Functions type fits?

A. Standard Workflows
B. Express Workflows
C. Both cost the same at this volume
D. Neither — use SQS with Lambda only
Correct answer: B. Express Workflows are priced per execution and duration for high-volume, short-duration workloads and provide at-least-once semantics, unlike Standard's exactly-once but costlier per-transition pricing.

Q9. An order process requires a human approval step, must run for up to three days, and must be auditable with exactly-once execution. Which Step Functions type fits?

A. Express Workflows
B. Standard Workflows
C. A single long-running Lambda function
D. SQS with a visibility timeout
Correct answer: B. Standard Workflows support long-running executions, human approval wait states, and exactly-once semantics with a full execution history for audit.

Q10. A simple API endpoint only needs to write a request body directly into a DynamoDB table with no transformation logic. What avoids an unnecessary function?

A. A Lambda function with a minimal handler
B. An API Gateway native service integration to DynamoDB
C. An Application Load Balancer with a Lambda target
D. A Step Functions Express workflow
Correct answer: B. API Gateway can integrate directly with AWS services, bypassing Lambda entirely for simple pass-through, which removes a function's cost and cold-start surface.

Q11. A workload runs a licensed host-level security agent that cannot be containerized. Which compute option is viable?

A. Lambda
B. Fargate tasks
C. EC2 instances in an Auto Scaling Group
D. EKS Fargate profiles
Correct answer: C. A host-level agent needs access to the instance itself, which the serverless options do not expose — plain EC2 (or EC2-backed container capacity) is required.

Q12. A batch job runs for 40 minutes per invocation and can restart from a checkpoint. Which platform and pricing model fit?

A. Lambda with a 15-minute timeout increase
B. ECS or EC2 with Spot capacity
C. Lambda with provisioned concurrency
D. API Gateway with a Lambda integration
Correct answer: B. Lambda has a hard execution-duration ceiling, so a 40-minute job must run on containers or EC2; checkpointing makes it interruptible, which makes Spot the right pricing model.

Q13. A service runs continuously at a flat 3,000 requests per second around the clock with no idle periods. Which pricing approach is most cost-effective?

A. Lambda with on-demand billing
B. A committed baseline with Savings Plans or RIs
C. Spot Instances only
D. Fargate with per-second billing
Correct answer: B. A steady, predictable baseline with no idle time is exactly what commitments are for; per-request billing is excellent for spiky traffic and poor for a constant high-volume load.

Q14. A workload runs for ten minutes every hour and is idle the rest of the time. What is the cost-optimal model?

A. A continuously running EC2 fleet sized for the peak
B. Per-request billing with Lambda or Fargate
C. Reserved Instances for the full hour
D. A dedicated host
Correct answer: B. Paying for a full hour of instance time to serve ten minutes of work is the waste the scenario describes; per-request billing charges only for the work performed.

Q15. An ECS service's target-tracking policy is raising desired count, but new tasks sit in a pending state and never start. What is the most likely cause?

A. The service scaling policy is misconfigured
B. The capacity provider has no headroom to place the tasks
C. The task definition has the wrong IAM role
D. The ALB health check is failing
Correct answer: B. Pending tasks mean the service wants more tasks than the underlying capacity can place; the fix is capacity provider managed scaling or a larger node group maximum, not a change to the service policy.

Q16. A team wants each container task to have its own security group and IP address rather than sharing a host's port space. Which network mode provides this?

A. bridge
B. host
C. awsvpc
D. none
Correct answer: C. awsvpc mode gives every task its own ENI, private IP, and security group, enabling per-task security instead of shared host-level rules.

Q17. A company runs dozens of small independent containerized services and has no Kubernetes requirement or existing Kubernetes tooling. Which orchestrator minimizes operational overhead?

A. Amazon EKS
B. Amazon ECS
C. Self-managed Kubernetes on EC2
D. AWS Batch
Correct answer: B. ECS has no per-cluster control-plane charge and no Kubernetes upgrade cycle to manage, which makes it the lower-overhead choice when portability is not a requirement.

Q18. A workload needs a 500 GB local scratch volume for temporary processing files and runs for several hours at a time. Which platform fits?

A. Lambda with an EFS mount
B. Fargate tasks with ephemeral storage
C. EC2 instances with instance store or attached EBS
D. API Gateway with a Lambda integration
Correct answer: C. A large persistent local scratch volume plus multi-hour execution eliminates the serverless options; EC2 with instance store or attached EBS is the fit.

Q19. A team wants to shift 10% of traffic to a new container version, watch error rates for an hour, then shift the rest. Which combination supports this?

A. Route 53 weighted routing only
B. ALB weighted target groups with CloudWatch alarms on the new target group
C. A second ALB in a different region
D. An Auto Scaling Group with a warm pool
Correct answer: B. Weighted target groups perform the traffic split at the load balancer, and per-target-group metrics let you evaluate the canary before completing the shift.

Q20. A function processes SQS messages in batches, and one bad message in a batch causes the entire batch to be retried repeatedly. What addresses this?

A. Increase the function timeout
B. Enable partial batch failure reporting and configure a dead-letter queue
C. Reduce the batch size to one
D. Increase reserved concurrency
Correct answer: B. Partial batch failure reporting lets the function return only the failed message identifiers so the rest of the batch is not retried, and a DLQ isolates the poison message.

Q21. A team needs a hard guarantee that a function can always run at least 50 concurrent executions even when other functions in the account are consuming the regional limit. What provides this?

A. Provisioned concurrency of 50
B. Reserved concurrency of 50
C. A larger memory setting
D. A longer timeout
Correct answer: B. Reserved concurrency both caps and guarantees a function's share of the account's concurrency pool, reserving it away from other functions.

Q22. A workload's traffic is unpredictable, bursts to thousands of requests per second, and drops to zero for hours. The team wants no idle cost. Which pairing fits?

A. EC2 Auto Scaling with Reserved Instances
B. Lambda behind API Gateway with per-request billing
C. A dedicated host
D. EKS with a fixed node group
Correct answer: B. Per-request billing charges only for work performed, which is exactly what a workload with long idle periods and unpredictable bursts needs.

Q23. A team wants to run a containerized web service without managing any EC2 instances, but the service needs a sidecar container that must run on every node. What is the constraint conflict?

A. None — Fargate supports sidecars on every node
B. Fargate has no shared nodes, so a per-node sidecar is impossible; managed node groups are required
C. Sidecars require Lambda layers
D. Sidecars require a dedicated host
Correct answer: B. Fargate's per-task isolation means there is no node to attach a per-node sidecar to; a node-based launch type is required for that pattern.

Q24. A company wants maximum flexibility to change instance families and regions over a one-year commitment while still receiving a discount on a steady baseline. What should they purchase?

A. EC2 Instance Savings Plans
B. Compute Savings Plans
C. Standard Reserved Instances
D. Spot Instances
Correct answer: B. Compute Savings Plans apply across any instance family, region, and OS, and even to Fargate and Lambda usage, trading a slightly lower discount for maximum flexibility.

Q25. A batch workload is fault-tolerant and can be restarted, and the team wants the lowest possible compute cost with protection against a single capacity pool being unavailable. What should they configure?

A. On-Demand instances in a single Availability Zone
B. A mixed-instance Auto Scaling Group or Spot Fleet with multiple instance types and pools
C. Reserved Instances for the batch fleet
D. A dedicated host
Correct answer: B. Spot with diversification across instance types and capacity pools lowers cost and reduces the chance that a single pool's interruption affects the whole fleet; the workload's restartability makes the interruption notice tolerable.

Preview

Everything in this week assumed that the compute layer is the hard part — that once you have chosen a platform and a scaling model, the data layer is a solved problem you configure once. That assumption holds right up until the workload has to survive the loss of an Availability Zone, or serve reads from a second continent, or accept writes in two regions at the same time. The open question this week leaves unresolved is what happens to the database when the compute layer is already multi-AZ and the storage underneath it is not.

Day 22 answers that for the relational case with Aurora, whose storage layer is replicated six ways across three Availability Zones and self-heals without a failover event, and then extends it across regions with Aurora Global Database — replication that happens in the storage layer rather than through binlog shipping, which is why the cross-region lag is measured in sub-second terms and a secondary region can be promoted in under a minute. The interesting part is not the numbers but the mechanism: understanding why storage-layer replication behaves differently from a read replica is what separates the two answers on an exam question that mentions both.

Sources