Amazon ECS Fargate Architecture & Task Definitions
Recap: From Network Plumbing to Compute Placement
Week 2 closed on a decision framework rather than a service: the networking lab exam asked you to choose between TGW, peering, and PrivateLink for a given reachability requirement, to justify a Direct Connect resiliency pattern against a stated RTO, and to place Route 53 Resolver endpoints correctly for hybrid DNS resolution. That cluster is one of the highest-weight areas on the exam, and the reason it is weighted so heavily is that connectivity decisions are expensive to reverse — a CIDR overlap discovered after a peering mesh is built, or a DNS forwarding rule that only works in one direction, is a multi-week remediation rather than a config change.
Today extends that same reasoning one layer up the stack. Once the network path exists — a shared VPC subnet, a TGW attachment, a PrivateLink endpoint — the next question is what actually runs at the far end of it, and how much of the host you are willing to own. Fargate is the answer that says "none of it," and the interesting part of the exam question is never whether Fargate is serverless; it is which of the task definition's fields encode the security and networking decisions you just spent two weeks learning to make. The awsvpc network mode is where those two threads meet: it is the mechanism that gives a container the same first-class network identity that an EC2 instance has always had.
Foundations You'll Need Today
Containers, and Why They Are Not Just Small Servers
A container is a way of packaging an application together with everything it needs to run — its code, its libraries, its runtime — into a single unit called an image. When you run that image, you get a container: an isolated process that believes it is alone on a machine, even though it is sharing the underlying operating system kernel with other containers. The problem this solves is the classic "it works on my laptop" gap. A virtual machine solves the same problem by shipping an entire operating system, which is heavy — gigabytes of image, a full boot sequence, and a lot of duplicated work when you run twenty of them. A container ships only the application layer, so it starts in seconds and you can pack many of them onto one machine. The word "host" appears throughout today's material and simply means the machine the containers are actually running on. The entire Fargate-versus-EC2 decision is a question about who owns that host.
VPCs, Subnets, ENIs, and Security Groups
A VPC (Virtual Private Cloud) is your own private network inside AWS, defined by a range of IP addresses you choose — something like 10.0.0.0/16, which is a shorthand for "all addresses starting with 10.0." That range is carved into subnets, each of which lives in one Availability Zone (a physically separate data center cluster) and holds a slice of the addresses. When you launch anything into a VPC — an EC2 instance, a database, or a container — it gets a network interface, called an ENI, which is the thing that actually holds an IP address from the subnet and connects the resource to the network. A security group is a firewall attached to that interface: a list of rules saying which traffic is allowed in and out. Two facts matter for today. First, every resource that gets an ENI consumes one IP address from its subnet, so a subnet has a hard ceiling on how many things can live in it. Second, a security group rule can reference another security group rather than an IP range, which is how you say "allow traffic from the load balancer" without hardcoding addresses. When today's material says a Fargate task "has its own ENI," it means the task is a first-class citizen of the VPC with its own address and its own firewall, exactly like an EC2 instance.
IAM Roles and What It Means to "Assume" One
An IAM role is a set of permissions that is not attached to a person or a password. Instead, it is designed to be assumed — temporarily taken on — by something that needs to call AWS APIs: an EC2 instance, a Lambda function, a container, or even a user in another account. When something assumes a role, AWS hands it short-lived credentials that expire on their own, which is why roles are preferred over long-lived access keys that can leak. A role has two halves that are easy to conflate. The permissions policy says what the role is allowed to do once you have it. The trust policy says who is allowed to assume it in the first place. Today's material depends on a distinction between two roles that both belong to the same task: one that the platform uses on the task's behalf, and one that the application code inside the container uses. If you do not know that a role is something you assume rather than something you log into, that distinction will read as arbitrary rather than as two genuinely different identities.
Load Balancers, Target Groups, and Health Checks
A load balancer sits in front of your application and spreads incoming requests across however many copies of it are running, so that no single copy is overwhelmed and so that traffic keeps flowing when one copy dies. In AWS, an Application Load Balancer works with a target group, which is simply the list of places the load balancer is allowed to send traffic — and it needs to know which of those places are actually healthy. It finds out by health checks: on a schedule, it sends a request to a path you configure, such as /health, and treats a successful response as proof that the target is alive. A target that fails enough checks in a row is taken out of rotation until it recovers. This matters today because a container platform uses the same mechanism to decide whether a newly started task is ready to receive traffic, and a mismatch between the health check path and what the container actually serves is one of the most common reasons a deployment never finishes.
With that grounding — containers that need a host, a VPC that gives each one an address and a firewall, roles that are assumed rather than logged into, and a load balancer that only sends traffic to targets it has verified — here is why Fargate exists and what problem it actually solves.
1. Why This Is on the Exam
The SAP-C02 exam treats container compute as a placement decision, not a product tour. The scenario almost always arrives with a constraint attached — a team with no Kubernetes experience, a workload with unpredictable burst, a compliance requirement that no shared host may run two tenants' code, or a mandate to reduce the number of EC2 instances the platform team patches. Fargate is the option that satisfies several of those constraints simultaneously, and the exam expects you to know precisely which ones it satisfies and which it quietly does not.
The architectural problem Fargate solves is the split of responsibility in a container platform. Running containers on EC2 means you own the AMI lifecycle, the container runtime version, the agent, the instance sizing, the bin-packing of tasks onto hosts, and the capacity buffer that keeps a scale-out from failing while instances boot. None of that is application logic, and all of it is a source of incidents. Fargate removes the host from your inventory entirely: you declare CPU and memory at the task level, and AWS provisions an isolated compute allocation for that task. The trade is that you lose host-level control — no SSH, no custom AMI, no DaemonSets, no GPU, no persistent local storage beyond ephemeral task storage — and you pay a per-task premium over well-utilized EC2.
This maps most directly to Domain 2 (Design Resilient Architectures) and Domain 4 (Design Cost-Optimized Architectures), with a recurring appearance in Domain 3 (Design High-Performing Architectures) when the scenario is about scaling behavior. The resilient-architecture framing is usually about blast radius: a task that gets its own ENI and security group cannot be reached by a neighbor's misconfigured port mapping, and a task that dies takes only itself down. The cost framing is the inverse: Fargate is cheap at low and spiky utilization and expensive at sustained high utilization, so the exam wants you to notice whether the workload is steady-state or bursty before you pick it.
There is also a governance angle that shows up in multi-account scenarios. Because a Fargate task has no host, there is no long-lived compute resource for a workload team to patch, and no shared kernel for a noisy neighbor to exploit. For organizations that have just spent Phase 1 building account separation and SCP guardrails, Fargate is the compute primitive that keeps the isolation boundary at the task rather than at the instance — which is a much easier boundary to reason about in an audit.
2. Mechanism: What Actually Happens When a Task Starts
A task definition is a versioned, immutable blueprint. It declares one or more container definitions, each with an image, a command, port mappings, environment variables, secrets references, log configuration, health check, and resource requirements. At the task level it declares the CPU and memory for the whole task, the network mode, the execution role, the task role, and any volumes. When you register a new revision, the old revision remains addressable by ARN, which is what makes rollback a matter of pointing a service at a previous revision rather than rebuilding anything.
When the ECS scheduler places a task on Fargate, it does not find a host with spare capacity. It asks the Fargate capacity provider to provision a micro-VM dedicated to that task, sized to the CPU and memory in the task definition, and then starts the containers inside it. The task's containers share that micro-VM's network namespace and, in awsvpc mode, share a single elastic network interface that is attached to the subnets and security groups you specified. That ENI is the task's identity on the VPC: it has a private IP from the subnet CIDR, it appears in flow logs, and it is what a security group rule or a VPC endpoint policy evaluates against.
Two IAM roles are in play and they are routinely confused. The task execution role is used by the ECS agent itself — to pull the image from ECR, to fetch secrets from Secrets Manager or SSM Parameter Store, and to write logs to CloudWatch Logs. The task role is assumed by the application code inside the container, and it is the one that should be scoped to least privilege for whatever AWS APIs the workload calls. A task that can pull its image but cannot read the S3 bucket it needs is a task execution role that is fine and a task role that is missing a statement; the reverse produces a task that never starts.
Placement is governed by the service's desired count and its deployment configuration, not by a scheduler you tune. For a service with a load balancer attached, ECS registers the task's ENI and port with the target group, waits for the target to pass health checks, and only then counts the task toward the deployment's minimum healthy percent. That ordering is why a Fargate service can roll out without dropping requests even though every task is replaced: the new tasks are healthy before the old ones are drained. The service also owns the scaling relationship — an Application Auto Scaling target attached to the service adjusts desired count in response to CloudWatch metrics or a schedule, and each new desired count is a new set of micro-VMs.
3. The Core Decision Boundary: Who Owns the Host
Every Fargate scenario question reduces to one fork: does the workload require control over the host, or does it only require control over the container? If the answer is "only the container," Fargate is almost always correct, because it removes an entire class of operational work and an entire class of failure. If the answer is "the host," Fargate is disqualified outright and no amount of cost or simplicity argument recovers it — the exam is testing whether you recognize the disqualifying requirement before you start optimizing.
The disqualifying requirements are specific and worth memorizing as a set. DaemonSets and any per-node agent pattern are impossible because there is no node that persists across tasks. GPU workloads are unsupported. Privileged containers, host networking, and host port mapping are unavailable. Persistent local storage beyond the task's ephemeral volume is not available, so anything that needs a local cache that survives a task restart must move to EFS or a service. Custom AMIs, custom kernel parameters, and instance-level agents are gone. And the pricing model changes shape: you pay for the CPU and memory you reserve for the lifetime of the task, rounded up to the nearest supported combination, whether or not the container uses it.
The other side of the fork is subtler and is where most scenario questions actually live. Fargate is not automatically cheaper than EC2 just because there are no instances to manage. At low utilization — a service that idles most of the day, or a batch job that runs for minutes — Fargate wins decisively because you pay only while the task exists. At sustained high utilization, a well-packed EC2 Auto Scaling group with Savings Plans applied will beat Fargate on raw compute cost, and the exam will sometimes hand you a steady-state, high-utilization workload specifically to see whether you reflexively pick serverless.
| Requirement in the scenario | Fargate | EC2 launch type |
|---|---|---|
| No instances to patch or right-size | Yes — no host in your inventory | No — you own AMI, agent, and sizing |
| Per-task network isolation | Yes — awsvpc is the only mode | Optional — awsvpc, bridge, or host |
| DaemonSet / per-node agent | Not possible | Supported |
| GPU or privileged containers | Not supported | Supported on GPU instance families |
| Persistent local storage | Ephemeral only; use EFS | Instance store or EBS volumes |
| Cost at sustained high utilization | Higher per vCPU-hour | Lower with Savings Plans and good bin-packing |
| Cost at spiky or low utilization | Lower — pay only while running | Higher — idle capacity still billed |
| Scale-out latency | Seconds to provision a micro-VM | Minutes if a new instance must boot |
4. Configuration Modes and Their Tradeoffs
The task definition's network mode is the first knob and, on Fargate, it is not really a choice: awsvpc is the only supported mode. That matters because it means every Fargate task consumes an IP address from the subnet it is placed in, and the subnet's available IP count becomes a hard ceiling on how many tasks can run there. A service that scales to 200 tasks across two subnets needs 100 free addresses in each, and a /24 that also hosts other resources will run out. This is the single most common capacity-planning mistake in Fargate designs, and it is invisible until a scale-out event fails with a resource initialization error.
CPU and memory are declared at the task level and must be a valid combination — Fargate does not accept arbitrary values, and the ratio of CPU to memory is constrained. The practical consequence is that you cannot fine-tune a task to exactly the resources it needs; you pick from a menu, and the menu's granularity means you often reserve more than the workload uses. Memory is the harder constraint to get right, because exceeding the task's memory limit kills the container with an out-of-memory exit rather than throttling it, so the failure is abrupt and the metric to watch is the memory utilization against the task limit, not the container's own reported usage.
The launch type and capacity provider strategy determine what happens when the service needs more tasks. With Fargate, the capacity provider is FARGATE or FARGATE_SPOT, and a strategy can weight between them — for example, a base of on-demand tasks that always exist plus a weighted share of Spot tasks for burst. Fargate Spot tasks can be reclaimed with a two-minute warning, delivered as a SIGTERM to the container, so a service that mixes them must handle graceful shutdown and must not be the only copy of a stateful process. The tradeoff is direct: Spot capacity is materially cheaper, and the price is that a fraction of your tasks will be interrupted on AWS's schedule rather than yours.
Logging and secrets are configured per container and are the fields most often left at defaults in a hurry. The awslogs driver ships stdout and stderr to a CloudWatch log group, and the log group must exist or the task fails to start — a detail that turns a missing log group into a deployment outage rather than a missing log. Secrets can be injected as environment variables from Secrets Manager or SSM Parameter Store, resolved by the task execution role at start time, which keeps credentials out of the image and out of the task definition's plaintext. The tradeoff is that the value is resolved once at task start, so a rotated secret requires a new task to take effect — which is fine for a rolling deployment and a problem for a long-lived task.
5. Sizing, Limits and Quotas
Fargate's sizing model is a fixed menu rather than a continuous range, and the numbers matter because they determine both cost and whether a workload fits at all. Task-level CPU is expressed in vCPU units and memory in GB, and only certain pairings are valid — the general shape is that each vCPU step permits a bounded range of memory, so a 0.25 vCPU task cannot be given 30 GB. The largest task sizes support multiple vCPUs and tens of gigabytes of memory, which is enough for most application containers but is a real ceiling for in-memory analytics or large JVM heaps.
Ephemeral storage is the limit that surprises people. Each Fargate task gets a small default amount of ephemeral storage for the container filesystem, and it can be increased up to a documented maximum at task-definition time. Anything written there is lost when the task stops, and anything that needs to be shared between containers in the same task or persisted across tasks must go to EFS or an external service. A container that writes a large temporary file — an image transform, a report generation, a database dump — will hit this limit and fail in a way that looks like a disk error inside the application rather than a platform limit.
Quotas operate at several layers and each one produces a different error. At the account and Region level there are limits on the number of tasks a service can run and on the rate at which you can register task definitions or call the ECS API. At the network level, the subnet's available IP addresses cap concurrent tasks in awsvpc mode. At the load balancer level, the target group has its own limits and the ALB has a limit on targets per target group. And at the Fargate capacity level, Spot capacity is not guaranteed — a FARGATE_SPOT-only service can fail to place tasks during a capacity event, which is exactly why the recommended pattern is a base of on-demand tasks with Spot as the burst layer.
| Dimension | What to check | Failure symptom when exceeded |
|---|---|---|
| Task CPU / memory pairing | Must be a supported combination | Task definition registration rejected |
| Ephemeral storage | Default is small; raise it explicitly if needed | Container exits with a disk/write error |
| Subnet IP addresses | One ENI per task in awsvpc mode | Resource initialization error on scale-out |
| Service task count | Account/Region service quota | Tasks stay in PENDING or fail to place |
| Fargate Spot capacity | Not guaranteed; use a base of on-demand | Placement failure during capacity events |
| ALB targets per target group | Load balancer quota | Registration failures during deployment |
6. Failure Modes and What They Look Like in Production
The most common Fargate failure is a task that starts and immediately stops, and the diagnostic path is a fixed sequence. First, read the stopped task's reason and the container's exit code — a non-zero exit with a short runtime is almost always an application or configuration error, not a platform problem. Second, check the stopped reason for a platform-level message such as an out-of-memory kill or an inability to pull the image. Third, look at the CloudWatch log stream for the task, which exists only if the log configuration was correct and the log group existed. A task that stops with no log stream at all is usually a task execution role that cannot write logs, which is a different fix from an application crash.
Image pull failures are the second family and they cluster around permissions and networking. If the image is in ECR in another account, the task execution role needs permission to pull it and the repository needs a policy allowing that principal. If the task runs in a private subnet with no NAT gateway and no VPC endpoints for ECR and S3, the pull will time out — ECR image layers are stored in S3, so a private-subnet Fargate task needs both the ECR interface endpoints and the S3 gateway endpoint, or a NAT path. This is a classic exam scenario: a service that works in a public subnet and fails to start in a private one, with no application change.
Health-check failures present differently and are more insidious because the task is running fine. If the target group's health check path, port, or protocol does not match what the container actually serves, ECS will start tasks, fail to register them as healthy, and then roll the deployment back — repeatedly, in a loop, with the service never reaching steady state. The signal is a deployment circuit breaker trip or a service that oscillates between desired and running counts. The first diagnostic move is to check the target group's health check configuration against the container's actual listening port and path, and to confirm the security group on the task allows traffic from the load balancer's security group on that port.
Capacity and placement failures are the third family and they are the ones that look like the platform is broken when it is not. A scale-out that fails because the subnet has no free IPs, a FARGATE_SPOT service that cannot place tasks during a capacity event, or a service that hits an account-level task quota all present as tasks stuck in PENDING or a service that never reaches its desired count. The distinguishing feature is that nothing is crashing — the tasks that do run are healthy, and the gap is between desired and running. That gap, not the task logs, is where the investigation starts.
7. The Operational and SRE Angle
Fargate changes the shape of the on-call runbook because the failure surface moves from the host to the task and the deployment. There is no instance to SSH into, no disk to fill, no kernel to patch, and no agent to restart — which removes a large fraction of traditional compute incidents. What remains is a smaller but sharper set: tasks that will not start, deployments that will not stabilize, and scaling that will not keep up. Each of those has a metric and an alarm, and the runbook for each is short enough to be genuinely useful at 3 a.m.
The metrics worth alarming on are the service-level ones rather than the task-level ones. Running task count versus desired count is the primary signal: a persistent gap means placement or health-check trouble, and it is the earliest indicator that a deployment is not converging. CPU and memory utilization at the service level tell you whether the scaling target is set sensibly, and memory utilization approaching the task limit is the leading indicator of an out-of-memory kill. On the load balancer side, target response time and the unhealthy host count catch the case where tasks are running but not serving. For a service behind an ALB, the combination of a rising unhealthy host count and a flat running task count is the signature of a health-check mismatch rather than a capacity problem.
Deployment safety is where Fargate's operational model pays off, and it is worth configuring deliberately. The deployment circuit breaker, with rollback enabled, makes a failed deployment self-healing: if the new tasks do not reach a healthy state, ECS rolls back to the previous task definition revision automatically instead of leaving the service in a half-deployed state. Combined with a minimum healthy percent that keeps capacity above the level the SLO requires, this turns a bad image push into a brief, self-correcting event rather than an incident. The tradeoff is that the circuit breaker needs a health check that actually reflects application health — a health check that returns 200 as soon as the process is listening will let a broken deployment through.
From an SLO perspective, the interesting property of Fargate is that scale-out is fast enough to be part of the availability story rather than a background concern. A task can be provisioned in seconds, so a service that scales on a leading indicator — queue depth, request rate, or a custom metric — can absorb a traffic step change without the multi-minute instance-boot lag that an EC2 Auto Scaling group has to plan around. That does not remove the need for headroom, because the scaling policy still has to fire and the new tasks still have to pass health checks, but it shortens the window in which a spike is visible to users. The corollary is that a scaling policy with a long evaluation period wastes the advantage; for latency-sensitive services, shorter periods with a lower threshold are usually the right trade.
8. Edge Cases and Exam Gotchas
The gotchas cluster around the boundary between what Fargate hides and what it does not. It hides the host, but it does not hide the network: a Fargate task still needs a subnet with free IPs, still needs a route to whatever it must reach, and still needs security group rules that permit that traffic. Scenarios that describe a task that cannot reach a database in another VPC are testing whether you remember that the task's ENI is subject to the same routing and security group rules as any other resource in that subnet — there is no special Fargate networking path.
The second cluster is about the two IAM roles. A task that fails to start because it cannot pull its image is a task execution role problem. A task that starts and then gets AccessDenied from an AWS API is a task role problem. A task that starts but produces no logs is usually a task execution role that lacks the log-write permission, or a log group that does not exist. The exam will describe the symptom and expect you to name the role, so it is worth being able to state the split in one sentence: execution role for the platform's actions on the task's behalf, task role for the application's actions.
The third cluster is about persistence and state. Fargate tasks are ephemeral by design, so any scenario that requires a container to keep data across restarts is pointing at EFS, S3, or a database — never at the task's local storage. A related trap is the session: a service that stores sessions in task memory will lose them on every deployment, and the fix is an external session store, which is exactly the ElastiCache material from Week 4. The exam likes this because it connects two domains with one scenario.
The fourth cluster is cost-shaped. Fargate pricing is per-second for the CPU and memory reserved, with a minimum billing duration, so a task that runs for a few seconds is billed for the minimum. A workload that spins up a task per request will pay for that minimum on every request, which is usually the point at which Lambda becomes the better answer. Conversely, a long-running task that idles at low CPU still pays for its full reservation, which is the point at which EC2 with a Savings Plan becomes the better answer. Recognizing which side of that line a scenario sits on is the whole game.
9. Fargate vs. the Services It Gets Confused With
Fargate is most often confused with three things: the EC2 launch type for the same orchestrator, Lambda, and EKS with Fargate profiles. The EC2 launch type comparison is about who owns the host and is covered above; the practical rule is that if the scenario mentions patching instances, custom AMIs, GPU, or sustained high utilization, you are on EC2, and if it mentions operational overhead, spiky load, or per-task isolation, you are on Fargate. The two can coexist in one cluster, and a mixed capacity provider strategy is a legitimate answer when a scenario wants a cheap baseline plus elastic burst.
The Lambda comparison is about execution model rather than packaging. Lambda is event-driven, has a hard maximum execution duration, and scales per invocation; Fargate runs a long-lived process that you scale by task count. A request-response workload with short, bursty invocations is Lambda's shape. A workload that holds a connection pool, runs a background loop, or needs a process that lives for hours is Fargate's shape. The exam sometimes offers both for a workload that could technically run on either, and the tiebreaker is usually the connection or the duration limit.
The EKS comparison is the one tomorrow's material covers in depth, but the boundary is worth stating now: Fargate is a capacity option for both ECS and EKS, and choosing EKS with Fargate profiles buys you Kubernetes' API and ecosystem at the cost of the Fargate limitations plus the operational weight of a Kubernetes control plane. If the team already runs Kubernetes, that trade is usually worth it; if they do not, ECS with Fargate is the shorter path to production.
| Option | Pick it when… | Avoid it when… |
|---|---|---|
| ECS on Fargate | No host management, spiky load, per-task isolation, small platform team | GPU, DaemonSets, custom AMI, sustained high utilization |
| ECS on EC2 | Steady high utilization, host-level control, GPU, cost-optimized with Savings Plans | Team cannot absorb AMI/agent/patching work |
| Lambda | Event-driven, short invocations, per-request scaling, no idle cost | Long-running processes, connection pools, >15 min execution |
| EKS with Fargate profiles | Existing Kubernetes tooling and portability requirement | No Kubernetes experience; DaemonSets required |
| EKS with managed node groups | Kubernetes plus host-level control and DaemonSets | Team wants zero node management |
Hands-on Lab: ALB Path Routing in Front of a Least-Privilege Fargate Service
Objective. Stand up an ECS Fargate service behind an Application Load Balancer, route to it by path, and give the task a dedicated IAM role scoped to exactly one S3 prefix. The point of the exercise is to feel where the two IAM roles diverge and to see the awsvpc ENI show up as a first-class network identity.
Step 1 — Create the network and the log group. In a VPC with at least two private subnets in different Availability Zones, create a security group for the ALB that allows inbound HTTP from your own address range, and a separate security group for the tasks that allows inbound on the container port only from the ALB's security group. Create a CloudWatch log group named for the service before you register anything — a missing log group is a task-start failure, and creating it first removes that variable from the lab.
Step 2 — Create the two roles. Create a task execution role with the AWS-managed policy for ECS task execution, which covers ECR pulls and log writes. Then create a task role with a single inline policy allowing s3:GetObject and s3:ListBucket on one bucket and one prefix, and nothing else. Note the difference in what each role is for; you will verify both in step 6.
Step 3 — Register the task definition. Register a Fargate task definition with awsvpc network mode, a supported CPU/memory combination, the execution role, the task role, and one container definition using a small public image. Configure the awslogs driver against the log group from step 1, and map the container port. Confirm the task definition registers without a CPU/memory pairing error — if it rejects, you have picked an unsupported combination, which is itself the lesson.
Step 4 — Create the target group and the ALB rules. Create an IP-type target group (Fargate tasks register by IP, not instance ID) with a health check path that the container actually serves. Create the ALB with a listener on port 80, a default rule that returns a fixed response, and a path-based rule for /app/* that forwards to the target group. The default rule returning a fixed response is deliberate: it lets you prove the path rule is doing the routing rather than everything falling through.
Step 5 — Create the service and watch the deployment. Create the ECS service with the Fargate capacity provider, the desired count, the two private subnets, the task security group, and the load balancer configuration pointing at the target group. Watch the service events: you should see the task move from PROVISIONING to PENDING to RUNNING, then register with the target group, then pass health checks. If it loops, the events will name the reason — an image pull failure, a health check failure, or a placement failure — and each maps to a different fix.
Step 6 — Verify the isolation and the roles. Confirm the task's ENI appears in the VPC with a private IP from the subnet, and that the security group on it permits only the ALB. Then exec into the container and call the S3 API against the permitted prefix (should succeed) and against a different bucket (should fail with AccessDenied). That failure is the task role working correctly. Finally, break the log group name in a new task definition revision and deploy it to see the task fail to start — the fastest way to internalize the execution role's scope.
Step 7 — Clean up. Delete the service, the ALB, the target group, the task definition revisions, the log group, and the two roles. Leaving the service running is the most common way to accumulate a surprising bill from a lab.
Scenario Question Drills
Q1. Why does Fargate's awsvpc network mode assign each task its own ENI?
Q2. A Fargate service in a private subnet fails to start with an image pull timeout, but the same task definition works in a public subnet. What is the most likely cause?
Q3. A workload requires a per-node log collection agent on every compute node. Which compute option is disqualified?
Q4. A task starts successfully but the application receives AccessDenied when calling S3. Which role needs to change?
Q5. A service scales from 20 to 200 tasks during a daily peak, and scale-out starts failing with a resource initialization error. What is the most likely cause?
Q6. A container writes a large temporary file during processing and exits with a disk-related error. What is the Fargate-specific explanation?
Q7. A service is deployed with a new task definition and never reaches steady state; tasks start, fail health checks, and are replaced in a loop. What is the first thing to check?
Q8. A team wants a cheap baseline of always-on tasks plus elastic burst capacity that can tolerate interruption. What should the service use?
Q9. A task is killed with an out-of-memory exit even though the container's own reported memory usage looks moderate. Why?
Q10. A workload runs a long-lived process that maintains a database connection pool and runs for many hours. Which compute option fits best?
Q11. A Fargate task starts but produces no log output in CloudWatch. What is the most likely cause?
Q12. A service stores user sessions in task memory. After every deployment, users are logged out. What is the correct fix?
Q13. A team wants a failed deployment to roll back automatically instead of leaving the service half-deployed. What should they configure?
Q14. A workload needs a GPU for inference. Which statement is correct?
Q15. A steady-state production service runs at consistently high CPU utilization around the clock. Cost is the primary constraint. What is the most likely correct choice?
Peek into Tomorrow
Everything above assumed ECS as the orchestrator, which let us treat the task definition as the unit of truth and the service as the thing that keeps a desired count true. That assumption is doing a lot of quiet work, and it is worth asking what changes when the orchestrator is Kubernetes instead. The task definition has no direct equivalent — a Kubernetes manifest splits the same information across a Pod spec, a Deployment, a Service, and a ServiceAccount — and the IAM story changes shape entirely, because a pod does not assume a task role by default. The mechanism that bridges that gap is IRSA, and it is the kind of detail that a scenario question can hinge on without ever naming it.
The more interesting unresolved question is about capacity. Fargate profiles exist on EKS too, and they carry the same limitations we just catalogued — no DaemonSets, no host access, a per-pod premium — but on EKS those limitations collide with assumptions the Kubernetes ecosystem makes everywhere. Log collectors, service meshes, and node-level autoscalers all assume a node exists. So the real question tomorrow has to answer is not "Fargate or nodes" but "which parts of this cluster can tolerate having no node at all," and how a managed node group changes that answer by giving you an AWS-managed Auto Scaling group you still do not have to patch by hand.
Sources
- Amazon ECS Developer Guide — AWS Fargate
- Amazon ECS Developer Guide — Task Definitions
- Amazon ECS Developer Guide — Task Networking (awsvpc mode)
- Amazon ECS Developer Guide — IAM Roles for Tasks
- Amazon ECS Developer Guide — Fargate Task Ephemeral Storage
- Amazon ECS Developer Guide — Rolling Update Deployment and Circuit Breaker
- Amazon ECS Developer Guide — Service Auto Scaling
- Amazon ECS Best Practices Guide — Security and IAM
- Amazon ECS Developer Guide — Service Load Balancing
- AWS Whitepaper — Fault Isolation Boundaries