Day 16 of 70 · Week 3
Day 16 / 70 Week 3 of 14 Phase 2: Compute, Containers & Global Databases

Amazon EKS Architecture & Managed Node Groups vs Fargate

🕑 ~58 min read · 3 services covered
EKS Managed Node Groups EKS Fargate Profiles

Recap: Where We Left Off

Day 15 established that Fargate on ECS removes the EC2 layer entirely: a task definition declares CPU, memory, and an IAM task role, and the scheduler places each task into its own isolated compute allocation. The awsvpc network mode then gives every task its own ENI and security group, which is what makes per-task network policy meaningful rather than decorative. That model is clean, and it is also the source of the trap this day sets up in a different service.

The trap is that "serverless containers" is a phrase that travels across services without carrying its constraints with it. On ECS, Fargate's isolation is mostly a gift — you get per-task security groups and no host patching. On EKS, the same launch type sits underneath a Kubernetes control plane that assumes nodes exist, and a large fraction of the Kubernetes ecosystem is built on that assumption. DaemonSets, host-path volumes, node-level autoscaling, and node affinity all quietly stop working. The exam does not ask you to recite that Fargate is serverless; it asks you to notice when a stated requirement — a log agent on every node, a GPU, a privileged container — makes Fargate structurally impossible. Same trap, different service.

Foundations You'll Need Today

Today's material sits on top of four ideas that the rest of this curriculum treats as background knowledge. If any of them are new, read this section first — the rest of the day will make far more sense with it in place.

Containers, and Why They Are Not Virtual Machines

A virtual machine is a complete simulated computer. It has its own operating system kernel, its own boot process, and its own copy of everything the OS needs. That completeness is what makes VMs portable and isolated, and it is also what makes them heavy: you pay for a full OS boot and a full OS's worth of memory before your application even starts. A container takes a different approach. Instead of simulating a computer, it uses features built into the host's existing operating system to fence off a process — or a small group of processes — so that it sees only its own files, its own network, and its own view of the system. The application inside a container believes it is alone on a machine, but it is actually sharing the host's kernel with every other container on that host.

That sharing is the whole point and the whole risk. Because there is no second kernel to boot, a container starts in a fraction of a second and costs almost nothing in overhead, which is why you can run dozens of them on a machine that would struggle to host three VMs. But "sharing a kernel" also means that a flaw in the kernel is a flaw shared by every container on that host. When this day talks about a "strong isolation boundary" or "no two tenants share a kernel," that is the distinction being drawn: containers on the same machine are separated by software rules, while separate machines are separated by hardware. Keep that in mind, because it is the reason one of today's compute options is chosen for security rather than for cost.

Kubernetes: Control Plane, Data Plane, and Nodes

Kubernetes is a system for running containers across many machines and keeping them running. It is split into two halves that do very different jobs. The control plane is the brain: it stores the desired state of everything you have asked for, decides where each container should run, and watches constantly to make reality match your description. The data plane is the muscle: the actual machines where containers execute. In Kubernetes, those machines are called nodes, and each node runs a small agent called the kubelet whose job is to receive instructions from the control plane and start or stop containers accordingly.

Two more terms appear throughout today's material. The scheduler is the part of the control plane that picks which node a given container should land on, based on how much CPU and memory it asked for and what the nodes have free. A DaemonSet is a special kind of workload that says "run one copy of this container on every node, always." Log collectors, monitoring agents, and security scanners are usually deployed this way, because they need to be present on every machine rather than running as a single shared service. Hold onto that definition — the fact that a DaemonSet needs a node to attach to is the single most important constraint in today's decision.

VPC Networking: Subnets, IP Addresses, and Elastic Network Interfaces

A Virtual Private Cloud (VPC) is your own private network inside AWS. It is divided into subnets, each of which owns a range of IP addresses — for example, a subnet might own every address from 10.0.1.0 through 10.0.1.255. Anything that needs to communicate on the network needs an address from one of those ranges, and the number of addresses a subnet owns is finite. This matters more than it sounds, because running out of addresses in a subnet is a real and common failure: new machines simply cannot be created until space is freed.

An Elastic Network Interface (ENI) is the virtual equivalent of a network card. Every EC2 instance has at least one, and each ENI holds one primary IP address plus a number of additional addresses it is allowed to carry. That number is not unlimited — it depends on the size of the instance, and larger instances get more. This is the detail that makes today's pod-scheduling gotcha possible: if every container needs its own real IP address, and each machine can only hold so many addresses, then a machine can run out of addresses long before it runs out of CPU or memory. A security group, for completeness, is a set of firewall rules attached to an ENI that decides which network traffic is allowed in and out. It is the AWS equivalent of a firewall rule set, and it operates per network interface rather than per machine.

IAM Roles vs. Kubernetes Permissions

AWS Identity and Access Management (IAM) is how AWS decides who is allowed to call which AWS APIs. An IAM role is a set of permissions that can be assumed by a person, an application, or a service, and it is the standard way to grant access without handing out long-lived passwords. When you hear that someone has "full EKS permissions," it means their IAM role allows them to call the AWS APIs that create, modify, and delete EKS clusters.

Kubernetes has its own, entirely separate permission system called RBAC (role-based access control), which decides what a user is allowed to do inside a cluster — list pods, read logs, delete deployments. These two systems do not talk to each other automatically. Having permission to call the AWS API that manages a cluster does not grant you any ability to interact with the workloads running in it, and vice versa. Something has to explicitly connect an AWS identity to a Kubernetes identity, and today's material covers what happens when that connection is missing. It is a frequent source of confusion precisely because both systems use the word "role" and both are described as "permissions."

With that grounding, here's why EKS exists, what problem it actually solves, and why the choice of where your containers run is a genuine architectural decision rather than a configuration detail.

1. Why EKS Compute Choice Is on the Exam

EKS appears on SAP-C02 primarily as a decision problem rather than a configuration problem. The exam rarely asks you to write a manifest or recall a kubectl flag. It gives you a workload description — a team with existing Kubernetes tooling, a batch job with unpredictable arrival, a compliance requirement that no two tenants share a kernel — and asks which compute substrate you would place it on. The four answers are almost always self-managed nodes, managed node groups, Fargate profiles, and "not EKS at all, use ECS."

That framing maps to Domain 2 (Design Resilient Architectures) and Domain 3 (Design High-Performing Architectures), with a recurring cost thread from Domain 4. The resilience angle is about what happens when a node dies and how much of your capacity is self-healing. The performance angle is about scheduling latency, bin-packing efficiency, and whether your pods can actually get the resources they request. The cost angle is the one candidates most often get wrong: Fargate is frequently described as "more expensive," which is true per-vCPU-hour but not necessarily true per-workload once you account for the idle capacity that node groups carry.

The reason this is worth a full day rather than a paragraph is that the wrong answer is usually plausible. A scenario describing a small, bursty, low-utilization workload with no operational team reads like a Fargate scenario, and often is — but if the same scenario mentions a sidecar log shipper deployed as a DaemonSet, Fargate is disqualified and the correct answer flips to managed node groups. The exam is testing whether you hold both the cost intuition and the structural constraints in your head at once.

There is also a governance dimension that connects back to Week 1. EKS clusters are frequently multi-tenant, and the isolation model you choose determines what a tenant can reach. Fargate gives you a per-pod VM boundary, which is a stronger isolation primitive than a shared node with cgroups and namespaces. When a scenario mentions untrusted or mutually suspicious tenants, that boundary becomes an architectural requirement rather than a cost line, and it pushes the answer toward Fargate even when the utilization math argues the other way.

2. How EKS Actually Works Underneath

EKS splits the Kubernetes control plane from the data plane and manages only the former. The API server, etcd, the scheduler, and the controller manager run in an AWS-owned account, spread across multiple Availability Zones, with the etcd volume replicated and the API server endpoints fronted by an elastic network interface that you can restrict with security groups. You never see those instances, never patch them, and never get SSH access. What you get back is a standard Kubernetes API endpoint and a certificate authority, which means every upstream Kubernetes tool — kubectl, Helm, Argo CD, the operator ecosystem — works unmodified.

The data plane is where the choice lives. A node in EKS is an EC2 instance that has been bootstrapped with the kubelet, the container runtime, and the AWS VPC CNI plugin, then joined to the cluster via a bootstrap script that authenticates against the cluster's endpoint. Managed node groups automate that bootstrap: you declare an instance type, a desired/min/max size, subnets, and optionally a launch template, and EKS provisions an Auto Scaling Group behind the scenes, runs the bootstrap, and registers the nodes. Self-managed nodes are the same thing without the automation — you own the Auto Scaling Group, the AMI, the bootstrap, and the lifecycle. Fargate profiles take a third path: instead of joining a node, each pod scheduled into a matching namespace and label selector gets its own micro-VM, with the kubelet and runtime injected by AWS and no node object that you can see or schedule against.

The networking model is the part that most often surprises people coming from ECS. EKS uses the AWS VPC CNI by default, which allocates a real VPC IP address to every pod from the subnet's CIDR range. That is different from overlay networks like Calico or Flannel, where pods get addresses from a private range that is then encapsulated. The practical consequence is that pod IPs are routable inside the VPC, which makes VPC security groups, VPC Flow Logs, and Network Firewall inspection work on pod traffic without any extra plumbing. The cost is IP address consumption: each node pre-allocates a pool of secondary IPs based on its instance type's ENI limits, and a cluster with many small nodes can exhaust a subnet faster than the node count suggests.

Authentication and authorization are split in a way that trips people up. The cluster's IAM role controls who can call the EKS API — creating node groups, updating the cluster — but it does not control what happens inside Kubernetes. Inside the cluster, the aws-auth ConfigMap (or, on newer clusters, EKS access entries) maps IAM principals to Kubernetes users and groups, and then standard Kubernetes RBAC decides what those users can do. A very common production incident is an engineer with full EKS API permissions who cannot run kubectl get pods because nobody added them to aws-auth. The two layers are independent, and the exam does occasionally probe whether you know that.

3. The Core Decision Boundary: Nodes or No Nodes

Every EKS compute question reduces to a single fork: does this workload need to be aware of, or share, a node? If the answer is yes — because it runs a per-node agent, needs a GPU, mounts a host path, requires privileged mode, or depends on a node-level autoscaler — then Fargate is off the table and you are choosing between managed and self-managed node groups. If the answer is no, Fargate becomes viable and the decision shifts to a cost and operational-overhead comparison.

The reason this fork is so clean is that Fargate's isolation is implemented by removing the node from your view entirely. There is no shared kernel between pods, no node filesystem to mount, no device to pass through, and no node object for a DaemonSet to target. That is exactly what makes it a strong security boundary and exactly what makes it incompatible with a large slice of the Kubernetes ecosystem. When a scenario describes a requirement that only makes sense in the presence of a node, you have your answer.

Requirement in the scenarioSelf-managed nodesManaged node groupsFargate profiles
DaemonSet on every nodeYesYesNo — no node to target
GPU / specialized hardwareYesYes (GPU instance types)No
Privileged containers, hostPath mountsYesYesNo
Custom AMI / kernel tuningYesYes (launch template AMI)No
Strong per-pod isolation for untrusted tenantsWeak (shared kernel)Weak (shared kernel)Strong (per-pod micro-VM)
No node patching or capacity planningNoPartialYes
Bin-packing efficiency at high utilizationHighHighLower — per-pod overhead
Spot / mixed capacity with graceful drainManualNativeNot applicable

Read the table as a filter rather than a scorecard. The first four rows are hard disqualifiers for Fargate; the last four are soft preferences that only matter once Fargate is still in the running. A scenario that mentions a DaemonSet and a tight budget is not asking you to weigh cost against isolation — it is telling you the answer is a node group and the cost discussion is a distractor.

4. Configuration Modes and What Each Costs You

Managed node groups are the default recommendation for most production clusters, and the reason is that they remove the two most error-prone parts of running Kubernetes on EC2: the AMI lifecycle and the join process. You specify an instance type and a size range, and EKS handles the Auto Scaling Group, the bootstrap, and — if you opt in — the AMI updates. The managed AMI update path is worth understanding because it is not a rolling replacement of your nodes in the naive sense. EKS creates a new node group with the updated AMI, cordons and drains the old nodes one at a time, and deletes the old group when the migration completes. That drain respects PodDisruptionBudgets, which means a badly configured PDB can stall the update indefinitely.

Self-managed nodes exist for the cases where you need control that managed node groups do not expose. Custom AMIs with hardened kernels, specific instance types that the managed path does not yet support, unusual bootstrap logic, or a requirement to run the nodes in a way that conflicts with the managed lifecycle. The cost is that you now own the AMI pipeline, the bootstrap script, the Auto Scaling Group configuration, and the node upgrade process. In practice, teams that choose self-managed nodes usually do so because of a specific constraint, not because they prefer the operational burden.

Fargate profiles are configured as a namespace plus an optional label selector. Any pod that matches gets scheduled onto Fargate; anything that does not match falls back to node-based scheduling. That selector is the mechanism that lets you run a mixed cluster — a Fargate profile for the stateless web tier and a managed node group for the DaemonSet-dependent observability stack — which is a common and defensible production pattern. The tradeoff is that Fargate pods are billed per vCPU and per GB of memory per second, rounded up, with a minimum billing increment, and the pod spec must fit within the supported CPU/memory combinations. You cannot request an arbitrary 3.5 vCPU; you pick from a fixed set.

There is a fourth mode that the exam sometimes implies without naming: EKS Auto Mode, which shifts node provisioning, scaling, and patching to AWS-managed infrastructure while still giving you nodes. It sits between managed node groups and Fargate in the operational-overhead spectrum. If a scenario asks for "no node management but DaemonSets must work," that is the shape of the answer, and it is worth knowing that the middle ground exists rather than treating the choice as strictly binary.

5. Sizing, Limits, and Quotas

The numbers that matter for EKS compute decisions fall into three buckets: what Fargate will accept, what a node can hold, and what the cluster as a whole will allow. Fargate pod sizing is constrained to a defined set of vCPU and memory pairs — the smallest is a quarter vCPU with half a gigabyte, and the largest is 16 vCPU with 120 GB — and the memory you can request scales with the vCPU you choose. A pod that requests a combination outside that set is rejected at admission, not at runtime, which is a useful thing to know when a scenario describes a pod that "fails to schedule" with no obvious resource pressure.

Node capacity is governed by the instance type's ENI and IP limits, not just its CPU and memory. Because the VPC CNI assigns a real VPC IP to every pod, the number of pods a node can host is capped by how many secondary IP addresses its ENIs can hold. A small instance type may have plenty of spare CPU but run out of pod slots first. This is the single most common cause of "the node has capacity but pods are stuck Pending," and the fix is either a larger instance type or enabling prefix delegation, which lets the CNI allocate /28 prefixes instead of individual IPs and raises the pod ceiling substantially.

DimensionWhat constrains itFirst thing to check
Fargate pod sizeFixed vCPU/memory combinationsIs the requested pair in the supported set?
Pods per nodeENI secondary IP limits for the instance typePrefix delegation enabled? Instance type large enough?
Subnet IP exhaustionVPC CNI warm pool per nodeSubnet CIDR size vs. node count and pod density
Cluster node ceilingService quota per cluster and per RegionQuota dashboard before a large scale-out
Fargate concurrent podsAccount-level quotaRequest an increase ahead of a launch

Quotas deserve a specific mention because they are the classic launch-day failure. EKS has per-cluster and per-Region limits on nodes, and Fargate has an account-level limit on concurrent pods. Neither is visible in the cluster's own metrics, and both are raised by support request rather than automatically. A scenario that describes a planned tenfold scale-out for a product launch is often really asking whether you would pre-emptively request a quota increase, and the answer is yes.

6. Failure Modes and What They Look Like in Production

The most common EKS failure is not a crash — it is a pod that never starts. A pod stuck in Pending with an "Insufficient pods" or "Too many pods" event means the scheduler found a node with CPU and memory but no available pod slots, which points at the CNI IP ceiling rather than at resource pressure. A pod stuck in Pending with "no nodes available to schedule pods" in a Fargate profile usually means the namespace or label selector does not match, so the pod is looking for nodes that do not exist. Both look identical in a dashboard that only shows pod count, which is why the first diagnostic move is always kubectl describe pod and reading the Events section rather than looking at aggregate metrics.

Node-level failures present differently. When an EC2 instance backing a node group is terminated — by a Spot interruption, an unhealthy status check, or an AZ event — the node goes NotReady and the pods on it are rescheduled, but only after the node's taint toleration period expires. That default is five minutes, which means a Spot-heavy node group can have a visible five-minute gap in capacity after every interruption unless you tune the toleration. The symptom is a periodic dip in available replicas that correlates with Spot reclaim events, and the fix is either a shorter toleration or a mixed-instance strategy that keeps enough On-Demand baseline to absorb the gap.

Fargate failures tend to be admission-time rather than runtime. A pod that requests a resource combination outside the supported set, mounts a hostPath, or declares a DaemonSet will fail to schedule with an event that names the constraint. Because there is no node to inspect, the diagnostic surface is smaller — you read the pod events and the Fargate profile configuration, and there is nothing else to look at. That is a genuine advantage operationally, but it also means that when something does go wrong at the infrastructure layer, you have no visibility into it and must wait for AWS.

Control-plane failures are rare but worth recognizing. The EKS control plane is multi-AZ and managed, so an API server outage is an AWS-side event, but you can still cause your own by misconfiguring the cluster security group or the public/private endpoint access settings. A cluster with only private endpoint access and no VPC connectivity from your bastion is unreachable, and the symptom is a kubectl timeout rather than an error. That is a configuration failure, not a service failure, and it is a common self-inflicted wound during initial setup.

7. The Operational and SRE Angle

From an SRE perspective, the compute choice determines what you monitor and what you can automate. On node-based clusters, the metrics that matter are node-level: CPU and memory reservation versus allocation, the ratio of requested to allocatable resources, and the count of pods per node against the CNI ceiling. A cluster that is 40% utilized by actual consumption but 95% reserved by requests is effectively full, and the only way to see that is to track requests and limits separately from usage. Container Insights gives you this, but the default dashboards emphasize usage, so the reservation view has to be built deliberately.

On Fargate, node metrics do not exist, so the monitoring shifts entirely to pod-level and application-level signals. That is simpler in one sense and harder in another: you lose the early-warning signal that a node is filling up, and you gain a billing model where a runaway pod is a cost event rather than a capacity event. The alarm that matters most on Fargate is not CPU utilization but pod count against the account quota, because that is the limit you will hit first during a scale-out.

The runbook shape differs too. A node-based incident runbook starts with identifying the affected node, cordoning it, draining it, and letting the Auto Scaling Group replace it. A Fargate incident runbook has no node step — it starts with the pod events and the profile configuration, and the remediation is usually a spec change rather than an infrastructure action. Teams that run mixed clusters need both runbooks, and the triage step is determining which substrate the affected workload is on before opening either one.

For SLO purposes, the substrate affects your error budget math in a subtle way. Node-based clusters have a capacity headroom that absorbs short bursts without any user-visible impact, which makes them forgiving of traffic spikes. Fargate scales per pod and is bounded by the account quota and the scheduler's reaction time, so a sudden spike can produce a brief period of unschedulable pods that shows up as elevated latency or 5xx before the new pods are ready. If your SLO is tight, that reaction time is part of your error budget whether you planned for it or not.

8. Edge Cases and Exam Gotchas

The single most-tested gotcha is the DaemonSet incompatibility, and it is worth stating precisely: Fargate pods run in isolated micro-VMs with no shared node, so any workload that requires a persistent per-node agent cannot run there. That covers log shippers, node exporters, service meshes that install a per-node component, and security agents. The exam will describe the requirement in workload terms — "collect logs from every node" — rather than naming DaemonSets, so the translation step is the actual test.

The second gotcha is the cost inversion. Fargate is more expensive per vCPU-hour than an equivalent EC2 instance, but a node group that runs at 15% average utilization is paying for 85% idle capacity. For spiky, low-duty-cycle workloads, Fargate can be cheaper in absolute terms despite the higher unit price. The exam rarely asks you to compute this, but it does ask you to recognize that "Fargate is more expensive" is not a universally valid reason to reject it.

The third is the aws-auth versus IAM distinction. Having eks:* in an IAM policy does not grant you any Kubernetes RBAC permissions. A scenario where an administrator "cannot access the cluster" despite full IAM permissions is almost always an aws-auth or access-entry mapping problem, and the fix is on the Kubernetes side, not the IAM side.

The fourth is the mixed-cluster pattern. Fargate profiles and node groups coexist in the same cluster, and the profile's namespace and label selector determine which pods land where. A pod that matches no profile and has no node selector will land on nodes; a pod that matches a profile will land on Fargate even if nodes have capacity. Getting the selector wrong is a common cause of "why is this pod on Fargate when I have nodes?"

The fifth is the pod-slot ceiling. A node with free CPU and memory can still refuse to schedule pods because the CNI has exhausted its allocatable IPs. If a scenario describes pods stuck Pending on a node that appears to have capacity, the answer is almost always the ENI/IP limit or prefix delegation, not a resource request problem.

9. EKS Compute Options vs. the Alternatives

The comparison that matters most is EKS against ECS, because the exam frequently offers both as answers to the same scenario. The deciding factor is almost never technical capability — both run containers well — but organizational. If the team already runs Kubernetes elsewhere, has existing manifests and Helm charts, or needs portability across clouds, EKS preserves that investment. If the team is AWS-native and wants the lowest operational overhead, ECS with Fargate is simpler and has fewer moving parts. Choosing EKS for a team with no Kubernetes experience is a common wrong answer.

Within EKS, the three compute options are not mutually exclusive and the best production answer is often a combination. The table below is the decision rule in compact form.

OptionPick it when…Avoid it when…
Managed node groupsYou need DaemonSets, GPUs, custom AMIs, or high bin-packing efficiency, and want AWS to handle the AMI lifecycleYou need kernel-level control or an instance type the managed path does not support
Self-managed nodesYou need a hardened custom AMI, unusual bootstrap logic, or a specific instance typeYou have no capacity to own the AMI pipeline and node upgrade process
Fargate profilesPods are stateless, need strong isolation, and the workload is spiky or low-duty-cycleAny pod needs a node-level agent, GPU, hostPath, or privileged mode
ECS on FargateThe team is AWS-native, has no Kubernetes investment, and wants minimal operational surfaceYou need Kubernetes APIs, CRDs, or multi-cloud portability

The rule of thumb worth carrying into the exam: start from the workload's structural requirements, not from its cost profile. If nothing in the description requires a node, Fargate is a legitimate answer and the cost discussion becomes a tiebreaker. If anything in the description requires a node, the cost discussion is irrelevant and the answer is a node group. The exam's distractors are built to make you skip that first step.

Hands-On Lab: Choosing a Substrate for a Bursty Batch Namespace

The scenario is a namespace running batch jobs that arrive unpredictably — some hours it is idle, some hours it runs forty concurrent jobs, and the jobs are short-lived and stateless. The team has no interest in patching nodes. Your job is to determine empirically whether a managed node group or a Fargate profile is the better fit, and to find the point at which each one breaks.

1. Establish the baseline cluster. Create an EKS cluster with one managed node group of a modest instance type, sized to a minimum of two nodes across two Availability Zones. Confirm the cluster is healthy with kubectl get nodes and note the allocatable pod count per node from kubectl describe node. That number is your ceiling for the node-based path, and it is almost always lower than the instance's CPU would suggest.

2. Deploy the batch workload on nodes. Apply a Job manifest with a modest CPU request and a parallelism of ten. Watch the pods schedule and record how long from apply to all pods Running. Then scale parallelism to forty and repeat. Note whether the cluster autoscaler adds nodes, how long that takes, and whether any pods sit Pending during the ramp. The time-to-capacity figure you record here is the number that matters for the comparison.

3. Create a Fargate profile for the namespace. Add a profile matching the batch namespace with a label selector that matches the job pods. Confirm that new pods land on Fargate by checking that they have no node name and that kubectl get nodes shows no new nodes. Re-run the same parallelism ramp and record the same time-to-capacity figure.

4. Deliberately break Fargate. Add a DaemonSet to the namespace and observe what happens — it will not schedule, and the event will name the constraint. Then modify the Job to request a CPU/memory combination outside the supported Fargate set and observe the admission failure. Record both error messages verbatim; they are the diagnostic signatures you will recognize in production.

5. Deliberately break the node path. On the node-based namespace, deploy enough pods to exhaust the CNI IP allocation on a single node while leaving CPU and memory free. Observe the Pending state and the event message. Then enable prefix delegation on the VPC CNI add-on and re-test to confirm the ceiling moves.

6. Compare cost and record the crossover. Using the utilization you observed — which for a bursty batch workload will be low on average — estimate the monthly cost of the node group at its minimum size versus the Fargate cost for the same total job-seconds. Write down the duty cycle at which the two cross over. That number is the answer to the "is Fargate more expensive" question for this specific workload, and it is usually lower than people expect.

7. Write the decision memo. In one page, state which substrate you would choose for this namespace and why, citing the time-to-capacity figures, the structural constraints you hit, and the crossover duty cycle. Include the mixed-cluster option explicitly: a Fargate profile for the batch namespace plus a small managed node group for cluster-level DaemonSets is a legitimate answer, and the memo should say why you did or did not choose it.

Scenario Question Drills

Q1. A platform team runs a log-collection agent as a DaemonSet across an EKS cluster. They want to move the application tier to Fargate to reduce node management. What happens to the DaemonSet?

A. It runs normally on the Fargate micro-VMs
B. It fails to schedule, because Fargate pods have no shared node for a per-node agent to target
C. It converts automatically into a sidecar in each pod
D. It runs only on the control plane nodes
Correct answer: B. Fargate pods run in isolated micro-VMs with no shared node, so DaemonSets — which require a persistent per-node agent — cannot be scheduled. The application tier can move to Fargate, but the DaemonSet must stay on a node group.

Q2. A team has deep existing Kubernetes tooling, Helm charts, and a requirement to run the same manifests on another cloud. Which AWS compute platform best fits?

A. Amazon ECS with the EC2 launch type
B. Amazon EKS
C. AWS Lambda with container images
D. AWS Batch on Fargate
Correct answer: B. EKS runs standard upstream Kubernetes, so existing manifests, Helm charts, and CRDs work unmodified and the same artifacts are portable to other Kubernetes environments. ECS is AWS-proprietary and would require rewriting the deployment tooling.

Q3. Pods are stuck in Pending on a node that shows plenty of free CPU and memory. The scheduler event reads "Too many pods." What is the most likely cause?

A. The node's kubelet has crashed
B. The VPC CNI has exhausted the node's allocatable secondary IP addresses, capping pod slots below what CPU and memory would allow
C. The pods have an incorrect nodeSelector
D. The cluster's service quota for nodes has been reached
Correct answer: B. Because the AWS VPC CNI assigns a real VPC IP to every pod, the number of pods a node can host is capped by its ENI secondary IP limits. Enabling prefix delegation raises that ceiling substantially.

Q4. A workload runs at roughly 12% average CPU utilization with sharp, unpredictable spikes. The team has no capacity to patch nodes. Which substrate is most likely to be both operationally and financially appropriate?

A. A large managed node group sized for peak
B. EKS Fargate profiles, since per-pod billing avoids paying for idle node capacity
C. Self-managed nodes with Reserved Instances
D. A single large EC2 instance running kubeadm
Correct answer: B. At low duty cycle, a node group sized for peak pays for idle capacity most of the time. Fargate's higher unit price is offset by billing only for what runs, and it removes node patching entirely — provided no pod needs node-level features.

Q5. An engineer has an IAM policy granting eks:* on all resources but receives "error: You must be logged in to the server (Unauthorized)" from kubectl. What is the most likely cause?

A. The cluster's security group is blocking port 443
B. The IAM principal is not mapped to a Kubernetes user or group in aws-auth (or via EKS access entries), so Kubernetes RBAC denies the request
C. The kubeconfig is pointing at the wrong Region
D. The EKS control plane is down
Correct answer: B. IAM controls access to the EKS API; Kubernetes RBAC controls what happens inside the cluster. The two layers are independent, and an unmapped IAM principal has no Kubernetes identity.

Q6. A machine-learning team needs GPU-backed pods in an EKS cluster. Which compute option can satisfy this?

A. EKS Fargate profiles
B. Managed node groups with GPU instance types
C. Either option works identically
D. Neither — EKS does not support GPUs
Correct answer: B. Fargate does not expose GPUs or other specialized hardware. GPU workloads require node-based compute, typically a managed node group using GPU instance types with the appropriate device plugin.

Q7. A cluster runs a Fargate profile for the web namespace and a managed node group for everything else. A pod in the web namespace has no matching label selector for the profile. Where does it run?

A. It fails to schedule
B. It falls back to the node group, since Fargate profiles only claim pods that match both namespace and selector
C. It runs on Fargate regardless of the selector
D. It runs on the control plane
Correct answer: B. A Fargate profile claims a pod only when both the namespace and the label selector match. Non-matching pods fall through to node-based scheduling, which is what makes mixed clusters work.

Q8. A Spot-heavy managed node group shows a recurring five-minute dip in available replicas that correlates with Spot reclaim events. What is the most likely explanation?

A. The Auto Scaling Group is misconfigured
B. Pods are waiting out the default node NotReady toleration period before being rescheduled
C. The load balancer health checks are too aggressive
D. The VPC CNI is leaking IP addresses
Correct answer: B. When a node goes NotReady, pods wait out the toleration period (five minutes by default) before being evicted and rescheduled. Tuning the toleration or keeping an On-Demand baseline absorbs the gap.

Q9. A regulated workload requires that no two tenants share a kernel. Which EKS compute option provides the strongest isolation boundary?

A. Managed node groups with pod security policies
B. EKS Fargate profiles, where each pod runs in its own micro-VM
C. Self-managed nodes with namespaces
D. Network policies alone
Correct answer: B. Fargate gives each pod its own micro-VM, so there is no shared kernel between tenants. Node-based options share a kernel across all pods on the node, which namespaces and network policies do not change.

Q10. A pod spec requests 3.5 vCPU and 7 GB of memory and is rejected before it ever runs. Why?

A. The cluster has insufficient capacity
B. Fargate only accepts a fixed set of vCPU/memory combinations, and that pair is not one of them
C. The pod needs a node selector
D. The IAM task role is missing
Correct answer: B. Fargate pod sizing is constrained to a defined set of vCPU and memory pairs. A request outside that set is rejected at admission rather than failing at runtime.

Q11. A team wants AWS to handle node provisioning, scaling, and patching but still needs DaemonSets to work. Which option fits?

A. EKS Fargate profiles
B. EKS Auto Mode, which manages nodes on your behalf while still providing node objects
C. Self-managed nodes with a custom AMI
D. ECS on Fargate
Correct answer: B. EKS Auto Mode shifts node provisioning, scaling, and patching to AWS-managed infrastructure while still giving you nodes, so DaemonSets continue to work. Fargate cannot satisfy the DaemonSet requirement.

Q12. A product launch is expected to increase pod count tenfold in a single Region. What should be done before the launch?

A. Nothing — EKS scales automatically without limits
B. Request service quota increases for cluster nodes and, if using Fargate, for concurrent Fargate pods
C. Switch the cluster to a single Availability Zone
D. Disable the cluster autoscaler
Correct answer: B. EKS has per-cluster and per-Region node quotas, and Fargate has an account-level concurrent-pod quota. Neither is visible in cluster metrics and both require a support request to raise.

Q13. A managed node group AMI update appears to hang indefinitely, with old nodes never draining. What is the most likely cause?

A. The new AMI is corrupt
B. A PodDisruptionBudget is preventing the drain from evicting pods
C. The cluster's IAM role lacks permissions
D. The VPC CNI is misconfigured
Correct answer: B. The managed AMI update path cordons and drains nodes one at a time, respecting PodDisruptionBudgets. A PDB that cannot be satisfied will stall the drain indefinitely.

Q14. A team is AWS-native, has no Kubernetes experience, and wants the lowest possible operational overhead for a containerized web application. Which choice is most appropriate?

A. Amazon EKS with managed node groups
B. Amazon ECS on Fargate
C. Self-managed Kubernetes on EC2
D. Amazon EKS with Fargate profiles
Correct answer: B. With no existing Kubernetes investment, ECS on Fargate has fewer moving parts and a smaller operational surface than EKS. Choosing EKS for a team without Kubernetes experience is a common wrong answer.

Q15. A cluster is configured with private endpoint access only, and an engineer's kubectl commands time out rather than returning an authorization error. What does this indicate?

A. The engineer lacks Kubernetes RBAC permissions
B. The engineer has no network path to the private API endpoint, which is a connectivity problem rather than an authorization one
C. The control plane is in a different Region
D. The kubeconfig certificate has expired
Correct answer: B. A timeout rather than an Unauthorized error points at network reachability. With private-only endpoint access, the client must be inside the VPC or connected via VPN/Direct Connect to reach the API server.

Peek into Tomorrow

Everything in this day assumed that capacity appears when you ask for it. A node group scales out, a Fargate pod schedules, and the workload starts. What we have not examined is the gap between the moment a scaling decision is made and the moment the new capacity can actually serve traffic — and on node-based compute that gap is not small. An EC2 instance has to boot, the kubelet has to join the cluster, the container runtime has to pull images, and the application has to initialize before it is useful. For a workload with a four-minute bootstrap, a scale-out triggered by a traffic spike arrives four minutes after the spike, which is often too late to matter.

Tomorrow's material addresses that gap directly, and it does so with two mechanisms that are easy to confuse. One pauses an instance at a defined point in its lifecycle so a script can run — registering it with a load balancer on the way in, or draining connections on the way out — which is about correctness rather than speed. The other keeps a pool of instances already past the expensive part of startup, so scale-out promotes something that is nearly ready instead of beginning from a cold boot. The open question worth carrying forward is which of those two you reach for when the problem is latency rather than correctness, and what the tradeoff is when you keep pre-initialized capacity sitting idle.

Sources