Amazon EKS Architecture & Managed Node Groups vs Fargate
Recap: Where We Left Off
Day 15 established that Fargate on ECS removes the EC2 layer entirely: a task definition declares CPU, memory, and an IAM task role, and the scheduler places each task into its own isolated compute allocation. The awsvpc network mode then gives every task its own ENI and security group, which is what makes per-task network policy meaningful rather than decorative. That model is clean, and it is also the source of the trap this day sets up in a different service.
The trap is that "serverless containers" is a phrase that travels across services without carrying its constraints with it. On ECS, Fargate's isolation is mostly a gift — you get per-task security groups and no host patching. On EKS, the same launch type sits underneath a Kubernetes control plane that assumes nodes exist, and a large fraction of the Kubernetes ecosystem is built on that assumption. DaemonSets, host-path volumes, node-level autoscaling, and node affinity all quietly stop working. The exam does not ask you to recite that Fargate is serverless; it asks you to notice when a stated requirement — a log agent on every node, a GPU, a privileged container — makes Fargate structurally impossible. Same trap, different service.
Foundations You'll Need Today
Today's material sits on top of four ideas that the rest of this curriculum treats as background knowledge. If any of them are new, read this section first — the rest of the day will make far more sense with it in place.
Containers, and Why They Are Not Virtual Machines
A virtual machine is a complete simulated computer. It has its own operating system kernel, its own boot process, and its own copy of everything the OS needs. That completeness is what makes VMs portable and isolated, and it is also what makes them heavy: you pay for a full OS boot and a full OS's worth of memory before your application even starts. A container takes a different approach. Instead of simulating a computer, it uses features built into the host's existing operating system to fence off a process — or a small group of processes — so that it sees only its own files, its own network, and its own view of the system. The application inside a container believes it is alone on a machine, but it is actually sharing the host's kernel with every other container on that host.
That sharing is the whole point and the whole risk. Because there is no second kernel to boot, a container starts in a fraction of a second and costs almost nothing in overhead, which is why you can run dozens of them on a machine that would struggle to host three VMs. But "sharing a kernel" also means that a flaw in the kernel is a flaw shared by every container on that host. When this day talks about a "strong isolation boundary" or "no two tenants share a kernel," that is the distinction being drawn: containers on the same machine are separated by software rules, while separate machines are separated by hardware. Keep that in mind, because it is the reason one of today's compute options is chosen for security rather than for cost.
Kubernetes: Control Plane, Data Plane, and Nodes
Kubernetes is a system for running containers across many machines and keeping them running. It is split into two halves that do very different jobs. The control plane is the brain: it stores the desired state of everything you have asked for, decides where each container should run, and watches constantly to make reality match your description. The data plane is the muscle: the actual machines where containers execute. In Kubernetes, those machines are called nodes, and each node runs a small agent called the kubelet whose job is to receive instructions from the control plane and start or stop containers accordingly.
Two more terms appear throughout today's material. The scheduler is the part of the control plane that picks which node a given container should land on, based on how much CPU and memory it asked for and what the nodes have free. A DaemonSet is a special kind of workload that says "run one copy of this container on every node, always." Log collectors, monitoring agents, and security scanners are usually deployed this way, because they need to be present on every machine rather than running as a single shared service. Hold onto that definition — the fact that a DaemonSet needs a node to attach to is the single most important constraint in today's decision.
VPC Networking: Subnets, IP Addresses, and Elastic Network Interfaces
A Virtual Private Cloud (VPC) is your own private network inside AWS. It is divided into subnets, each of which owns a range of IP addresses — for example, a subnet might own every address from 10.0.1.0 through 10.0.1.255. Anything that needs to communicate on the network needs an address from one of those ranges, and the number of addresses a subnet owns is finite. This matters more than it sounds, because running out of addresses in a subnet is a real and common failure: new machines simply cannot be created until space is freed.
An Elastic Network Interface (ENI) is the virtual equivalent of a network card. Every EC2 instance has at least one, and each ENI holds one primary IP address plus a number of additional addresses it is allowed to carry. That number is not unlimited — it depends on the size of the instance, and larger instances get more. This is the detail that makes today's pod-scheduling gotcha possible: if every container needs its own real IP address, and each machine can only hold so many addresses, then a machine can run out of addresses long before it runs out of CPU or memory. A security group, for completeness, is a set of firewall rules attached to an ENI that decides which network traffic is allowed in and out. It is the AWS equivalent of a firewall rule set, and it operates per network interface rather than per machine.
IAM Roles vs. Kubernetes Permissions
AWS Identity and Access Management (IAM) is how AWS decides who is allowed to call which AWS APIs. An IAM role is a set of permissions that can be assumed by a person, an application, or a service, and it is the standard way to grant access without handing out long-lived passwords. When you hear that someone has "full EKS permissions," it means their IAM role allows them to call the AWS APIs that create, modify, and delete EKS clusters.
Kubernetes has its own, entirely separate permission system called RBAC (role-based access control), which decides what a user is allowed to do inside a cluster — list pods, read logs, delete deployments. These two systems do not talk to each other automatically. Having permission to call the AWS API that manages a cluster does not grant you any ability to interact with the workloads running in it, and vice versa. Something has to explicitly connect an AWS identity to a Kubernetes identity, and today's material covers what happens when that connection is missing. It is a frequent source of confusion precisely because both systems use the word "role" and both are described as "permissions."
With that grounding, here's why EKS exists, what problem it actually solves, and why the choice of where your containers run is a genuine architectural decision rather than a configuration detail.
1. Why EKS Compute Choice Is on the Exam
EKS appears on SAP-C02 primarily as a decision problem rather than a configuration problem. The exam rarely asks you to write a manifest or recall a kubectl flag. It gives you a workload description — a team with existing Kubernetes tooling, a batch job with unpredictable arrival, a compliance requirement that no two tenants share a kernel — and asks which compute substrate you would place it on. The four answers are almost always self-managed nodes, managed node groups, Fargate profiles, and "not EKS at all, use ECS."
That framing maps to Domain 2 (Design Resilient Architectures) and Domain 3 (Design High-Performing Architectures), with a recurring cost thread from Domain 4. The resilience angle is about what happens when a node dies and how much of your capacity is self-healing. The performance angle is about scheduling latency, bin-packing efficiency, and whether your pods can actually get the resources they request. The cost angle is the one candidates most often get wrong: Fargate is frequently described as "more expensive," which is true per-vCPU-hour but not necessarily true per-workload once you account for the idle capacity that node groups carry.
The reason this is worth a full day rather than a paragraph is that the wrong answer is usually plausible. A scenario describing a small, bursty, low-utilization workload with no operational team reads like a Fargate scenario, and often is — but if the same scenario mentions a sidecar log shipper deployed as a DaemonSet, Fargate is disqualified and the correct answer flips to managed node groups. The exam is testing whether you hold both the cost intuition and the structural constraints in your head at once.
There is also a governance dimension that connects back to Week 1. EKS clusters are frequently multi-tenant, and the isolation model you choose determines what a tenant can reach. Fargate gives you a per-pod VM boundary, which is a stronger isolation primitive than a shared node with cgroups and namespaces. When a scenario mentions untrusted or mutually suspicious tenants, that boundary becomes an architectural requirement rather than a cost line, and it pushes the answer toward Fargate even when the utilization math argues the other way.
2. How EKS Actually Works Underneath
EKS splits the Kubernetes control plane from the data plane and manages only the former. The API server, etcd, the scheduler, and the controller manager run in an AWS-owned account, spread across multiple Availability Zones, with the etcd volume replicated and the API server endpoints fronted by an elastic network interface that you can restrict with security groups. You never see those instances, never patch them, and never get SSH access. What you get back is a standard Kubernetes API endpoint and a certificate authority, which means every upstream Kubernetes tool — kubectl, Helm, Argo CD, the operator ecosystem — works unmodified.
The data plane is where the choice lives. A node in EKS is an EC2 instance that has been bootstrapped with the kubelet, the container runtime, and the AWS VPC CNI plugin, then joined to the cluster via a bootstrap script that authenticates against the cluster's endpoint. Managed node groups automate that bootstrap: you declare an instance type, a desired/min/max size, subnets, and optionally a launch template, and EKS provisions an Auto Scaling Group behind the scenes, runs the bootstrap, and registers the nodes. Self-managed nodes are the same thing without the automation — you own the Auto Scaling Group, the AMI, the bootstrap, and the lifecycle. Fargate profiles take a third path: instead of joining a node, each pod scheduled into a matching namespace and label selector gets its own micro-VM, with the kubelet and runtime injected by AWS and no node object that you can see or schedule against.
The networking model is the part that most often surprises people coming from ECS. EKS uses the AWS VPC CNI by default, which allocates a real VPC IP address to every pod from the subnet's CIDR range. That is different from overlay networks like Calico or Flannel, where pods get addresses from a private range that is then encapsulated. The practical consequence is that pod IPs are routable inside the VPC, which makes VPC security groups, VPC Flow Logs, and Network Firewall inspection work on pod traffic without any extra plumbing. The cost is IP address consumption: each node pre-allocates a pool of secondary IPs based on its instance type's ENI limits, and a cluster with many small nodes can exhaust a subnet faster than the node count suggests.
Authentication and authorization are split in a way that trips people up. The cluster's IAM role controls who can call the EKS API — creating node groups, updating the cluster — but it does not control what happens inside Kubernetes. Inside the cluster, the aws-auth ConfigMap (or, on newer clusters, EKS access entries) maps IAM principals to Kubernetes users and groups, and then standard Kubernetes RBAC decides what those users can do. A very common production incident is an engineer with full EKS API permissions who cannot run kubectl get pods because nobody added them to aws-auth. The two layers are independent, and the exam does occasionally probe whether you know that.
3. The Core Decision Boundary: Nodes or No Nodes
Every EKS compute question reduces to a single fork: does this workload need to be aware of, or share, a node? If the answer is yes — because it runs a per-node agent, needs a GPU, mounts a host path, requires privileged mode, or depends on a node-level autoscaler — then Fargate is off the table and you are choosing between managed and self-managed node groups. If the answer is no, Fargate becomes viable and the decision shifts to a cost and operational-overhead comparison.
The reason this fork is so clean is that Fargate's isolation is implemented by removing the node from your view entirely. There is no shared kernel between pods, no node filesystem to mount, no device to pass through, and no node object for a DaemonSet to target. That is exactly what makes it a strong security boundary and exactly what makes it incompatible with a large slice of the Kubernetes ecosystem. When a scenario describes a requirement that only makes sense in the presence of a node, you have your answer.
| Requirement in the scenario | Self-managed nodes | Managed node groups | Fargate profiles |
|---|---|---|---|
| DaemonSet on every node | Yes | Yes | No — no node to target |
| GPU / specialized hardware | Yes | Yes (GPU instance types) | No |
| Privileged containers, hostPath mounts | Yes | Yes | No |
| Custom AMI / kernel tuning | Yes | Yes (launch template AMI) | No |
| Strong per-pod isolation for untrusted tenants | Weak (shared kernel) | Weak (shared kernel) | Strong (per-pod micro-VM) |
| No node patching or capacity planning | No | Partial | Yes |
| Bin-packing efficiency at high utilization | High | High | Lower — per-pod overhead |
| Spot / mixed capacity with graceful drain | Manual | Native | Not applicable |
Read the table as a filter rather than a scorecard. The first four rows are hard disqualifiers for Fargate; the last four are soft preferences that only matter once Fargate is still in the running. A scenario that mentions a DaemonSet and a tight budget is not asking you to weigh cost against isolation — it is telling you the answer is a node group and the cost discussion is a distractor.
4. Configuration Modes and What Each Costs You
Managed node groups are the default recommendation for most production clusters, and the reason is that they remove the two most error-prone parts of running Kubernetes on EC2: the AMI lifecycle and the join process. You specify an instance type and a size range, and EKS handles the Auto Scaling Group, the bootstrap, and — if you opt in — the AMI updates. The managed AMI update path is worth understanding because it is not a rolling replacement of your nodes in the naive sense. EKS creates a new node group with the updated AMI, cordons and drains the old nodes one at a time, and deletes the old group when the migration completes. That drain respects PodDisruptionBudgets, which means a badly configured PDB can stall the update indefinitely.
Self-managed nodes exist for the cases where you need control that managed node groups do not expose. Custom AMIs with hardened kernels, specific instance types that the managed path does not yet support, unusual bootstrap logic, or a requirement to run the nodes in a way that conflicts with the managed lifecycle. The cost is that you now own the AMI pipeline, the bootstrap script, the Auto Scaling Group configuration, and the node upgrade process. In practice, teams that choose self-managed nodes usually do so because of a specific constraint, not because they prefer the operational burden.
Fargate profiles are configured as a namespace plus an optional label selector. Any pod that matches gets scheduled onto Fargate; anything that does not match falls back to node-based scheduling. That selector is the mechanism that lets you run a mixed cluster — a Fargate profile for the stateless web tier and a managed node group for the DaemonSet-dependent observability stack — which is a common and defensible production pattern. The tradeoff is that Fargate pods are billed per vCPU and per GB of memory per second, rounded up, with a minimum billing increment, and the pod spec must fit within the supported CPU/memory combinations. You cannot request an arbitrary 3.5 vCPU; you pick from a fixed set.
There is a fourth mode that the exam sometimes implies without naming: EKS Auto Mode, which shifts node provisioning, scaling, and patching to AWS-managed infrastructure while still giving you nodes. It sits between managed node groups and Fargate in the operational-overhead spectrum. If a scenario asks for "no node management but DaemonSets must work," that is the shape of the answer, and it is worth knowing that the middle ground exists rather than treating the choice as strictly binary.
5. Sizing, Limits, and Quotas
The numbers that matter for EKS compute decisions fall into three buckets: what Fargate will accept, what a node can hold, and what the cluster as a whole will allow. Fargate pod sizing is constrained to a defined set of vCPU and memory pairs — the smallest is a quarter vCPU with half a gigabyte, and the largest is 16 vCPU with 120 GB — and the memory you can request scales with the vCPU you choose. A pod that requests a combination outside that set is rejected at admission, not at runtime, which is a useful thing to know when a scenario describes a pod that "fails to schedule" with no obvious resource pressure.
Node capacity is governed by the instance type's ENI and IP limits, not just its CPU and memory. Because the VPC CNI assigns a real VPC IP to every pod, the number of pods a node can host is capped by how many secondary IP addresses its ENIs can hold. A small instance type may have plenty of spare CPU but run out of pod slots first. This is the single most common cause of "the node has capacity but pods are stuck Pending," and the fix is either a larger instance type or enabling prefix delegation, which lets the CNI allocate /28 prefixes instead of individual IPs and raises the pod ceiling substantially.
| Dimension | What constrains it | First thing to check |
|---|---|---|
| Fargate pod size | Fixed vCPU/memory combinations | Is the requested pair in the supported set? |
| Pods per node | ENI secondary IP limits for the instance type | Prefix delegation enabled? Instance type large enough? |
| Subnet IP exhaustion | VPC CNI warm pool per node | Subnet CIDR size vs. node count and pod density |
| Cluster node ceiling | Service quota per cluster and per Region | Quota dashboard before a large scale-out |
| Fargate concurrent pods | Account-level quota | Request an increase ahead of a launch |
Quotas deserve a specific mention because they are the classic launch-day failure. EKS has per-cluster and per-Region limits on nodes, and Fargate has an account-level limit on concurrent pods. Neither is visible in the cluster's own metrics, and both are raised by support request rather than automatically. A scenario that describes a planned tenfold scale-out for a product launch is often really asking whether you would pre-emptively request a quota increase, and the answer is yes.
6. Failure Modes and What They Look Like in Production
The most common EKS failure is not a crash — it is a pod that never starts. A pod stuck in Pending with an "Insufficient pods" or "Too many pods" event means the scheduler found a node with CPU and memory but no available pod slots, which points at the CNI IP ceiling rather than at resource pressure. A pod stuck in Pending with "no nodes available to schedule pods" in a Fargate profile usually means the namespace or label selector does not match, so the pod is looking for nodes that do not exist. Both look identical in a dashboard that only shows pod count, which is why the first diagnostic move is always kubectl describe pod and reading the Events section rather than looking at aggregate metrics.
Node-level failures present differently. When an EC2 instance backing a node group is terminated — by a Spot interruption, an unhealthy status check, or an AZ event — the node goes NotReady and the pods on it are rescheduled, but only after the node's taint toleration period expires. That default is five minutes, which means a Spot-heavy node group can have a visible five-minute gap in capacity after every interruption unless you tune the toleration. The symptom is a periodic dip in available replicas that correlates with Spot reclaim events, and the fix is either a shorter toleration or a mixed-instance strategy that keeps enough On-Demand baseline to absorb the gap.
Fargate failures tend to be admission-time rather than runtime. A pod that requests a resource combination outside the supported set, mounts a hostPath, or declares a DaemonSet will fail to schedule with an event that names the constraint. Because there is no node to inspect, the diagnostic surface is smaller — you read the pod events and the Fargate profile configuration, and there is nothing else to look at. That is a genuine advantage operationally, but it also means that when something does go wrong at the infrastructure layer, you have no visibility into it and must wait for AWS.
Control-plane failures are rare but worth recognizing. The EKS control plane is multi-AZ and managed, so an API server outage is an AWS-side event, but you can still cause your own by misconfiguring the cluster security group or the public/private endpoint access settings. A cluster with only private endpoint access and no VPC connectivity from your bastion is unreachable, and the symptom is a kubectl timeout rather than an error. That is a configuration failure, not a service failure, and it is a common self-inflicted wound during initial setup.
7. The Operational and SRE Angle
From an SRE perspective, the compute choice determines what you monitor and what you can automate. On node-based clusters, the metrics that matter are node-level: CPU and memory reservation versus allocation, the ratio of requested to allocatable resources, and the count of pods per node against the CNI ceiling. A cluster that is 40% utilized by actual consumption but 95% reserved by requests is effectively full, and the only way to see that is to track requests and limits separately from usage. Container Insights gives you this, but the default dashboards emphasize usage, so the reservation view has to be built deliberately.
On Fargate, node metrics do not exist, so the monitoring shifts entirely to pod-level and application-level signals. That is simpler in one sense and harder in another: you lose the early-warning signal that a node is filling up, and you gain a billing model where a runaway pod is a cost event rather than a capacity event. The alarm that matters most on Fargate is not CPU utilization but pod count against the account quota, because that is the limit you will hit first during a scale-out.
The runbook shape differs too. A node-based incident runbook starts with identifying the affected node, cordoning it, draining it, and letting the Auto Scaling Group replace it. A Fargate incident runbook has no node step — it starts with the pod events and the profile configuration, and the remediation is usually a spec change rather than an infrastructure action. Teams that run mixed clusters need both runbooks, and the triage step is determining which substrate the affected workload is on before opening either one.
For SLO purposes, the substrate affects your error budget math in a subtle way. Node-based clusters have a capacity headroom that absorbs short bursts without any user-visible impact, which makes them forgiving of traffic spikes. Fargate scales per pod and is bounded by the account quota and the scheduler's reaction time, so a sudden spike can produce a brief period of unschedulable pods that shows up as elevated latency or 5xx before the new pods are ready. If your SLO is tight, that reaction time is part of your error budget whether you planned for it or not.
8. Edge Cases and Exam Gotchas
The single most-tested gotcha is the DaemonSet incompatibility, and it is worth stating precisely: Fargate pods run in isolated micro-VMs with no shared node, so any workload that requires a persistent per-node agent cannot run there. That covers log shippers, node exporters, service meshes that install a per-node component, and security agents. The exam will describe the requirement in workload terms — "collect logs from every node" — rather than naming DaemonSets, so the translation step is the actual test.
The second gotcha is the cost inversion. Fargate is more expensive per vCPU-hour than an equivalent EC2 instance, but a node group that runs at 15% average utilization is paying for 85% idle capacity. For spiky, low-duty-cycle workloads, Fargate can be cheaper in absolute terms despite the higher unit price. The exam rarely asks you to compute this, but it does ask you to recognize that "Fargate is more expensive" is not a universally valid reason to reject it.
The third is the aws-auth versus IAM distinction. Having eks:* in an IAM policy does not grant you any Kubernetes RBAC permissions. A scenario where an administrator "cannot access the cluster" despite full IAM permissions is almost always an aws-auth or access-entry mapping problem, and the fix is on the Kubernetes side, not the IAM side.
The fourth is the mixed-cluster pattern. Fargate profiles and node groups coexist in the same cluster, and the profile's namespace and label selector determine which pods land where. A pod that matches no profile and has no node selector will land on nodes; a pod that matches a profile will land on Fargate even if nodes have capacity. Getting the selector wrong is a common cause of "why is this pod on Fargate when I have nodes?"
The fifth is the pod-slot ceiling. A node with free CPU and memory can still refuse to schedule pods because the CNI has exhausted its allocatable IPs. If a scenario describes pods stuck Pending on a node that appears to have capacity, the answer is almost always the ENI/IP limit or prefix delegation, not a resource request problem.
9. EKS Compute Options vs. the Alternatives
The comparison that matters most is EKS against ECS, because the exam frequently offers both as answers to the same scenario. The deciding factor is almost never technical capability — both run containers well — but organizational. If the team already runs Kubernetes elsewhere, has existing manifests and Helm charts, or needs portability across clouds, EKS preserves that investment. If the team is AWS-native and wants the lowest operational overhead, ECS with Fargate is simpler and has fewer moving parts. Choosing EKS for a team with no Kubernetes experience is a common wrong answer.
Within EKS, the three compute options are not mutually exclusive and the best production answer is often a combination. The table below is the decision rule in compact form.
| Option | Pick it when… | Avoid it when… |
|---|---|---|
| Managed node groups | You need DaemonSets, GPUs, custom AMIs, or high bin-packing efficiency, and want AWS to handle the AMI lifecycle | You need kernel-level control or an instance type the managed path does not support |
| Self-managed nodes | You need a hardened custom AMI, unusual bootstrap logic, or a specific instance type | You have no capacity to own the AMI pipeline and node upgrade process |
| Fargate profiles | Pods are stateless, need strong isolation, and the workload is spiky or low-duty-cycle | Any pod needs a node-level agent, GPU, hostPath, or privileged mode |
| ECS on Fargate | The team is AWS-native, has no Kubernetes investment, and wants minimal operational surface | You need Kubernetes APIs, CRDs, or multi-cloud portability |
The rule of thumb worth carrying into the exam: start from the workload's structural requirements, not from its cost profile. If nothing in the description requires a node, Fargate is a legitimate answer and the cost discussion becomes a tiebreaker. If anything in the description requires a node, the cost discussion is irrelevant and the answer is a node group. The exam's distractors are built to make you skip that first step.
Hands-On Lab: Choosing a Substrate for a Bursty Batch Namespace
The scenario is a namespace running batch jobs that arrive unpredictably — some hours it is idle, some hours it runs forty concurrent jobs, and the jobs are short-lived and stateless. The team has no interest in patching nodes. Your job is to determine empirically whether a managed node group or a Fargate profile is the better fit, and to find the point at which each one breaks.
1. Establish the baseline cluster. Create an EKS cluster with one managed node group of a modest instance type, sized to a minimum of two nodes across two Availability Zones. Confirm the cluster is healthy with kubectl get nodes and note the allocatable pod count per node from kubectl describe node. That number is your ceiling for the node-based path, and it is almost always lower than the instance's CPU would suggest.
2. Deploy the batch workload on nodes. Apply a Job manifest with a modest CPU request and a parallelism of ten. Watch the pods schedule and record how long from apply to all pods Running. Then scale parallelism to forty and repeat. Note whether the cluster autoscaler adds nodes, how long that takes, and whether any pods sit Pending during the ramp. The time-to-capacity figure you record here is the number that matters for the comparison.
3. Create a Fargate profile for the namespace. Add a profile matching the batch namespace with a label selector that matches the job pods. Confirm that new pods land on Fargate by checking that they have no node name and that kubectl get nodes shows no new nodes. Re-run the same parallelism ramp and record the same time-to-capacity figure.
4. Deliberately break Fargate. Add a DaemonSet to the namespace and observe what happens — it will not schedule, and the event will name the constraint. Then modify the Job to request a CPU/memory combination outside the supported Fargate set and observe the admission failure. Record both error messages verbatim; they are the diagnostic signatures you will recognize in production.
5. Deliberately break the node path. On the node-based namespace, deploy enough pods to exhaust the CNI IP allocation on a single node while leaving CPU and memory free. Observe the Pending state and the event message. Then enable prefix delegation on the VPC CNI add-on and re-test to confirm the ceiling moves.
6. Compare cost and record the crossover. Using the utilization you observed — which for a bursty batch workload will be low on average — estimate the monthly cost of the node group at its minimum size versus the Fargate cost for the same total job-seconds. Write down the duty cycle at which the two cross over. That number is the answer to the "is Fargate more expensive" question for this specific workload, and it is usually lower than people expect.
7. Write the decision memo. In one page, state which substrate you would choose for this namespace and why, citing the time-to-capacity figures, the structural constraints you hit, and the crossover duty cycle. Include the mixed-cluster option explicitly: a Fargate profile for the batch namespace plus a small managed node group for cluster-level DaemonSets is a legitimate answer, and the memo should say why you did or did not choose it.
Scenario Question Drills
Q1. A platform team runs a log-collection agent as a DaemonSet across an EKS cluster. They want to move the application tier to Fargate to reduce node management. What happens to the DaemonSet?
Q2. A team has deep existing Kubernetes tooling, Helm charts, and a requirement to run the same manifests on another cloud. Which AWS compute platform best fits?
Q3. Pods are stuck in Pending on a node that shows plenty of free CPU and memory. The scheduler event reads "Too many pods." What is the most likely cause?
Q4. A workload runs at roughly 12% average CPU utilization with sharp, unpredictable spikes. The team has no capacity to patch nodes. Which substrate is most likely to be both operationally and financially appropriate?
Q5. An engineer has an IAM policy granting eks:* on all resources but receives "error: You must be logged in to the server (Unauthorized)" from kubectl. What is the most likely cause?
Q6. A machine-learning team needs GPU-backed pods in an EKS cluster. Which compute option can satisfy this?
Q7. A cluster runs a Fargate profile for the web namespace and a managed node group for everything else. A pod in the web namespace has no matching label selector for the profile. Where does it run?
Q8. A Spot-heavy managed node group shows a recurring five-minute dip in available replicas that correlates with Spot reclaim events. What is the most likely explanation?
Q9. A regulated workload requires that no two tenants share a kernel. Which EKS compute option provides the strongest isolation boundary?
Q10. A pod spec requests 3.5 vCPU and 7 GB of memory and is rejected before it ever runs. Why?
Q11. A team wants AWS to handle node provisioning, scaling, and patching but still needs DaemonSets to work. Which option fits?
Q12. A product launch is expected to increase pod count tenfold in a single Region. What should be done before the launch?
Q13. A managed node group AMI update appears to hang indefinitely, with old nodes never draining. What is the most likely cause?
Q14. A team is AWS-native, has no Kubernetes experience, and wants the lowest possible operational overhead for a containerized web application. Which choice is most appropriate?
Q15. A cluster is configured with private endpoint access only, and an engineer's kubectl commands time out rather than returning an authorization error. What does this indicate?
Peek into Tomorrow
Everything in this day assumed that capacity appears when you ask for it. A node group scales out, a Fargate pod schedules, and the workload starts. What we have not examined is the gap between the moment a scaling decision is made and the moment the new capacity can actually serve traffic — and on node-based compute that gap is not small. An EC2 instance has to boot, the kubelet has to join the cluster, the container runtime has to pull images, and the application has to initialize before it is useful. For a workload with a four-minute bootstrap, a scale-out triggered by a traffic spike arrives four minutes after the spike, which is often too late to matter.
Tomorrow's material addresses that gap directly, and it does so with two mechanisms that are easy to confuse. One pauses an instance at a defined point in its lifecycle so a script can run — registering it with a load balancer on the way in, or draining connections on the way out — which is about correctness rather than speed. The other keeps a pool of instances already past the expensive part of startup, so scale-out promotes something that is nearly ready instead of beginning from a cold boot. The open question worth carrying forward is which of those two you reach for when the problem is latency rather than correctness, and what the tradeoff is when you keep pre-initialized capacity sitting idle.
Sources
- Amazon EKS User Guide — What is Amazon EKS
- Amazon EKS User Guide — Managed node groups
- Amazon EKS User Guide — AWS Fargate for EKS
- Amazon EKS User Guide — Pod configuration for Fargate
- Amazon EKS User Guide — VPC CNI plugin and prefix delegation
- Amazon EKS User Guide — Granting Kubernetes access (aws-auth and access entries)
- Amazon EKS User Guide — Service quotas
- Amazon EKS User Guide — Amazon EKS optimized AMIs
- AWS Whitepaper — Overview of Amazon EKS Architecture