Serverless Architectures — Step Functions & API Gateway
Recap: Where We Left Off
Day 19 ended on the two concurrency controls that decide whether a Lambda-based system stays up under load: reserved concurrency, which caps and guarantees a function's maximum simultaneous executions, and provisioned concurrency, which pre-warms execution environments so cold starts stop showing up in your p99. It also covered event source mapping, the polling layer that batches records from Kinesis, DynamoDB Streams, and SQS into Lambda invocations, along with the per-source concurrency limits and the partial batch failure and DLQ handling that keeps a single poisoned record from stalling an entire shard.
Today extends that material rather than replacing it. Reserved concurrency answers "how many copies of this function may run at once," but it says nothing about what those copies are doing or in what order. The moment a business process spans more than one function — validate, reserve inventory, charge a card, wait for a human approval, notify — the concurrency knobs stop being the interesting part of the design. Step Functions and API Gateway are the layer that sits above the individual function and owns sequencing, retries, and the request contract. The same reserved-concurrency reasoning still applies underneath; it just becomes one input to a larger orchestration decision.
Foundations You'll Need Today
Today's material sits on top of a few ideas that the rest of this page will use constantly without stopping to define them. None of them are complicated, but if any one of them is fuzzy, the arguments about Step Functions and API Gateway will read as a list of product names rather than a set of design tradeoffs. Here is the grounding.
What a Lambda function is, and what "invoking" one means
A Lambda function is a piece of code you upload to AWS and then never think about servers again. You do not rent a machine, install a runtime, or keep a process running. Instead, you hand AWS the code and say "run this when something calls it." That call is what "invoking" means: some other thing — a user's browser hitting an endpoint, a file landing in a storage bucket, a timer firing — triggers the function, AWS spins up a small execution environment, runs your code once with the input it was given, and tears the environment down. You are billed for the milliseconds the code actually ran, not for idle time.
Two consequences of that model come up repeatedly today. The first is that because environments are created on demand, the very first invocation after a quiet period has to wait for one to be built, which is the "cold start" the recap mentions. The second is that AWS will happily run many copies of your function at the same time if many things invoke it at once — that is "concurrency," and it is why the previous day's material about capping and pre-warming those copies matters. Keep both in mind: a Lambda is not a long-running program you talk to, it is a short-lived unit of work that gets started, does one job, and stops.
Synchronous versus asynchronous calls
When one piece of software calls another, there are two fundamentally different ways to do it. In a synchronous call, the caller sends the request and then waits, doing nothing else, until a response comes back. This is how a web browser talks to a website: you click, the page waits, the answer arrives, the page renders. It is simple to reason about because the caller always knows the outcome immediately. The cost is that the caller is blocked for the entire duration, and if the answer never comes, the caller is stuck waiting until something gives up.
In an asynchronous call, the caller hands off the request and moves on without waiting for the result. The work still happens, but the caller learns about the outcome later, if at all — through a notification, a status check, or a record written somewhere. This is how you avoid one slow component freezing everything upstream of it. Almost every failure mode described later on this page is really a story about the mismatch between these two styles: a caller that expects a synchronous answer, a backend that takes too long to give one, and a timeout that fires in between. When you read about a 504 error or a lost response that causes a duplicate charge, that is the seam between synchronous and asynchronous behavior.
APIs, HTTP, and status codes
An API — application programming interface — is just a defined way for one piece of software to ask another piece of software to do something. Rather than a human clicking buttons, a program sends a structured request and gets a structured response back. The overwhelming majority of web APIs speak HTTP, the same protocol your browser uses, which means a request has a method (GET to read something, POST to create something, and so on), a URL identifying what you are asking about, and optionally a body carrying data. The response comes back with a status code, a three-digit number that summarizes what happened.
You do not need to memorize the whole list, but three families matter today. Codes starting with 2 (like 200) mean success. Codes starting with 4 (like 429) mean the caller did something the server would not accept — commonly, sending requests too fast. Codes starting with 5 (like 504) mean the server itself failed to produce an answer, often because something behind it took too long. When this page says API Gateway returns a 504 or that throttling shows up as 429s, it is describing which of those families the client sees, and that distinction is usually the first clue in diagnosing a problem.
Queues, and why decoupling matters
A queue is a holding area for messages. One program puts a message in; another program takes it out and acts on it. The two never talk to each other directly, and that indirection is the entire point. The producer can keep producing even if the consumer is slow, down, or not yet written, because the queue absorbs the backlog. The consumer can process at its own pace, and if it crashes mid-message, the message can be redelivered rather than lost. This is what people mean by "decoupling": the two sides are connected through a buffer instead of a direct call, so neither one's problems immediately become the other's.
Today's page contrasts several services that all sit in this space but behave differently. A queue like SQS holds messages until something consumes them and can guarantee ordering if you configure it to. A streaming service like Kinesis keeps an ordered log that multiple readers can replay. EventBridge is a router rather than a queue: it inspects each event and forwards it to whichever consumers have registered an interest, without the producer knowing who they are. The distinctions between those three are exactly the kind of thing the exam tests, and they only make sense once the basic queue idea is clear.
With that grounding, here's why Step Functions and API Gateway exist and what problem they actually solve.
1. Why This Is on the Exam
SAP-C02 scenario questions rarely ask you to define Step Functions or API Gateway. They ask you to choose between them and something else — a chain of Lambda functions calling each other directly, an SQS queue with a consumer, an Application Load Balancer in front of Fargate, or a single monolithic function that does everything. The exam is testing whether you can recognize the architectural smell of a workflow that has outgrown ad-hoc function-to-function invocation, and whether you know which of the two serverless front-door services actually fits a given requirement.
The reason this shows up so often is that the failure mode is subtle and expensive. A team builds a three-step process as three Lambdas where each one invokes the next. It works in development. In production, a transient failure in step two means step one has already committed its side effect, and there is no record anywhere of which executions are half-finished. Retries get bolted on inside each function, and now the retry logic is duplicated three times with three different backoff strategies. The exam presents this as a scenario — "the team needs visibility into which orders are stuck and the ability to retry a failed step without re-running the whole process" — and the correct answer is almost always to move the sequencing into a state machine.
This maps most directly to Domain 2 (Design Resilient Architectures) and Domain 3 (Design High-Performing Architectures), with a recurring cost angle from Domain 4. Resilience because orchestration is how you get idempotent retries and durable execution state. Performance because API Gateway's integration types and caching decisions determine whether a request costs you a Lambda invocation at all. Cost because Standard versus Express workflow pricing, and REST versus HTTP API pricing, differ by enough to change the answer on a high-volume scenario. When a question mentions millions of executions, sub-second latency requirements, or a need for exactly-once semantics, it is pointing at one of these forks.
2. How the Orchestration Layer Actually Works
A Step Functions state machine is a JSON document — the Amazon States Language definition — that describes a set of states and the transitions between them. The service maintains the execution history for every run, and that history is the whole point. Each state transition is recorded, so when an execution fails you can see exactly which state it died in, what the input to that state was, and what the error was. The state machine itself is not a running process; it is a durable coordinator that invokes other things and waits. That distinction matters because it means an execution can pause for hours or days without consuming compute, which is what makes human-approval steps practical.
States come in a handful of types that cover most orchestration needs. A Task state does work — invoke a Lambda, run an ECS task, call an API Gateway endpoint, publish to SNS, start a nested state machine, or call almost any AWS API directly through the optimized integrations. A Choice state branches on the input. A Wait state pauses for a duration or until a timestamp. A Parallel state fans out and waits for all branches. A Map state iterates over an array, either inline or in Distributed mode for large-scale parallel processing. A Pass state transforms data without doing work. Fail and Succeed terminate the execution. The retry and catch fields on a Task state are where the resilience story lives: Retry defines which error names to retry, how many times, with what interval and backoff rate, and Catch defines where to transition when retries are exhausted.
API Gateway sits at the other end of the request path. It is a managed front door that terminates TLS, authenticates the caller, validates and transforms the request, and routes it to a backend. The backend can be a Lambda function, but it does not have to be — API Gateway has direct integrations to other AWS services, which means a request that only needs to write an item to DynamoDB can be served without ever invoking a function. That is the detail most candidates miss. When a scenario describes a simple pass-through endpoint with no business logic, the cheapest and lowest-latency answer is often a direct service integration rather than a Lambda proxy. EventBridge is the third piece: an event bus that routes events from AWS services and your own applications to targets based on rules, which is how you decouple a producer from an unknown number of consumers without the producer knowing any of them exist.
3. The Core Decision Boundary: Standard vs. Express
The single fork that most Step Functions scenario questions hinge on is the workflow type, and it is a genuine tradeoff rather than a strict upgrade in either direction. Standard Workflows are built for long-running, auditable, exactly-once processes. They keep a full execution history in the console and API for up to 90 days, they support the full set of state types including the synchronous callback pattern used for human approvals, and they are priced per state transition. Express Workflows are built for high-volume, short-duration, event-processing workloads. They are priced by number of executions and by duration, they can run at very high request rates, and they emit logs and metrics to CloudWatch rather than maintaining a queryable execution history. The semantics are at-least-once, which means your downstream operations need to be idempotent.
The exam phrasing that selects Standard is almost always about durability or human interaction: "the process may wait several days for approval," "the team needs to audit every step of each execution," "the workflow must not run a step twice." The phrasing that selects Express is about volume and cost: "millions of executions per day," "short-lived event processing," "the workload can tolerate at-least-once delivery." A useful mental check is whether you would ever need to open the console and look at one specific execution from last week. If yes, that is Standard. If the answer is "we only care about aggregate success rates," Express is the cheaper fit.
| Dimension | Standard Workflows | Express Workflows |
|---|---|---|
| Execution semantics | Exactly-once | At-least-once |
| Max duration | Up to one year | Up to five minutes |
| Execution history | Queryable via console and API | CloudWatch Logs only |
| Pricing model | Per state transition | Per execution plus duration |
| Request rate | Moderate, rate-limited | Very high throughput |
| Callback pattern (human approval) | Supported | Not supported |
| Best fit | Business processes, approvals, audits | Event processing, streaming, IoT |
4. Configuration Modes and Their Tradeoffs
On the API Gateway side, the first configuration decision is REST API versus HTTP API, and the naming is genuinely unhelpful because both serve HTTP traffic. REST APIs are the older, feature-complete product. They support API keys and usage plans, request and response validation against JSON Schema models, AWS WAF integration, private endpoints, response caching, and the full set of request/response transformation mapping templates. HTTP APIs are the newer, leaner product built for lower latency and lower cost. They support JWT authorizers and Lambda authorizers, but they drop API keys, usage plans, per-method caching, and request validation. The tradeoff is roughly "features versus price and latency," and the exam will describe a requirement for usage plans or WAF and expect you to land on REST.
The second decision is the integration type, and this is where cost savings hide. A Lambda proxy integration passes the entire request to a function and expects a specific response shape back. A Lambda custom integration lets you use mapping templates to reshape both directions. An AWS service integration calls another AWS service directly — SQS, SNS, DynamoDB, Kinesis, Step Functions — with no function in the path. A mock integration returns a static response without touching any backend, which is useful for health checks and for stubbing endpoints during development. An HTTP integration proxies to an arbitrary HTTP endpoint, which is how you put API Gateway in front of an existing on-premises service or a third-party API.
Step Functions has its own set of configuration knobs worth knowing. The Map state has two modes: Inline, which runs up to 40 concurrent iterations within the parent execution and shares its history, and Distributed, which launches child executions and can process very large datasets with high concurrency. Distributed Map is the answer when a scenario describes processing a large S3 manifest or millions of items. The callback pattern — where a Task state passes a task token to an external system and waits for that system to call SendTaskSuccess or SendTaskFailure — is the mechanism behind human approval workflows and long-running external integrations. It is also the reason those workflows must be Standard: Express has a five-minute ceiling and no callback support.
5. Sizing, Limits and Quotas
Quotas matter here because several of them are the actual constraint in a scenario rather than a footnote. Step Functions Standard executions can run for up to one year, which is what makes multi-day approval workflows viable. Express executions are capped at five minutes, so any scenario describing a long-running process rules Express out immediately. The state machine definition itself has a size limit, and the execution history for a Standard workflow is capped at 25,000 events — a limit that long-running Map-heavy workflows can genuinely hit, which is one reason Distributed Map exists. The Map state's Inline mode allows up to 40 concurrent iterations, while Distributed Map scales far higher and is the correct answer for large-scale fan-out.
API Gateway quotas shape the front door. REST APIs have a default steady-state limit of 10,000 requests per second per region with a burst capacity, and this is adjustable through a support request. HTTP APIs have a higher default request rate. Payload size is capped at 10 MB for REST APIs, which is a hard limit that shows up in scenarios about file uploads — the standard answer is to have the client upload directly to S3 using a presigned URL and pass only the object key through the API. Integration timeout is 29 seconds for REST APIs, which is the number that kills any design where API Gateway synchronously waits on a long-running backend process. The correct pattern in that case is to return immediately and hand the work to Step Functions or a queue.
| Limit | Value | Why it matters in a scenario |
|---|---|---|
| Standard execution duration | Up to 1 year | Enables multi-day human approval waits |
| Express execution duration | Up to 5 minutes | Rules out long-running processes |
| Standard execution history | 25,000 events | Long Map-heavy workflows can exhaust it |
| Map state, Inline mode | Up to 40 concurrent iterations | Large fan-out needs Distributed Map |
| REST API payload | 10 MB | Large uploads go direct to S3 via presigned URL |
| REST API integration timeout | 29 seconds | Long backends must be asynchronous |
| REST API default rate | 10,000 RPS per region | Adjustable; HTTP APIs default higher |
6. Failure Modes and What They Look Like in Production
The most common production failure in a Step Functions workflow is a Task state that fails in a way the state machine does not handle, leaving the execution in a Failed state with no automatic recovery. The symptom is a growing count of failed executions in CloudWatch and a business process that has silently stopped for a subset of customers. The first diagnostic move is to open the execution history and read the failed state's error and cause fields, which usually name the downstream service and the specific error. The fix is almost always to add a Catch block that routes to a compensating action or a dead-letter path, rather than letting the execution terminate.
The second failure mode is the non-idempotent retry. A Task state invokes a Lambda that charges a credit card, the Lambda succeeds but the response is lost to a network timeout, Step Functions retries, and the customer is charged twice. This is not a Step Functions bug — it is a design gap, and the exam tests it. The mitigation is to make the downstream operation idempotent, typically by passing a deterministic idempotency key derived from the execution ID and having the downstream service deduplicate on it. Any scenario that mentions retries and a payment or provisioning step is probing whether you know this.
On the API Gateway side, the classic failure is the 504 integration timeout. A backend that takes longer than 29 seconds to respond produces a gateway timeout, and the client sees a failure even though the backend may eventually complete. The symptom is a spike in 5xx responses correlated with a specific endpoint, while the backend's own metrics look healthy. The fix is architectural: return a 202 with a job identifier, do the work asynchronously, and let the client poll or receive a notification. The other common failure is throttling — 429 responses when request volume exceeds the account or stage limit — which is diagnosed by looking at the ThrottleCount and 4XXError metrics together. A third, quieter failure is a Lambda authorizer that caches a decision for longer than intended, so a revoked token keeps working until the cache TTL expires.
7. The Operational and SRE Angle
Step Functions emits metrics that map cleanly onto an SLO. ExecutionsStarted, ExecutionsSucceeded, ExecutionsFailed, ExecutionsTimedOut, and ExecutionsAborted are the core set, and the ratio of failed to started is the natural error-rate signal. For a business process, the more useful SLO is often expressed in terms of completion latency rather than raw error rate — "95% of orders reach a terminal state within ten minutes" — which you can approximate by alarming on ExecutionTime. The operational discipline that matters most is alerting on failed executions rather than waiting for a customer to report a stuck order, because a failed execution is silent by default.
API Gateway's operational surface is the standard four golden signals applied to the front door. Count is the request volume per stage and method. Latency is reported as IntegrationLatency (time spent in the backend) and Latency (total time including API Gateway overhead), and the gap between them tells you whether a slowdown is yours or the gateway's. Errors split into 4XXError (client-side, including throttling) and 5XXError (server-side, including integration timeouts). Saturation shows up as ThrottleCount. The runbook shape for a 5xx spike is: check whether IntegrationLatency rose first (backend problem) or whether 5XXError rose with flat latency (a backend returning errors quickly), then check whether the backend's own concurrency limits were hit.
For EventBridge, the operational concern is different again. EventBridge is a routing layer, and a rule that matches nothing fails silently — the event is simply dropped. The metrics to watch are Invocations, FailedInvocations, and ThrottledRules, and the discipline is to have a catch-all rule or a dead-letter queue on every rule that matters, so that an event which fails to reach its target is recoverable rather than lost. Because EventBridge is asynchronous and decoupled by design, the producer has no idea whether the consumer succeeded, which is exactly the property that makes it useful and exactly the property that makes silent loss possible.
8. Edge Cases and Exam Gotchas
The gotcha that catches the most candidates is assuming Express Workflows are simply a cheaper Standard. They are not a drop-in replacement: Express cannot do the callback pattern, cannot run longer than five minutes, and does not retain a queryable execution history. A scenario that mentions a human approval step is a Standard scenario, full stop, regardless of volume. Conversely, a scenario that mentions millions of short executions and a tight cost target is an Express scenario, and answering Standard because "it's more reliable" is the trap.
The second cluster of gotchas is about API Gateway's limits being hard rather than soft. The 10 MB payload limit and the 29-second integration timeout are not adjustable, so any design that depends on exceeding them is wrong by construction. The correct patterns — presigned S3 uploads for large payloads, asynchronous job submission for long-running work — are the answers the exam is looking for. A related trap is choosing REST when HTTP API would do, or vice versa: if the scenario needs API keys, usage plans, request validation, or WAF, it needs REST; if it only needs JWT authorization and low latency at high volume, HTTP API is the cheaper correct answer.
Finally, be precise about what EventBridge is for. It is not a queue and it does not guarantee ordering. It is a content-based router for events, and its value is that a producer can publish without knowing its consumers. If a scenario requires strict ordering or exactly-once processing of messages, that is SQS FIFO or Kinesis, not EventBridge. If a scenario requires that a new consumer be added without changing the producer, that is EventBridge. And if a scenario describes a scheduled job, EventBridge Scheduler is the modern answer rather than a CloudWatch Events cron rule, though both appear in older material.
9. This vs. the Services It Gets Confused With
The most frequent confusion is Step Functions versus a chain of Lambda functions invoking each other. Direct invocation is simpler to build and has lower per-step latency, but it has no durable execution state, no built-in retry with backoff, no visibility into which step failed, and no way to pause. Step Functions adds a coordination layer and a small amount of latency per transition in exchange for all of those properties. The rule of thumb: if the process has more than two or three steps, involves a wait, or needs an audit trail, use a state machine. If it is a single transformation with no branching, a plain Lambda is fine and cheaper.
The second confusion is API Gateway versus an Application Load Balancer. Both are front doors, but they serve different backends. ALB is the right choice for containerized or EC2 backends, for WebSocket-style long-lived connections at the load balancer level, and when you need the richer health-check and target-group semantics. API Gateway is the right choice when the backend is Lambda or another AWS service, when you need per-request authorization and throttling at the API level, or when you want usage plans and API keys. A useful discriminator: if the scenario mentions Fargate or EC2 targets, ALB; if it mentions Lambda functions or direct service integrations, API Gateway.
| Requirement | Pick | Why |
|---|---|---|
| Multi-step process with retries and audit trail | Step Functions | Durable execution history and built-in Retry/Catch |
| Single transformation, no branching | Lambda | Lower latency, no orchestration overhead |
| Human approval that may wait days | Step Functions Standard | Callback pattern plus one-year execution limit |
| Millions of short event executions, cost-sensitive | Step Functions Express | Per-execution pricing, high throughput |
| HTTP front door for Lambda or AWS services | API Gateway | Native integrations, auth, throttling, usage plans |
| HTTP front door for Fargate or EC2 | Application Load Balancer | Target groups, health checks, container-native |
| Fan out one event to unknown consumers | EventBridge | Content-based routing, producer decoupled from consumers |
| Strict ordering, exactly-once message processing | SQS FIFO or Kinesis | EventBridge does not guarantee ordering |
Hands-on Lab: A Standard Workflow with Human Approval and Retry/Catch
The goal of this lab is to build a Step Functions Standard state machine that models an order-approval process: it validates an order, attempts a downstream charge, retries transient failures with backoff, catches permanent failures into a compensation path, and pauses for a human approval before fulfilment. You will wire the approval using the callback pattern with a task token, and you will verify the retry and catch behavior by deliberately failing the charge step.
1. Create the downstream functions. Write three small Lambda functions. validateOrder takes an order object and returns it with a validated: true flag, or throws a ValidationError if the amount is missing. chargeCard simulates a payment call: it should throw a TransientError on roughly the first invocation for a given order ID (store a counter in a DynamoDB table or use an environment variable to force the failure) and succeed on retry. notifyCustomer writes a message to an SNS topic. Give each function an IAM role scoped to only what it needs.
2. Define the state machine. In the Step Functions console, author a state machine in Amazon States Language. Start with a Task state that invokes validateOrder. Add a Choice state after it that branches to a Fail state if validated is false. The next Task state invokes chargeCard and carries a Retry block with ErrorEquals: ["TransientError"], IntervalSeconds: 2, MaxAttempts: 3, and BackoffRate: 2.0. Add a Catch block on the same state with ErrorEquals: ["States.ALL"] that transitions to a compensation Task state which marks the order as failed.
3. Add the human approval wait. After a successful charge, add a Task state that uses the .waitForTaskToken integration pattern. The state should invoke a Lambda that publishes the task token and the order details to an SNS topic or writes them to a DynamoDB table for an approver to pick up. The state machine will pause here indefinitely — this is the behavior that requires Standard Workflows. Set a TimeoutSeconds on the state so that an unanswered approval eventually fails rather than hanging forever.
4. Complete the approval loop. Build a second small Lambda, submitApproval, that takes a task token and a decision and calls SendTaskSuccess or SendTaskFailure on the Step Functions API. Invoke it manually from the console or via a test API Gateway endpoint to simulate an approver clicking approve or reject. On success, the workflow should proceed to notifyCustomer and then a Succeed state; on failure, it should route to the compensation path.
5. Verify the failure paths. Start an execution and confirm the charge step retries three times with increasing intervals before either succeeding or catching. Start a second execution with an invalid order and confirm it fails at the Choice state without ever reaching the charge step. Start a third and let the approval timeout expire, then inspect the execution history to see the timeout recorded as a distinct event.
6. Inspect and instrument. Open the execution history for each run and identify the exact state where each one terminated. Then create a CloudWatch alarm on the ExecutionsFailed metric for this state machine and confirm it fires when you deliberately trigger the compensation path. Finally, add a Catch block that publishes the failed execution's input to a dead-letter SQS queue so that a stuck order can be replayed manually.
Scenario Question Drills
Q1. A high-volume IoT pipeline needs to run millions of short workflow executions per day as cheaply as possible, and can tolerate at-least-once execution. Which Step Functions type fits?
Q2. An insurance claims process must pause for up to five days while a human adjuster reviews a claim, and every step of every claim must be auditable for compliance. Which configuration is required?
Q3. A REST API in API Gateway must enforce per-customer request quotas and reject requests that do not match a defined JSON schema before they reach the backend. Which API type is required?
Q4. A client needs to upload a 200 MB video through an API. The current design posts the file to an API Gateway endpoint backed by Lambda, and uploads fail. What is the correct fix?
Q5. A Step Functions workflow charges a customer's credit card, and the team has configured a Retry block on the charge state. During a network blip, a customer is charged twice. What is the root cause?
Q6. An endpoint backed by a Lambda function consistently returns 504 errors under load, while the Lambda's own error rate and duration metrics look normal. What is the most likely cause?
Q7. A team needs to process a manifest of 5 million S3 objects, applying the same transformation to each, with high concurrency and per-item error isolation. Which Step Functions feature fits?
Q8. A producer application publishes order events, and the team wants to add a new analytics consumer next quarter without modifying or redeploying the producer. What should route the events?
Q9. A simple endpoint needs to write an item to DynamoDB and return a 200. There is no business logic, no transformation, and no other backend call. What is the most cost-effective design?
Q10. A workflow's execution history shows it terminated in a Failed state, and the team wants to know which step failed and why. Where is that information?
Q11. An EventBridge rule that should forward critical security events to a Lambda function appears to be dropping some events, and no errors are visible. What is the most likely explanation?
Q12. A team wants to expose an existing on-premises HTTP service through API Gateway without rewriting it as a Lambda function. Which integration type should they use?
Q13. A Step Functions workflow invokes a Lambda that occasionally throws a throttling error. The team wants the workflow to retry with increasing intervals and, after three attempts, route to a notification path rather than failing. Which two state machine fields accomplish this?
Q14. A scenario requires strict ordering of messages and exactly-once processing for financial transactions. Which service should handle the messages?
Q15. A team is choosing between an Application Load Balancer and API Gateway for a new service. The backend will be a Fargate task that needs health checks and target-group-based routing. Which should they choose?
Peek into Tomorrow
Everything covered today assumed you had already decided that the workload belongs on a serverless orchestration layer. That assumption is doing a lot of work, and it is exactly the assumption tomorrow's synthesis is designed to stress. Step Functions and API Gateway are one point in a much larger compute decision space that also includes ECS on Fargate, EKS with managed node groups, plain Lambda invoked directly, and EC2 Auto Scaling groups. Each of those has a different answer to the same three questions: how much operational overhead does the team absorb, what does it cost at the expected volume, and how much control do you give up in exchange for the managed layer.
The unresolved question is where the boundary actually sits. A team with deep Kubernetes experience and existing manifests will make a different call than a team of three developers with no container platform experience, even for the same workload. Cost curves cross at different volumes — a Fargate service that is cheap at low traffic can become more expensive than a right-sized EC2 fleet at steady high traffic, while a Lambda-based design inverts that relationship. Tomorrow consolidates those tradeoffs into a decision matrix and drills them with scenario questions, which is the point at which the individual service knowledge from this week has to become a repeatable judgment call.
Sources
- AWS Step Functions Developer Guide — What Is AWS Step Functions?
- AWS Step Functions — Standard vs. Express Workflows
- AWS Step Functions — Service Integration Patterns (including callback with task token)
- AWS Step Functions — Error Handling (Retry and Catch)
- Amazon API Gateway Developer Guide — What Is Amazon API Gateway?
- Amazon API Gateway — Choosing Between HTTP APIs and REST APIs
- Amazon API Gateway — Quotas and Important Notes
- Amazon EventBridge User Guide — What Is Amazon EventBridge?
- AWS Well-Architected Framework — Reliability Pillar