Well-Architected Reliability Pillar Deep Dive
Recap: From Backup Mechanics to Reliability Design
Yesterday's work on AWS Backup established a concrete mechanism: policy-based backup schedules that sweep EBS volumes, RDS instances, DynamoDB tables, and EFS file systems from one console, with cross-region and cross-account copy for isolation and Vault Lock enforcing WORM immutability so that even a compromised administrator cannot delete a backup before its retention expires. That is a strong answer to the ransomware question, and it is the answer most exam scenarios about "protect backups from a compromised account" are looking for.
What that material deliberately did not answer is the larger question of where backup sits inside a reliability strategy. Backup is one of four failure-management activities, and it is the cheapest and slowest of them. Today extends yesterday's material by widening the frame from "how do I protect the data" to "how do I design a workload that survives failure at all" — the Well-Architected Reliability pillar, which treats backup as the last line of defense rather than the first. The pillar's four areas — foundations, workload architecture, change management, and failure management — are the vocabulary the exam uses to describe every resilience scenario you will see from here forward.
Foundations You'll Need Today
Today's material is written in the vocabulary of reliability engineering, and that vocabulary assumes a handful of ideas that are easy to take for granted once you have used AWS for a while. None of them are complicated, but each one is load-bearing for the arguments that follow, so it is worth spending a few minutes on them before the pillar's four areas start making sense.
Availability Zones and Why "Multi-AZ" Is a Real Design Choice
An AWS region is a geographic area — Northern Virginia, Ireland, Singapore — and inside each region AWS operates several physically separate data centers called Availability Zones, or AZs. "Physically separate" is the important part: each AZ has its own power, cooling, and network connections, and they are far enough apart that a fire, flood, or power failure in one is very unlikely to affect another. When you launch a server in AWS, you pick a region and an AZ within it, and that choice determines which physical building your server actually runs in.
This matters because a single server in a single AZ has a single point of failure: if that building loses power, your application is down. Running the same application in two or three AZs at once — which is what "Multi-AZ" means — means that the loss of any one building degrades capacity but does not take the application offline. The trade-off is that you are paying for more servers than you strictly need at any given moment, which is why the pillar treats Multi-AZ as a decision to be justified by an availability target rather than as a default. The same logic extends to a related concept you will see in today's failure-mode discussion: a NAT gateway is the AWS-managed component that lets servers in a private subnet reach the internet, and if all your servers in every AZ route through a NAT gateway that lives in only one AZ, then you have Multi-AZ compute with a single-AZ dependency — redundancy on paper, not in practice.
Service Quotas: The Limits You Don't See Until You Hit Them
Every AWS account has limits on how much of each service it can use in each region — a maximum number of servers, a maximum number of database instances, a maximum rate of function invocations. These are called service quotas, and they exist to protect both you and AWS from runaway usage. The two properties worth remembering are that quotas are per-account and per-region, which means two accounts in the same region have independent limits, and the same account in two regions has independent limits in each.
Quotas are adjustable: you can request an increase through the Service Quotas console, and AWS usually grants reasonable requests within a day or two. The reason this belongs in a reliability discussion is that a workload which scales up during a traffic spike and hits a quota will simply fail to scale, and the resulting errors look exactly like an application bug. The fix is not architectural — it is to know which quotas your design will approach and to request increases before launch rather than during an incident.
RTO and RPO: Two Numbers That Define Every Recovery Decision
When a business says it needs to "survive a disaster," that statement is not actionable until it is expressed as two numbers. Recovery Time Objective, or RTO, is how long the business can tolerate being down — if the RTO is four hours, then a recovery process that takes three hours is acceptable and one that takes six is not. Recovery Point Objective, or RPO, is how much data the business can tolerate losing, measured in time — an RPO of five minutes means that after a failure, the business can accept losing the last five minutes of transactions, but not the last hour.
These two numbers are the single most useful input to any resilience design, because they determine which recovery pattern is proportionate. A long RTO and a long RPO can be satisfied by taking periodic backups and restoring them when needed, which is cheap. A short RPO requires continuously copying data as it changes, which is more expensive. A short RTO requires having infrastructure already running and ready to take over, which is more expensive still. Today's material will repeatedly ask you to start from these two numbers and work outward, rather than starting from a preferred architecture and hoping it satisfies the requirement.
Backup vs. Replication: Two Different Jobs
Backup and replication are often mentioned in the same breath, but they solve different problems and have different failure characteristics. A backup is a periodic copy of your data taken at a point in time and stored somewhere separate — last night's database snapshot, for example. Its strength is that it protects against data corruption and accidental deletion, because you can go back to a copy from before the bad thing happened. Its weakness is that the gap between backups is the amount of data you can lose, so a nightly backup has an RPO measured in hours no matter how reliable the storage is.
Replication is the continuous copying of changes from one place to another as they happen — every write to the primary database is also sent to a second copy, usually within seconds or milliseconds. Its strength is a very short RPO, because the second copy is almost always current. Its weakness is that it faithfully replicates mistakes: if an application bug deletes a table, the deletion is replicated too, and the second copy is now just as broken as the first. This is why mature designs use both — replication for fast failover, backup for the ability to rewind — and why today's material treats them as complementary rather than as alternatives.
With that grounding, here is why the Reliability pillar exists as a structured discipline and what problem it actually solves.
1. Why the Reliability Pillar Is on the Exam
The Reliability pillar is the single largest source of scenario material on SAP-C02. The exam's second domain, "Design for New Solutions," and its resilience-adjacent questions in the first domain both lean on the same underlying question: given a stated business requirement, which architectural choice actually delivers the required availability at the lowest reasonable cost and operational burden? The pillar exists because AWS found that most outages are not caused by AWS infrastructure failing — they are caused by customers building workloads that cannot tolerate a component failing, or by changes that were deployed without a rollback path. The pillar is therefore written as a design discipline, not a service catalog.
What makes this examinable is that the pillar gives you a shared vocabulary for trade-offs that would otherwise be argued in adjectives. "Highly available" is not a requirement; "99.99% availability, RTO of 15 minutes, RPO of 5 minutes" is. Every reliability scenario on the exam is ultimately asking you to map a stated requirement onto one of the four areas and then pick the cheapest design that satisfies it. A question that says "the business cannot tolerate more than five minutes of data loss" is an RPO question, which is a failure-management question, which points at replication rather than backup. A question that says "deployments occasionally take down production" is a change-management question, which points at canary releases and automated rollback rather than at Multi-AZ.
The trap is that candidates reach for the most resilient-sounding answer rather than the one that matches the stated requirement. Multi-region active-active is almost never the right answer to a question that specifies a 15-minute RTO, because it costs an order of magnitude more than the pilot-light design that meets the requirement. The pillar's framing — design to the requirement, not to the maximum — is what the exam is testing when it offers you four technically valid architectures and only one that is proportionate.
2. How the Pillar Is Structured: Four Areas and Their Design Questions
The Reliability pillar is organized into four areas, and the exam expects you to know which area a given scenario belongs to because the remediation differs by area. Foundations covers the things you must get right before the workload exists: service quotas that will silently throttle you at scale, network topology that determines whether an AZ failure is survivable, and account structure that determines blast radius. Workload architecture covers how the workload itself is built — whether it is distributed across multiple AZs, whether it degrades gracefully when a dependency is slow, and whether it can scale horizontally rather than vertically. Change management covers how you modify the workload: automated deployment, monitoring of the deployment itself, and the ability to reverse a change quickly. Failure management covers what happens when something fails anyway: backup, disaster recovery, and — critically — tested recovery procedures.
The reason the four areas matter more than any individual best practice is that they form a progression of cost. Foundations are nearly free and prevent entire classes of failure. Workload architecture costs engineering effort and some infrastructure duplication. Change management costs tooling and discipline. Failure management is the most expensive per unit of protection, because you are paying for capacity that sits idle until something goes wrong. A mature design invests in the cheap areas first and only buys failure-management capacity to cover the residual risk that the first three areas cannot eliminate. This is why the pillar's guidance reads as a sequence rather than a checklist.
There is also a design principle that cuts across all four areas and shows up constantly in exam distractors: stop guessing capacity. The pillar's position is that you should scale horizontally and automatically rather than provisioning for peak, because a workload sized for peak is over-provisioned for the 95% of the time it is not at peak, and a workload sized for average fails at peak. This is the principle behind Auto Scaling groups, DynamoDB on-demand capacity, and Aurora Serverless — and it is why "provision a larger instance type" is almost always a wrong answer on a reliability question.
3. The Core Decision Boundary: Which Area Does This Scenario Belong To?
Almost every reliability scenario on the exam can be routed to one of the four areas by asking a single question: is the failure being described something that happens before the workload runs, during normal operation, when you change the workload, or when something breaks despite your design? The answer determines which set of remediations is even relevant, and picking the wrong area is the most common way to arrive at a technically valid but wrong answer. A candidate who reads "the application went down during a deployment" and reaches for Multi-AZ has misrouted the scenario — Multi-AZ protects against infrastructure failure, not against a bad deployment, and the correct remediation is a canary release with automated rollback.
The routing question is worth internalizing because the exam frequently presents a scenario with a plausible-sounding symptom and a root cause in a different area. "The application is intermittently slow under load" could be a workload-architecture problem (no horizontal scaling, a single-threaded component) or a foundations problem (a service quota being hit, a NAT gateway bottleneck). The diagnostic move is to ask what changes when the failure occurs: if it correlates with traffic volume, it is a scaling or quota issue; if it correlates with a deployment, it is change management; if it correlates with an AZ or region event, it is workload architecture or failure management.
| Scenario symptom | Area | First remediation to consider |
|---|---|---|
| Throttling at a predictable scale point | Foundations | Raise the service quota; request increases ahead of launch |
| Single AZ outage takes the workload down | Workload architecture | Multi-AZ deployment behind a load balancer |
| Deployment caused an outage | Change management | Canary/blue-green release with automated rollback |
| Data loss after a corruption event | Failure management | Backup with point-in-time recovery, tested restore |
| Slow dependency cascades into full outage | Workload architecture | Timeouts, circuit breakers, graceful degradation |
| Region-wide event exceeds RTO | Failure management | DR strategy matched to the stated RTO/RPO |
The table is not a lookup you should memorize as a mapping; it is a demonstration that the four areas partition the failure space cleanly. Once you can name the area, the candidate answers narrow to two or three, and the remaining choice is usually about cost proportionality — which is exactly the judgment the exam is scoring.
4. Design Tradeoffs Within Each Area
Within foundations, the tradeoff is between pre-provisioning headroom and paying for capacity you do not use. Service quotas are the clearest example: they are per-account, per-region limits that are generous by default but not infinite, and a workload that scales to thousands of instances or millions of Lambda invocations will hit them. The pillar's guidance is to treat quotas as a design input — know which quotas your architecture will approach, and request increases before launch rather than during an incident. The cost of doing this is essentially zero; the cost of not doing it is an outage that looks like an AWS problem but is actually a planning failure.
Within workload architecture, the central tradeoff is between coupling and complexity. A tightly coupled monolith is simpler to operate but fails as a unit; a distributed system can isolate failures but introduces partial-failure modes that did not exist before. The pillar's answer is not "always distribute" but "distribute where the failure isolation is worth the operational cost, and make the boundaries explicit." This is why the pillar emphasizes loose coupling via queues and asynchronous messaging: a synchronous call chain of five services has five places to fail and no way to degrade, while an asynchronous handoff lets the caller continue working while the downstream recovers.
Within change management, the tradeoff is between deployment velocity and blast radius. Deploying everything at once is fast and simple; deploying in small increments with automated rollback is slower per change but bounds the damage of a bad change. The pillar's position is that small, reversible changes are safer in aggregate even though each one takes longer, because the failure mode of a bad small change is a partial degradation rather than a full outage. Within failure management, the tradeoff is the familiar RTO/RPO-versus-cost spectrum: backup and restore is cheapest and slowest, pilot light is a middle ground, warm standby is faster and more expensive, and multi-site active-active is fastest and most expensive. The exam almost always gives you the RTO and RPO in the question precisely so you can pick the cheapest point on that spectrum that satisfies them.
5. Concrete Numbers: Quotas, Targets, and What They Imply
Reliability discussions become actionable only when they are quantified, and the exam rewards candidates who can attach numbers to design choices. The most important numbers are the availability targets themselves, because they determine how much redundancy is required. A 99.9% availability target permits roughly 43 minutes of downtime per month; 99.99% permits about 4.4 minutes; 99.999% permits about 26 seconds. Those figures are the reason a single-AZ deployment cannot credibly claim 99.99% — an AZ failure alone would consume the entire monthly budget — and why multi-region designs are reserved for targets at or above 99.99%.
Service quotas are the second category of numbers worth knowing. AWS publishes per-service, per-region quotas in the Service Quotas console and in each service's documentation, and the pillar's guidance is to treat them as design constraints rather than as trivia. The pattern that matters for the exam is not memorizing specific values but knowing that quotas exist, that they are adjustable by request, and that they are per-account and per-region — which means a multi-region design has independent quota headroom in each region, and a multi-account design has independent headroom in each account.
| Availability target | Allowed downtime per month | Typical design implication |
|---|---|---|
| 99.9% | ~43 minutes | Multi-AZ within a single region is sufficient |
| 99.95% | ~22 minutes | Multi-AZ plus automated recovery and health checks |
| 99.99% | ~4.4 minutes | Multi-AZ plus tested failover; multi-region for the data tier |
| 99.999% | ~26 seconds | Multi-region active-active with automated traffic steering |
The third category is RTO and RPO, which are business inputs rather than AWS limits, but which map onto concrete service capabilities. An RPO of zero requires synchronous replication, which is what Multi-AZ provides within a region. An RPO of seconds to minutes permits asynchronous replication, which is what cross-region read replicas and Aurora Global Database provide. An RPO of hours permits backup and restore. The exam will state one of these and expect you to select the replication mechanism that matches, so the mapping is worth having at hand.
6. Failure Modes and What They Look Like in Production
The most instructive failure mode in reliability work is the one that is invisible until it matters: a design that has redundancy on paper but a single point of failure in practice. The classic example is a Multi-AZ Auto Scaling group whose instances all depend on a single NAT gateway in one AZ, or a multi-AZ database fronted by an application that caches a connection to a specific replica. The infrastructure is redundant; the dependency graph is not. These failures present as an outage that "should not have happened" given the architecture diagram, and the first diagnostic move is to trace the actual request path rather than to trust the diagram.
The second common failure mode is the cascading failure, where a slow dependency consumes resources until the caller fails too. A downstream service that responds in 30 seconds instead of 50 milliseconds will exhaust the connection pool or thread pool of every caller, and the failure propagates upstream until the whole system is down even though only one component is actually degraded. The symptoms are latency climbing before errors appear, and thread or connection pool saturation on the callers rather than on the failing service. The first diagnostic move is to look at the caller's resource utilization, not the failing service's error rate, because the caller is where the exhaustion is happening.
The third is the quota exhaustion failure, which is particularly nasty because it looks like a capacity problem and is actually a configuration problem. A workload that scales out during a traffic spike and hits an EC2 or Lambda concurrency quota will fail to scale, and the resulting errors look identical to an application bug. The first diagnostic move is to check the relevant service quota against current usage, and the preventive move is to alarm on quota utilization as a leading indicator rather than waiting for the failure. The fourth is the untested-recovery failure: a backup that has never been restored, a failover that has never been executed, a runbook that has never been followed. These are not failures of the design but of the verification, and they are the reason the pillar treats tested recovery as a first-class requirement rather than a nice-to-have.
7. The Operational Angle: Monitoring, Alarms, and Runbook Shape
Reliability is not a property you design once and then possess; it is a property you observe continuously and defend operationally. The pillar's operational guidance is to instrument the workload against its stated objectives rather than against raw infrastructure metrics, because infrastructure metrics tell you that a server is healthy while the user experience is broken. The practical expression of this is to define service-level objectives for the user-facing behaviors that matter — request success rate, latency at the 99th percentile, end-to-end availability — and to alarm on the error budget those objectives imply rather than on CPU utilization.
The alarm design that follows from this is composite rather than single-metric. A latency alarm alone will page on transient spikes that self-resolve; a latency alarm combined with an error-rate alarm, using a composite alarm, pages only when both signals are present, which is a much better proxy for a real incident. The same logic applies to the choice of evaluation periods: a single 1-minute datapoint is noise, while a sustained breach over several periods is a signal. The pillar's framing is that alarms should be actionable — every page should correspond to a runbook entry, and every runbook entry should correspond to a page.
The runbook itself is the operational artifact the pillar cares most about, and its shape is diagnostic rather than prescriptive. A good runbook for a reliability incident starts with how to confirm the failure is real, then how to determine which area it belongs to, then the specific remediation for that area, and finally how to verify recovery. It should be executable by someone who did not build the system, which is the test that separates a real runbook from a note-to-self. The pillar's operational-excellence sibling makes the same point from the other direction: if a procedure is performed more than once, it should be automated, because manual procedures drift and drift is a reliability risk.
8. Edge Cases and Exam Gotchas
The most common gotcha is confusing high availability with disaster recovery. Multi-AZ is an HA pattern: it protects against the failure of a single Availability Zone within a region, and it does so with synchronous replication and automatic failover. It does nothing for a region-wide event, and it does not provide a separate copy of the data in a different geography. A question that mentions a regional outage, a geographic compliance requirement, or a data-residency constraint is a DR question, and the answer will involve cross-region replication or backup rather than Multi-AZ. Candidates lose points here constantly because Multi-AZ is the most familiar resilience pattern and therefore the most tempting.
The second gotcha is treating backup as a substitute for replication. Backup protects against data corruption and accidental deletion, and it has an RPO measured in hours because backups are periodic. Replication protects against infrastructure failure and has an RPO measured in seconds or zero. A scenario that specifies a five-minute RPO cannot be satisfied by backup alone, no matter how frequent the backup schedule, because the backup interval is the RPO. The exam will offer "increase backup frequency" as a distractor for exactly this reason.
The third gotcha is assuming that a managed service removes the need for reliability design. RDS Multi-AZ, Aurora, DynamoDB, and S3 all provide strong durability and availability guarantees, but they do not protect against a bad schema migration, a runaway query, an application bug that deletes data, or a misconfigured security group. The pillar's position is that managed services shift the reliability work rather than eliminating it, and the exam reflects this by presenting scenarios where the managed service is healthy and the customer's configuration is the failure. The fourth gotcha is the untested failover: a design that has never been exercised is an assumption, not a capability, and the pillar's answer to this is the Game Day, which is the subject of the next subsection's comparison.
9. Reliability vs. the Other Pillars, and Where Each One Stops
The Reliability pillar is frequently confused with the Operational Excellence pillar because both are concerned with how a workload behaves over time, and with the Cost Optimization pillar because resilience is expensive. The distinction is that Reliability asks whether the workload continues to work, Operational Excellence asks whether the team can operate and change it effectively, and Cost Optimization asks whether the resilience is proportionate to the requirement. A scenario about a deployment process that causes outages is an Operational Excellence scenario with reliability consequences; a scenario about over-provisioning for peak is a Cost Optimization scenario with reliability consequences. The exam expects you to name the primary pillar and then acknowledge the secondary one.
The practical rule for choosing between reliability patterns is to start from the stated RTO and RPO and work outward. If the requirement is hours, backup and restore is correct and anything more is waste. If it is tens of minutes, pilot light is the proportionate answer. If it is minutes, warm standby. If it is seconds or less, active-active. The exam will almost always give you the numbers, and the correct answer is the cheapest pattern that satisfies them — not the most resilient pattern available.
| Pattern | Typical RTO | Typical RPO | Pick it when… |
|---|---|---|---|
| Backup and restore | Hours to days | Hours | The business can tolerate a long outage and cost is the dominant constraint |
| Pilot light | Tens of minutes | Minutes | Core data must be current but standby infrastructure can be minimal |
| Warm standby | Minutes | Seconds to minutes | A scaled-down full replica can run continuously and be promoted quickly |
| Multi-site active-active | Seconds or less | Near zero | Near-zero RTO/RPO is required and budget is not the primary constraint |
The same proportionality logic applies within a single region. Multi-AZ is the right answer for a 99.9% target and the wrong answer for a 99.999% target, because no amount of intra-region redundancy can survive a regional event. The pillar's contribution to the exam is this habit of matching the design to the number, and it is the habit that separates a passing score from a marginal one.
Hands-On Lab: Running a Reliability Pillar Review (45 min)
This lab walks a sample three-tier web application through a structured Reliability pillar review. The goal is not to fix the architecture but to produce the artifact a real review produces: a list of findings, each tagged to one of the four areas, each with a stated remediation and a stated cost. Work through the steps in order and write down your findings as you go; the value is in the discipline of the review, not in the sample architecture.
- Draw the actual request path. Start from the user's browser and trace every hop to the database and back, including DNS, the load balancer, the application tier, any cache, and the data tier. Do not use the architecture diagram — reconstruct the path from the configuration. The purpose of this step is to find dependencies that the diagram omits, such as a shared NAT gateway, a single bastion host, or a hard-coded endpoint.
- Mark every single point of failure. For each hop, ask what happens if that component fails entirely and what happens if it becomes slow. A component that is redundant but whose callers all depend on one instance of it is still a single point of failure. Record the failure mode, not just the component name.
- Tag each finding to one of the four areas. Foundations, workload architecture, change management, or failure management. If a finding seems to belong to two areas, split it into two findings; the remediation differs by area and conflating them produces vague recommendations.
- Check the quotas. For every service in the path, look up the relevant per-account, per-region quota and compare it to the workload's expected peak. Record any quota the workload will approach within a factor of two, and note whether an increase has been requested.
- Review the change process. Ask how a deployment reaches production, whether it can be rolled back automatically, and whether the rollback has been exercised. A deployment process with no automated rollback is a change-management finding regardless of how good the architecture is.
- Review the recovery process. Ask when the last restore from backup was actually performed, whether the failover has been executed, and whether the runbook has been followed by someone other than its author. Untested recovery is a finding even when the backup configuration is correct.
- Quantify the requirement. Write down the availability target, RTO, and RPO the business has actually stated. If they have not been stated, that is the first finding, because every other decision depends on them.
- Match the design to the requirement. For each finding, propose the cheapest remediation that satisfies the stated requirement, and estimate its cost. A finding whose remediation costs more than the requirement justifies should be recorded as accepted risk rather than as a defect.
- Produce the findings list. Order the findings by the ratio of risk reduced to cost incurred, and identify the three that should be fixed first. This ordering is the deliverable; a review that produces an unordered list of everything wrong is not actionable.
When you have finished, compare your findings against the four-area structure. A review that produces only failure-management findings is usually a sign that the earlier areas were not examined closely enough, because foundations and workload-architecture problems are cheaper to fix and therefore should be found first.
Scenario Question Drills (20 min)
Q1. Which best exemplifies the Reliability pillar's 'failure management' best practice?
Q2. A workload runs in a single Availability Zone behind an Application Load Balancer and claims a 99.99% availability target. What is the primary reliability finding?
Q3. An application's Multi-AZ Auto Scaling group instances all route outbound traffic through a single NAT gateway in one AZ. An AZ failure takes the application down despite the Multi-AZ design. Which area does this finding belong to?
Q4. A downstream service begins responding in 30 seconds instead of 50 milliseconds, and within minutes the entire system is returning errors even though only that one service is degraded. What is the most likely mechanism?
Q5. A workload scales out during a traffic spike and begins returning errors that look like application bugs. CloudWatch shows the application is healthy. What should you check first?
Q6. A business requires an RPO of five minutes. The current design takes a nightly backup to S3. What is the correct remediation?
Q7. A team has documented DR runbooks and configured cross-region replication, but has never executed a failover. Which Reliability pillar area is the gap?
Q8. A workload must survive the loss of an entire AWS region with an RTO of 15 minutes and an RPO of 5 minutes, and the business wants to minimize standing infrastructure cost. Which DR pattern fits?
Q9. On-call engineers are paged repeatedly by isolated latency spikes that self-resolve within a minute. What alarm design reduces the noise while preserving sensitivity to real incidents?
Q10. A deployment process occasionally takes production down, and the team's response is to add more Multi-AZ redundancy. Which area does the actual problem belong to?
Q11. A team wants to verify that its automated recovery actually works before relying on it in an incident. Which Reliability pillar practice addresses this directly?
Q12. A workload uses RDS Multi-AZ and the team argues this satisfies the business's requirement to survive a regional outage. What is wrong with this reasoning?
Q13. A workload's architecture diagram shows redundancy at every tier, but an incident review finds that all application instances resolve a critical dependency through a hard-coded endpoint pointing at one instance. What is the first diagnostic move in a future incident?
Q14. A team is deciding whether to invest next in automated deployment rollback or in a warm standby region. The workload currently has no automated rollback and a stated RTO of four hours. Which investment is proportionate?
Q15. A reliability review produces a long list of findings, all tagged to failure management. What does this most likely indicate about the review?
Peek into Tomorrow
Everything in today's review assumed that the workload's design is the thing that fails. But the incident that started the review — the deployment that took production down — was not a design failure at all. It was a process failure, and no amount of Multi-AZ redundancy would have prevented it. That raises the question the Reliability pillar deliberately leaves open: if the workload is well designed and still goes down because of how the team changes it, what does the discipline of operating it actually look like?
The answer is not more architecture. It is operations as code — codifying the procedures that today live in runbooks and tribal knowledge, so that a change is a reviewed, reversible artifact rather than a sequence of manual steps. It is the practice of making small reversible changes rather than large risky ones, and of treating a blameless post-incident review as the mechanism that turns an outage into a permanent improvement. Tomorrow's material takes the runbook shape we sketched here and asks what happens when you automate it, and what a team's operating model looks like when the procedures themselves are version-controlled.
Sources
- AWS Well-Architected Framework — Reliability Pillar
- AWS Well-Architected Framework — General Design Principles
- Reliability Pillar — Design Principles
- Reliability Pillar — Best Practices (Foundations, Workload Architecture, Change Management, Failure Management)
- AWS Service Quotas Reference
- Amazon CloudWatch — Using Alarms
- Amazon CloudWatch — Composite Alarms
- AWS Whitepaper — Fault Isolation Boundaries