RTO/RPO Design Patterns & Backup Strategies (AWS Backup)
Recap: Where We Left Off
Day 38 covered Route 53 routing policies — simple, weighted, latency, failover, geolocation, geoproximity with bias, and multi-value answer routing — and the health checks that automatically pull unhealthy endpoints out of DNS responses. That material was a prerequisite for today's, and the dependency runs in a specific direction: every routing policy we discussed assumes there is a healthy endpoint somewhere to route to. Weighted routing for canary/A-B shifting, latency-based routing, and geoproximity with bias all answer the question "which of these live targets should this request go to?" None of them answer the prior question, "what happens when the target itself is gone — corrupted, deleted, or encrypted by an attacker?"
Health checks remove an endpoint from rotation, but they do not restore it. A failover routing policy moves traffic to a standby region, but it does not rebuild the data that region needs to serve. Today we close that gap. AWS Backup is the control plane that answers the recovery question underneath the routing question, and the RTO/RPO vocabulary we introduce here is what turns "we have backups" from an assertion into a number you can defend in a design review.
Foundations You'll Need Today
Today's material sits on top of four ideas that the rest of this page will use without stopping to define them. None of them are complicated on their own, but each one is doing real work in the arguments that follow, so it is worth getting them straight before we start.
Snapshots and Point-in-Time Recovery
A snapshot is a copy of a storage volume or database taken at one instant in time. It is not a live mirror of the data — it is a frozen picture of what the data looked like when the snapshot was taken, and it stays that way no matter what happens to the original afterward. That frozen quality is what makes snapshots useful for recovery: if the original is corrupted or deleted, the snapshot still holds the earlier, healthy version. The catch is that a snapshot only exists at the moments you take it. If you snapshot once a day at midnight and something goes wrong at 11 p.m., the best you can do is restore to midnight and lose almost a full day of work. Point-in-time recovery is the answer to that gap. Instead of taking discrete pictures, the database continuously records every change it makes, so you can ask it to rewind to any specific second within a retention window — say, 20 minutes before a bad deployment. The tradeoff is that continuous recording costs more and is only offered by certain managed services, which is exactly the boundary today's material keeps running into.
IAM Roles, Policies, and Cross-Account Trust
An IAM role is an identity that a service or a person can temporarily assume to get a set of permissions, rather than a permanent username and password. A policy is the document that spells out what those permissions actually are — which actions are allowed, on which resources. The part that matters today is how this works across account boundaries. By default, one AWS account cannot touch another account's resources at all; the boundary is absolute. To let account A write backups into account B, you need two things working together: a policy in account B that says "I trust account A's backup role to copy data into this vault," and a role in account A that has been granted the corresponding permission. If either half is missing or later removed, the operation simply stops working — and because the copy happens in the background, it stops working quietly. That two-sided handshake is the mechanism behind the isolated-backup-account pattern, and it is also the most common cause of the silent failure we will diagnose later.
Accounts, Organizations, and Service Control Policies
An AWS account is the fundamental unit of isolation in AWS — a hard wall around a set of resources, with its own users, permissions, and billing. AWS Organizations is the service that lets you group many accounts under one umbrella so you can manage them together, and the groups you sort them into are called organizational units, or OUs. A service control policy, or SCP, is a rule attached to an OU that sets a ceiling on what any account inside it is allowed to do, no matter what that account's own administrators want. SCPs do not grant permissions; they only take them away. This is the tool that makes "isolation" more than a word: putting a workload account in an OU with an SCP that denies the backup-deletion APIs means that even a fully compromised administrator inside that account cannot reach in and destroy the backups. Without the SCP, a separate account is just a separate account — the attacker still has a path.
Regions and Why Cross-Region Copy Matters
A region is a physically separate cluster of AWS data centers in a distinct geographic area, and resources in one region are independent of resources in another. That independence is the whole point of copying backups across regions: a failure that takes out an entire region — a large-scale power event, a natural disaster, a regional service disruption — leaves the copy in the second region untouched. Copying to another region is therefore a different and stronger guarantee than copying to another account in the same region, and the two are often combined. The cost is that the copy is stored and billed separately in the destination, and it is governed by the destination's own retention rules rather than the source's, which is why the lifecycle settings on a cross-region copy are a decision you have to make deliberately rather than inherit.
With that grounding, here is why AWS Backup exists and what problem it actually solves: it is the control plane that turns these four primitives — snapshots, cross-account trust, organizational guardrails, and regional isolation — into a single governed policy you can point at a fleet of resources and defend in a design review.
1. Why This Is on the Exam
SAP-C02 tests backup and recovery as a design discipline, not as a feature checklist. The exam's Domain 2 (Design for New Solutions) and Domain 4 (Continuous Improvement for Existing Solutions) both contain reliability objectives that hinge on whether you can select a recovery strategy that satisfies a stated business requirement at the lowest defensible cost. The recurring shape of these questions is a scenario that hands you two numbers — a recovery time objective and a recovery point objective — plus a budget constraint, and asks which architecture satisfies all three. Candidates who have memorized "AWS Backup does cross-region copy" but have not internalized the cost/RTO spectrum tend to pick the most resilient option available, which is frequently the wrong answer because it is the most expensive one.
The second reason this topic carries weight is that backup is where security and reliability requirements collide. Ransomware scenarios are now standard on professional-level exams because they break the assumption that the account owner is trustworthy. A backup strategy that a compromised administrator can delete is not a backup strategy; it is a second copy of the same risk. AWS Backup's Vault Lock feature exists specifically to make that class of attack survivable, and the exam expects you to know that immutability is a property of the vault, not of the backup plan.
Finally, this day sits at the boundary between two exam domains that are otherwise tested separately. The mechanics of backup plans, vaults, and selection criteria are reliability content. The cross-account copy pattern, the isolated backup account, and the SCP that protects it are governance content from Phase 1. Questions that span both — "design a backup architecture that a workload account cannot tamper with" — are exactly the multi-domain synthesis the exam favors, and they are the reason this material is worth more than its surface area suggests.
2. How AWS Backup Actually Works
AWS Backup is a policy-driven orchestration layer over the native snapshot and backup APIs of the services it supports. It does not replace those APIs; it schedules and coordinates them. The core object is the backup plan, which is a container for one or more backup rules. Each rule carries a schedule (a cron expression evaluated in UTC), a lifecycle block that governs when recovery points transition to cold storage and when they expire, and a target backup vault. The plan itself is inert until you attach resources to it, which you do through a resource assignment — either by tag, by resource ARN, or by resource type. Tag-based assignment is the pattern that scales, because a new EBS volume that inherits the right tag is protected the moment it is created, without anyone editing the plan.
The backup vault is the second core object and the one that carries the security properties. A vault is a logical container for recovery points, and it is where the access policy lives that governs who may read, restore, or delete from it. By default, a vault is deletable by anyone with the corresponding IAM permissions in the account. Vault Lock changes that. When you attach a vault lock policy in compliance mode, the retention settings on the recovery points inside become immutable for the duration of the lock — the vault itself cannot be deleted, and neither can the recovery points, until their retention period expires. This is a write-once-read-many guarantee enforced by the service, not by IAM, which is why it survives a compromised administrator.
Cross-region and cross-account copy are properties of a backup rule, not of the vault. A single rule can specify a destination vault in another region, another account, or both, and AWS Backup will copy each recovery point after it is created. The copy is asynchronous and independent of the source recovery point's lifecycle, which matters operationally: a copy that lands in a destination vault is governed by that vault's lifecycle settings, not the source's. This is the mechanism that makes the isolated-backup-account pattern work, and it is also the mechanism that most often surprises teams in production, because a copy can silently fail while the source backup succeeds.
3. The Core Decision Boundary: RTO and RPO
Every backup design question reduces to two numbers, and the exam will give you both. Recovery Time Objective is how long the business can tolerate the workload being unavailable — the clock starts at the moment of failure and stops when service is restored. Recovery Point Objective is how much data the business can tolerate losing, expressed as a duration measured backward from the failure. These are business inputs, not technical ones. An architect who proposes a solution before those numbers are stated is guessing, and the exam rewards candidates who recognize that the numbers, not the technology, determine the answer.
The reason RTO and RPO form a decision boundary rather than two independent axes is that they are jointly satisfied by a small number of architectural patterns, and each pattern has a characteristic cost profile. Backup and restore is the cheapest and slowest. Pilot light keeps data continuously replicated but minimal compute running. Warm standby runs a scaled-down but functional copy of the stack. Multi-site active-active serves production traffic from more than one region simultaneously. Moving down that list improves both RTO and RPO and increases standing cost, roughly monotonically. The exam question is almost always "which is the cheapest pattern that still meets the stated numbers," and the trap is choosing a more resilient pattern than the numbers require.
| Pattern | Typical RTO | Typical RPO | Standing cost | What is always running |
|---|---|---|---|---|
| Backup and restore | Hours to days | Hours | Lowest | Nothing but the backup service |
| Pilot light | Tens of minutes | Minutes | Low | Data replication only |
| Warm standby | Minutes | Seconds to minutes | Moderate | A scaled-down full stack |
| Multi-site active-active | Near zero | Near zero | Highest | Full production in 2+ regions |
The exam pattern to internalize is that RPO is usually the binding constraint, not RTO. A business that says "we can be down for an hour but we cannot lose more than five minutes of transactions" has ruled out backup and restore entirely, regardless of how cheap it is, because the backup schedule cannot produce a five-minute recovery point. Conversely, a business that says "we can lose a day of data but we need to be back in fifteen minutes" is describing a pattern where the data is cheap to replicate but the compute must be pre-warmed — which is pilot light, not warm standby.
4. Configuration Modes and Their Tradeoffs
Within AWS Backup itself, the knobs that matter are the schedule frequency, the lifecycle transitions, the copy destinations, and the vault lock mode. Schedule frequency is the direct determinant of RPO for anything backed by snapshot: a daily backup gives you a 24-hour RPO at best, an hourly backup gives you one hour, and if the requirement is tighter than that you are no longer designing a backup strategy — you are designing replication, and the answer moves to Aurora Global Database, DynamoDB Global Tables, or continuous replication via DMS. Recognizing that boundary is worth points, because the exam will offer you a backup-based answer to a replication-shaped problem.
Lifecycle transitions trade retrieval latency for storage cost. A recovery point can transition to cold storage after a defined number of days and expire after a longer one. The tradeoff is not free: cold storage recovery points take longer to restore, which directly worsens the effective RTO for any recovery that has to pull from cold tier. A plan that transitions everything to cold storage after seven days looks efficient on the bill and fails the first real recovery drill. The defensible pattern is to keep the most recent recovery points in warm storage for the window that matches your RTO, and transition older ones that exist for compliance rather than operational recovery.
Copy destinations are where the cost model gets interesting, because cross-region and cross-account copies are billed as separate storage in the destination. A plan that copies every daily recovery point to a second region for a year is paying for 365 additional recovery points in that region, and the exam will sometimes present this as an obviously correct answer when a lifecycle rule that expires the copies after 30 days would satisfy the same compliance requirement for a fraction of the cost. Vault lock mode is the last knob and the least negotiable: governance mode can be bypassed by a sufficiently privileged principal, compliance mode cannot be removed by anyone, including the account root, until the retention period expires. Compliance mode is the correct answer for any scenario that mentions regulatory retention or ransomware, and it is also irreversible, which is why the exam sometimes tests whether you understand that you cannot undo it.
5. Sizing, Limits and Quotas
The numbers below are the ones that show up in scenario questions, and each is traceable to the AWS Backup documentation or the service quotas page. They are worth knowing because several exam distractors are built from plausible-sounding but incorrect limits, and because quota exhaustion is a real production failure mode for large fleets.
| Item | Value | Why it matters |
|---|---|---|
| Backup plans per account per Region | 100 (soft) | Large orgs hit this and must consolidate plans |
| Backup vaults per account per Region | 100 (soft) | Per-team vaults consume this quickly |
| Recovery points per vault | Unlimited, but lifecycle expiry governs cost | Cost, not a hard cap, is the real limit |
| Concurrent backup jobs | Service-quota governed per resource type | Exceeded quotas cause jobs to queue, not fail |
| Vault lock minimum retention | Set by your policy; cannot be shortened once locked | Compliance mode is irreversible |
| Cross-region copy destinations | Any supported Region | Destination lifecycle is independent of source |
The operational implication of the soft quotas is that they are adjustable, but the request takes time, and a migration that onboards 200 accounts into a shared backup plan will hit the plan limit before it hits any storage limit. The more common sizing problem is not a quota at all but job concurrency: when a large fleet shares a single schedule, thousands of snapshot jobs start simultaneously, and the service throttles them into a queue. The backups still complete, but they complete later than the schedule implies, which quietly degrades your effective RPO during the window. Staggering schedules across resource groups is the standard mitigation, and it is the kind of detail that separates a design that works on paper from one that works at 3 a.m.
6. Failure Modes and What They Look Like in Production
The most dangerous backup failure is the silent one, and the most common silent failure is a copy job that stops succeeding while the source backup continues to report success. Because the copy is asynchronous and independent, the source job's green status tells you nothing about whether the destination vault received anything. The symptom appears only during a recovery attempt, which is the worst possible time to discover it. The first diagnostic move is to check the destination vault's recovery point count against the expected count, not to check the source plan's job history. If the counts diverge, the copy role's permissions or the destination vault policy is the usual cause.
The second failure mode is the backup that succeeds but cannot be restored within the RTO. This is almost always a lifecycle configuration problem: recovery points have transitioned to cold storage, and the restore path now includes a multi-hour retrieval before the restore even begins. The symptom is a recovery drill that passes on paper and fails on the clock. The diagnostic is to attempt a restore from the oldest recovery point you would actually rely on, not the newest, because the newest is the one least likely to have transitioned.
The third is the ransomware scenario itself, and it fails in a way that is easy to misread. If the attacker has compromised an account with backup permissions and the vault is not locked, the recovery points are deleted along with the data, and the backup system reports exactly what a healthy system reports: nothing, because there is nothing left to report. The absence of alarms is the symptom. The only defense is architectural — cross-account copy into a vault the compromised account cannot reach, with vault lock preventing deletion even by the destination account's own administrators. Any design that keeps the only copy of the backup inside the account being protected has not solved this problem, and the exam will test whether you notice that.
7. The Operational and SRE Angle
Backups are a reliability control that is only as good as its last verified restore, and the operational discipline that makes them trustworthy is the same discipline that makes any SLO trustworthy: measure the thing you actually care about, not the proxy. The proxy here is backup job success rate. The thing you care about is restore success rate and restore duration. A team that alarms on job failures and never measures restore time has an SLO it cannot defend, because the number that matters — how long recovery actually takes — is unmeasured until the incident.
The practical shape of this is a scheduled restore drill, run on a cadence, that restores a representative recovery point into an isolated environment and records the wall-clock time from decision to service. That number is your real RTO, and it is almost always worse than the design document claims, because it includes the human decision time, the restore job duration, and the application-level validation before traffic can be shifted back. Feeding that measured number back into the design is what turns an RTO target into an RTO capability.
On the monitoring side, the alarms worth having are: backup job failures by resource type, copy job failures specifically (these are the silent ones), vault recovery point count dropping below an expected floor, and vault lock policy removal attempts. That last one is a security signal as much as a reliability one, and it should page. The runbook shape follows from the failure modes above: first confirm the destination vault has the recovery points you expect, then confirm the restore path does not traverse cold storage, then execute the restore and time it. A runbook that starts with "restore the backup" and does not first verify the backup exists is a runbook that will fail during the incident it was written for.
8. Edge Cases and Exam Gotchas
The gotchas below are the ones that recur across practice exams, and most of them are variations on a single theme: the exam rewards you for knowing which layer owns a given guarantee. Immutability is owned by the vault lock, not the backup plan. Replication is owned by the copy rule, not the vault. Retention is owned by the lifecycle block, and the destination's lifecycle is independent of the source's. Confusing these layers is the most common way to lose points on this topic.
- Vault lock in compliance mode cannot be removed by anyone, including the account root, until retention expires — it is irreversible, and the exam tests whether you know that.
- Cross-account copy requires a destination vault policy that grants the source account permission; the copy fails silently if that policy is missing or later revoked.
- A recovery point copied to another Region is governed by the destination vault's lifecycle, so expiring the source does not expire the copy.
- Backup and restore cannot satisfy an RPO tighter than the backup schedule; if the requirement is minutes, the answer is replication, not backup.
- Cold storage transitions improve cost and worsen effective RTO — a plan that transitions everything is a plan that fails recovery drills.
- Tag-based resource assignment protects new resources automatically; ARN-based assignment does not, and is the usual cause of "the new volume was never backed up."
- An isolated backup account is only isolated if the SCP on the workload OU denies the backup APIs — otherwise the workload account can still reach in.
9. AWS Backup vs. the Services It Gets Confused With
AWS Backup is frequently confused with three other things: native service snapshots taken directly (EBS snapshots, RDS automated backups, DynamoDB point-in-time recovery), replication services (DMS, S3 Replication, Aurora Global Database), and the DR patterns themselves. The distinction that matters is that AWS Backup is a control plane over the first category, a complement to the second, and an implementation detail of the cheapest end of the third. It does not replace native snapshots — it schedules and governs them — and it does not provide replication, because replication is a continuous data-path concern and backup is a periodic point-in-time concern.
| Need | Right tool | Why not the others |
|---|---|---|
| Centralized policy, cross-account copy, immutable retention | AWS Backup | Native snapshots have no cross-account policy layer or vault lock |
| RPO measured in seconds | Replication (DMS CDC, Aurora Global DB, DynamoDB Global Tables) | Backup schedules cannot produce sub-minute recovery points |
| Single-service, single-account snapshot automation | Native service backup features | AWS Backup adds governance overhead you may not need |
| Ransomware-resistant retention | AWS Backup with Vault Lock in compliance mode, cross-account | IAM policies alone are bypassable by a compromised admin |
| Full DR with pre-warmed compute | Pilot light or warm standby, with AWS Backup underneath | Backup alone cannot meet a minutes-scale RTO |
The pick-X-when rules are short. Pick AWS Backup when the requirement is centralized governance, cross-account isolation, or immutable retention. Pick native service backups when the scope is a single account and a single service and you want the least machinery. Pick replication when the RPO is tighter than your backup schedule can express. Pick a DR pattern — pilot light, warm standby, active-active — when the RTO is tighter than a restore can deliver, and use AWS Backup as the data-layer foundation underneath it. The exam rarely asks you to choose between these in the abstract; it gives you numbers, and the numbers pick for you.
Hands-on Lab: A Ransomware-Resistant Backup Architecture (45 min)
This lab builds the pattern the exam rewards: a backup plan in a workload account, cross-account copy into an isolated backup account, and a vault lock that survives a compromised administrator. You will need two accounts in the same AWS Organization, and you will deliberately attempt to delete a locked recovery point to prove the guarantee holds.
- Create the isolated backup account's vault. In the backup account, create a vault named
central-backup-vault. Attach a vault access policy that grants the workload account's backup rolebackup:CopyIntoBackupVaultand nothing else. Note the vault ARN — you will need it in step 3. - Create the workload account's vault and plan. In the workload account, create a vault named
workload-vaultand a backup plan nameddaily-governedwith a single rule: daily schedule, 35-day retention, transition to cold storage at 30 days. Assign resources by tag — tag a test EBS volume withBackup=dailyand confirm the plan picks it up. - Add the cross-account copy. Edit the rule to add a copy action targeting the backup account's vault ARN. Run an on-demand backup and confirm a recovery point appears in both vaults. This is the step where most teams discover a missing destination vault policy, so verify the copy job status explicitly rather than assuming success from the source job.
- Lock the destination vault. In the backup account, attach a vault lock policy to
central-backup-vaultin compliance mode with a retention period of 35 days. Confirm the lock is active. - Attempt the attack. Still in the backup account, try to delete the copied recovery point and then try to delete the vault itself. Both should fail. Record the exact error. This is the evidence that the design works, and it is the artifact you would show an auditor.
- Break it deliberately. Remove the destination vault policy grant from step 1, run another on-demand backup, and observe that the source backup succeeds while the copy fails. This is the silent failure mode from section 6, reproduced on purpose.
- Measure the restore. Restore the copied recovery point into a new volume in the workload account and time it end to end. Compare that number against the 35-day retention window and ask whether the cold-storage transition at day 30 would have made this restore miss a 15-minute RTO.
Clean up by deleting the backup plan and the workload vault, then removing the vault lock from the destination vault — note that in compliance mode you cannot remove it early, so plan the retention period before you lock. The lab's real output is not the infrastructure; it is the measured restore time and the two failure reproductions, which together are the difference between a backup strategy you believe in and one you have tested.
Scenario Question Drills (20 min)
Q1. A company needs ransomware-resilient backups that cannot be deleted even by a compromised admin account. What should they configure?
Q2. A business states an RPO of five minutes and an RTO of one hour for a transactional database. Which approach satisfies the RPO?
Q3. A team's source backup jobs all report success, but a recovery attempt finds the destination vault empty. What is the most likely cause?
Q4. A compliance requirement mandates that backup retention cannot be shortened by any principal, including the account root. Which vault lock mode applies?
Q5. A workload can tolerate an RTO of 15 minutes and an RPO of 5 minutes, and the business wants to minimize standing infrastructure cost. Which DR pattern fits?
Q6. A backup plan transitions all recovery points to cold storage after seven days. Recovery drills begin failing their RTO targets. Why?
Q7. An organization wants new EBS volumes to be protected automatically the moment they are created, without editing the backup plan each time. What should they use?
Q8. A team copies every daily recovery point to a second region and retains them for a year. Compliance only requires 30 days of cross-region copies. What is the correct optimization?
Q9. A large fleet shares a single backup schedule, and backups consistently complete hours later than the schedule implies. What is the most likely cause and fix?
Q10. Which metric best represents whether a backup strategy actually meets its stated RTO?
Q11. An organization creates an isolated backup account but workload accounts can still call backup APIs against it. What is missing?
Q12. A team needs to restore a database to a point in time 20 minutes before a corruption event. Which capability applies?
Q13. Which statement about vault lock in compliance mode is correct?
Q14. A scenario requires near-zero RTO and RPO and states that budget is not the primary constraint. Which pairing fits?
Q15. Which of these is a genuine limitation of AWS Backup that the exam expects you to recognize?
Peek into Tomorrow
Everything in this day assumed the workload itself is worth recovering. We designed the backup layer, chose a DR pattern from the cost/RTO spectrum, and measured a restore — but we never asked whether the application can survive the failure long enough for any of that to matter. A restore that takes fifteen minutes is only useful if the application degrades gracefully for fifteen minutes instead of cascading into a full outage the moment one dependency disappears. That is a property of the workload's architecture, not of the backup plan, and no amount of vault locking compensates for a service that fails hard when a single downstream call times out.
Tomorrow's Well-Architected Reliability pillar review puts that question at the center. It treats service quotas and network topology as foundations, distributed system design and graceful degradation as workload architecture, and tested recovery procedures as the failure-management discipline that validates both. The open question this day leaves behind is the one the pillar is built to answer: given that recovery is possible, what has to be true about the workload for recovery to be sufficient?
Sources
- AWS Backup Developer Guide — What Is AWS Backup?
- AWS Backup — Vault Lock and Compliance Mode
- AWS Backup — Cross-Region and Cross-Account Backup
- AWS Backup — Creating a Scheduled Backup Plan
- AWS Service Quotas — AWS Backup Endpoints and Quotas
- AWS Well-Architected Framework — Reliability Pillar
- AWS Whitepaper — Disaster Recovery Options in the Cloud
- AWS Prescriptive Guidance — Backup and Recovery