Multi-Region Active-Active Architectures
Recap: From Tested Runbooks to Live Multi-Region Traffic
Day 34 framed the scheduled Game Day as the mechanism that converts DR assumptions into tested facts: you simulate an AZ outage or a region loss exercise, and you find out whether the runbook validation you wrote on paper survives contact with a real failure. That exercise is deliberately scoped to a single failure domain at a time, and its output is a set of validated recovery procedures with measured RTO. Today extends that work rather than repeating it. Once you have proven that a region can be evacuated and that the runbook holds, the obvious next question is whether you have to evacuate at all — whether both regions can serve live traffic simultaneously so that losing one is a capacity event rather than a recovery event.
That shift changes the failure model in a way the Game Day format does not prepare you for. A region loss exercise assumes a single authoritative copy of the data and a controlled promotion. Active-active assumes two or more regions accepting writes at the same time, which means the hard problem is no longer failover mechanics but conflict resolution, replication lag, and traffic steering that does not depend on DNS caching behavior. The runbook you validated yesterday still matters, but it becomes a degraded-mode procedure rather than the primary recovery path.
Foundations You'll Need Today
Regions and Availability Zones
AWS runs its infrastructure in geographically separate clusters called regions — Northern Virginia, Ireland, Singapore, and so on. Each region is an independent island: it has its own power, cooling, and network connections, and a failure in one region does not directly affect another. Inside each region are Availability Zones, which are separate data centers (or groups of them) with independent power and networking but connected to each other by very fast, low-latency links. The practical consequence is that an AZ failure is a routine event you design around cheaply, while a region failure is rare but severe, because the surviving region has no fast link to the failed one and cannot simply take over its storage. When this day says "multi-region," it means deliberately running the same workload in two or more of these independent islands at once, which is a much bigger commitment than spreading across AZs within one region.
DNS, TTL, and Why Cached Answers Are Slow to Change
When a user types a domain name, their computer asks a DNS resolver — usually run by their ISP or their company — to translate that name into an IP address. The resolver does not ask the authoritative source every time; it caches the answer for a period called the TTL (time to live), which the domain owner sets. If you change where a name points, resolvers that already have the old answer will keep handing it out until their cached copy expires. Worse, some resolvers and operating systems ignore short TTLs and hold onto answers longer than you asked them to. This is the mechanism behind the day's repeated warning that DNS-based failover is not instant: even if your health check detects a regional failure in seconds, clients whose resolver cached the old address will keep connecting to the dead region until that cache expires, and you cannot control when that happens.
Anycast IP Addresses
Normally an IP address identifies one specific machine in one specific place. An anycast address is different: the same IP address is announced from many locations around the world at once, and the internet's routing system delivers each user's traffic to whichever of those locations is nearest to them. The user does not choose and does not know which one they reached. This is how AWS Global Accelerator can hand you two IP addresses that never change while still serving users from dozens of edge locations: the address is stable, but the physical destination behind it is chosen dynamically by the network. It also means failover can happen without the client re-resolving anything, because the routing decision is made inside the network rather than by the client's DNS cache.
Asynchronous Replication and Replication Lag
When you copy data from one region to another, you can do it synchronously — the write does not count as complete until both copies are updated — or asynchronously, where the write completes immediately in the local region and the copy is shipped to the other region afterward. Synchronous replication across regions would make every write as slow as a round trip across an ocean, so multi-region systems almost always replicate asynchronously. The cost is replication lag: a window of time, usually well under a second but not zero, during which the two regions hold different versions of the same data. Any read that lands in the second region during that window sees the older value. This is why the day keeps returning to replication lag as the leading indicator of consistency problems — it is the size of the window in which the two regions disagree.
Conflict Resolution: What Happens When Two Regions Write at Once
If two regions both accept writes, sooner or later the same record will be modified in both places before either change has replicated to the other. Now there are two versions of the truth and no obvious way to combine them. A database that supports multi-region writes has to pick a rule for resolving this. The most common rule is last-writer-wins: the service compares timestamps and keeps whichever write happened later, silently discarding the other. That is simple and predictable, but it means one of the two updates is lost — which is fine for a user profile where the newest edit is genuinely the right one, and unacceptable for a financial transaction where both writes must be preserved. The alternative is to prevent the conflict from arising at all by ensuring only one region is ever allowed to write a given record, which turns a hard distributed-systems problem into a routing problem. That tradeoff is the central decision of today's topic.
With that grounding, here's why active-active is the hardest point on the disaster recovery spectrum and what problem it actually solves.
1. Why This Is on the Exam
Active-active is the top of the availability ladder, and SAP-C02 treats it as the answer to a specific class of scenario rather than a general best practice. The exam's resilient architectures domain repeatedly presents a workload with a stated RTO and RPO that are both effectively zero — a trading platform, a global consumer application, a payment authorization path — and then asks which architecture satisfies it. The distractor set is almost always a well-designed active-passive pattern: Multi-AZ with a cross-region read replica, a warm standby with Route 53 failover, or a pilot light. Each of those is a correct answer to a different question, and the exam is testing whether you can tell which question is being asked.
The reason active-active is hard to fake is that it is not a single service decision. It is a bundle of three independent decisions that must all be made consistently: how traffic reaches the right region, how data is replicated and reconciled across regions, and how the application behaves when the two regions disagree. Getting the traffic layer right while leaving a single-writer database in place produces an architecture that looks active-active on a whiteboard and fails the first time a region is lost, because the surviving region can read but cannot write. The exam is built to catch exactly that inconsistency, which is why the scenario questions tend to describe the data layer in detail and the traffic layer almost in passing.
There is also a cost dimension the exam expects you to reason about explicitly. Active-active is the most expensive point on the DR spectrum because you are paying for full production capacity in at least two regions at all times, plus cross-region data transfer, plus the operational overhead of a multi-region deployment pipeline. A scenario that says cost is the primary constraint and RTO is measured in minutes is not asking for active-active; it is asking you to recognize that the requirement does not justify the spend. The skill being tested is matching the architecture to the stated RTO/RPO rather than defaulting to the most resilient option available.
2. How Active-Active Actually Works
An active-active architecture has three layers that must each be multi-region, and it is worth walking them in order because the failure modes compound from the bottom up. At the traffic layer, clients need to reach a healthy region without depending on a DNS answer that may be cached far longer than your health check interval. At the data layer, every region that serves writes needs a local copy of the data that it can write to without coordinating with the other region on the request path. At the application layer, the code has to be written so that two concurrent writes to the same logical record produce a deterministic, acceptable outcome rather than a corrupted one.
The traffic layer is where most designs quietly break. Route 53 health checks and failover routing work, but the client's resolver caches the answer for the record's TTL, and some resolvers and operating systems ignore short TTLs entirely. That means a Route 53-only failover can take minutes to propagate to every client even though the health check itself detected the failure in seconds. AWS Global Accelerator sidesteps this by giving you two static anycast IP addresses that never change. Clients connect to the anycast address, the AWS edge network accepts the connection at the nearest point of presence, and Global Accelerator routes it over the AWS backbone to a healthy endpoint group. Because the client never re-resolves anything, failover happens at the network layer in seconds and is invisible to the application.
The data layer is where the real engineering lives. DynamoDB Global Tables replicate a table across regions with a multi-active, last-writer-wins conflict resolution model, so every replica accepts writes and the service reconciles divergent versions by comparing timestamps. Aurora Global Database in its multi-writer configuration allows write forwarding from secondary regions to the primary writer, which is a different model: the secondary region accepts the write and forwards it to the primary, so there is still a single authoritative writer and the conflict problem is avoided at the cost of cross-region write latency. Choosing between these two is the central data decision in any active-active design, and it is driven by whether the workload can tolerate last-writer-wins semantics.
Finally, the application layer has to be designed for the possibility that the same logical entity is being modified in two regions at once. The standard technique is to partition the write space so that conflicts cannot occur in the first place — route all writes for a given customer, tenant, or shard to a single home region, and use the other region for reads and for failover. This is sometimes called a "cell" or "sharded active-active" design, and it is what most production active-active systems actually do, because it converts an unsolvable distributed-systems problem into a routing problem.
3. The Core Decision Boundary: Conflict Model
Every active-active scenario reduces to one question: what happens when two regions write to the same record at the same time? The answer determines which data service you can use, and the data service determines how much of the application you have to rewrite. There are three viable answers, and they are not interchangeable.
The first is to accept last-writer-wins and design around it. DynamoDB Global Tables gives you this for free, and it is the right choice when the data is naturally convergent — session state, user preferences, counters that tolerate approximate values, or any record where the most recent write is genuinely the correct one. The second is to avoid concurrent writes entirely by partitioning the write space, which works with any data store including Aurora and even a relational database, and is the most common production pattern. The third is to use a database that provides a stronger cross-region consistency model, which in AWS practice means accepting the write-forwarding latency of Aurora Global Database multi-writer or moving to a purpose-built globally distributed database.
| Conflict model | Data service | Write latency | Application change | Use when |
|---|---|---|---|---|
| Last-writer-wins | DynamoDB Global Tables | Local (single-digit ms) | Low — accept convergence | Data is convergent or conflicts are tolerable |
| Write forwarding to a single writer | Aurora Global Database (multi-writer) | Cross-region round trip | Low — but latency-sensitive writes suffer | Relational model required, write volume is modest |
| Partitioned write ownership | Any store, including single-writer | Local for the owning region | High — routing layer required | Conflicts are unacceptable and the key space can be sharded |
The exam pattern here is that the scenario will describe the data semantics in a sentence or two and expect you to infer the model. "Users can update their profile from any region and the most recent update should win" is last-writer-wins. "Financial transactions must never be applied twice or out of order" is partitioned ownership or a single writer. "The application is a relational schema with foreign keys and the team cannot rewrite it" points at Aurora Global Database and forces you to accept the write-forwarding latency as the price of the relational model.
4. Traffic Steering Options and Their Tradeoffs
Once the data model is settled, the traffic layer offers several mechanisms that differ mainly in how quickly they react to a regional failure and how much control they give you over the split. The important distinction is between DNS-based steering, which is subject to resolver caching, and network-layer steering, which is not.
Route 53 latency-based routing sends each client to the region with the lowest measured network latency from that client's resolver. It is simple, it requires no additional infrastructure, and it is the right answer when the workload is read-heavy and a few minutes of degraded routing during a regional failure is acceptable. Route 53 geoproximity routing adds a bias parameter that lets you shift traffic toward or away from a region without changing the underlying geography, which is useful for gradually draining a region during a planned maintenance event. Both of these are DNS answers, so both inherit the TTL problem.
Global Accelerator operates at the network layer instead. It provisions two static anycast IP addresses, and clients connect to those addresses permanently. The AWS edge network accepts the TCP connection at the nearest point of presence and forwards it over the AWS global backbone to the endpoint group you have configured. Health checks on the endpoint group remove unhealthy endpoints, and because the client's connection is terminated at the edge rather than resolved through DNS, failover is measured in seconds and does not depend on any client-side caching behavior. The tradeoff is that Global Accelerator is an additional cost and it only fronts TCP/UDP endpoints — it is not a content cache, so it does not reduce origin load the way CloudFront does.
CloudFront is sometimes proposed as the traffic-steering answer, and it is worth being precise about why it usually is not. CloudFront is a content delivery network: it caches responses at edge locations and can be configured with origin groups for failover, but its failover is between origins within a distribution, and its caching behavior means that a regional failure may be masked by cached content rather than actually failed over. It is a good complement to active-active — serving static assets and cacheable API responses from the edge reduces the load that reaches either region — but it is not a substitute for Global Accelerator or Route 53 as the steering mechanism for dynamic, write-bearing traffic.
5. Sizing, Limits and Quotas
The numbers that matter in an active-active design are the ones that bound replication lag, failover time, and the cost of running full capacity twice. These are the figures worth committing to memory, and each is documented in the AWS service documentation linked in the sources below.
| Dimension | Value | Why it matters |
|---|---|---|
| Global Accelerator anycast IPs | 2 static IP addresses per accelerator | Clients never re-resolve, so failover is not TTL-bound |
| Global Accelerator endpoint groups | One per region, with health checks | Traffic dials let you weight regions without DNS changes |
| DynamoDB Global Tables replication | Typically under one second across regions | Bounds the staleness window for last-writer-wins reads |
| Aurora Global Database replication lag | Typically under one second | Secondary-region reads are near-current but not synchronous |
| Aurora Global Database promotion | Secondary promotable in under a minute | Sets the floor on RTO for the write-forwarding model |
| DynamoDB item size | 400 KB maximum | Constrains what can be stored in a single replicated record |
| Cross-region data transfer | Billed per GB in each direction | Replication traffic is a recurring cost, not a one-time one |
The cost model deserves its own sentence because it is the most common reason an active-active design gets rejected in review. You are paying for full production capacity in two regions, which roughly doubles compute and database spend, and you are paying for cross-region replication traffic continuously. For a write-heavy workload, the replication traffic can be a meaningful fraction of the total bill. The exam will sometimes present a scenario where the stated RTO is a few minutes and the budget is constrained, and the correct answer is to recognize that warm standby satisfies the RTO at a fraction of the cost — active-active is only justified when the RTO is genuinely near zero.
6. Failure Modes and What They Look Like in Production
Active-active systems fail in ways that are harder to diagnose than single-region systems, because the failure is often partial and asymmetric. The most common symptom is elevated latency in one region without any error rate increase, which usually means replication lag has grown and reads in the secondary region are returning stale data that the application then has to reconcile. The first diagnostic move is to check the replication metrics for the data layer — DynamoDB's replication latency metric or Aurora's AuroraGlobalDBReplicationLag — before looking at the application at all.
The second common failure is a write storm in one region that saturates the replication path. Because replication is asynchronous, a burst of writes in region A does not slow down region A; it slows down how quickly region B sees those writes. If the application in region B reads its own writes, users in region B will see their own updates disappear and reappear, which is the classic read-your-writes violation. The fix is either to route reads for a given user to the region that owns their writes, or to use a session-consistent read path, and the diagnostic signal is a spike in replication lag correlated with a spike in write throughput in the other region.
The third failure is the one that catches teams by surprise: a regional failure that does not actually fail over. Global Accelerator health checks are configured per endpoint group, and if the health check is pointed at a load balancer that is still returning 200s while the application behind it is broken, the accelerator will keep sending traffic to a region that cannot serve it. This is why the health check should target a deep health endpoint that exercises the application's dependencies, not just the load balancer's own health check. The symptom is a regional outage that produces no failover at all, and the first diagnostic move is to test the health check endpoint manually from outside the region.
Finally, there is the split-brain scenario that partitioned ownership is designed to prevent. If the routing layer that assigns write ownership fails or is misconfigured, two regions can both believe they own the same shard, and both will accept writes. With DynamoDB Global Tables this resolves to last-writer-wins and you lose data silently. With a relational store it can produce constraint violations or duplicate records. The mitigation is to make the ownership assignment itself highly available and to alarm on any write that arrives at a region that does not own the shard.
7. The Operational and SRE Angle
Operating an active-active system means monitoring three things that do not exist in a single-region deployment: replication lag, per-region traffic split, and the health of the steering layer itself. Replication lag is the leading indicator for almost every data-consistency problem, and it should have an alarm with a threshold well below the point at which the application's read-your-writes guarantee breaks. For DynamoDB Global Tables, that is the replication latency metric; for Aurora Global Database, it is the cross-region replication lag metric. Both should be alarmed per region pair, not aggregated, because a problem in one direction is invisible in a global average.
Traffic split is the second signal. Global Accelerator exposes traffic dials and per-endpoint-group metrics, and the expected steady state is a split that matches your capacity plan. A sudden shift in the split — all traffic landing in one region — is either a health check failure or a routing misconfiguration, and either way it means the other region is now carrying double its designed load. The alarm should fire on the split deviating from the expected range, not just on a region going to zero, because a partial shift is often the early symptom of a health check that is flapping.
The third signal is the steering layer's own health. Global Accelerator is a managed service with a high availability SLA, but the health checks you configure are yours, and a health check that is too shallow will fail to detect a real regional problem while a health check that is too aggressive will cause flapping and unnecessary failover. The runbook shape for an active-active system is therefore different from a failover runbook: instead of a sequence of promotion steps, it is a decision tree that starts with "is the traffic split correct?" and branches into "if not, is the health check wrong or is the region actually unhealthy?" The Game Day exercise from yesterday is the right way to validate that decision tree, and it should be run against a single region at a time so that the steering behavior is observable in isolation.
8. Edge Cases and Exam Gotchas
The gotchas in this topic cluster around the gap between what an architecture diagram shows and what the services actually do. The first and most heavily tested is that Global Tables uses last-writer-wins conflict resolution, which means concurrent writes to the same item in two regions do not merge — one of them is simply discarded. Any scenario that requires both writes to be preserved is not a Global Tables scenario, no matter how the rest of the requirements read.
The second gotcha is that Aurora Global Database multi-writer is not the same as multi-region writes with local latency. Secondary regions forward writes to the primary writer, so a write issued in the secondary region pays a cross-region round trip. A scenario that says "writes must complete in single-digit milliseconds in every region" is not satisfiable with Aurora Global Database, and the correct answer is either DynamoDB Global Tables or a partitioned design where each region owns its own shard.
The third is the DNS caching problem, which the exam tests by describing a requirement for "immediate" or "sub-second" failover. Route 53 failover routing cannot guarantee that, because the client's resolver controls the cache. Global Accelerator can, because the client's connection is terminated at the AWS edge and the routing decision happens inside the AWS network. If the scenario emphasizes speed of failover independent of client behavior, Global Accelerator is the answer.
The fourth is that active-active does not eliminate the need for a DR plan — it changes what the DR plan is for. In an active-active system, the failure you are planning for is not a region loss but a correlated failure across regions: a bad deployment pushed to both regions, a data corruption event replicated everywhere, or a control-plane failure that affects the steering layer. The exam occasionally presents this as a follow-up question, and the answer is usually that you need the same backup and point-in-time recovery discipline you would have in a single-region system, because replication faithfully copies corruption.
9. Active-Active vs. the Patterns It Gets Confused With
The exam's distractor set for active-active scenarios is drawn from the other points on the DR spectrum, and the way to keep them straight is to anchor on the RTO/RPO requirement and the cost constraint rather than on the service names. The table below is the decision boundary in its most compressed form.
| Pattern | Traffic | Data | Typical RTO | Pick when |
|---|---|---|---|---|
| Multi-Site Active-Active | Served from 2+ regions simultaneously | Multi-region writes with conflict resolution | Near zero | RTO/RPO are effectively zero and cost is not the primary constraint |
| Warm Standby | Failover via Route 53 or Global Accelerator | Replicated, scaled-down replica running | Minutes | RTO is minutes and you want to avoid full duplicate capacity cost |
| Pilot Light | Failover, then scale up | Core data replicated, minimal compute | Tens of minutes | RTO is tens of minutes and standing cost must be minimized |
| Backup & Restore | Rebuild in the recovery region | Backups only | Hours to days | RTO is measured in hours and cost is the dominant constraint |
The rule of thumb is that active-active is the only pattern where both regions serve live traffic in steady state. Every other pattern has a designated primary and a designated recovery region, and the difference between them is how much of the recovery region is already running. If a scenario describes a standby region that is running but not receiving traffic, it is warm standby, not active-active, even if the data is replicated in both directions. If it describes a region that is running at reduced capacity and scales up on failover, it is pilot light. The exam will often describe the data replication in a way that sounds active-active — bidirectional replication, low lag — while the traffic description makes clear that only one region is serving. Read the traffic sentence first.
Hands-On Lab: Designing an Active-Active Architecture (60 min)
This lab walks through the design decisions for an active-active deployment of a global read-heavy application with a small write path, using Global Accelerator for steering and DynamoDB Global Tables for data. The goal is not to deploy anything but to produce a written design that survives the conflict-resolution questions the exam will ask.
Step 1 — State the requirements explicitly. Write down the RTO and RPO targets, the read and write volume per region, and the data semantics for the write path. For this lab, assume an RTO of under 30 seconds, an RPO of under one second, and a write path where the most recent update to a user profile should win. That last sentence is the one that determines the data service, so write it down before choosing anything.
Step 2 — Choose the data layer and justify it. Given last-writer-wins semantics, DynamoDB Global Tables is the correct choice. Write a short paragraph explaining why Aurora Global Database would be wrong here: it would introduce cross-region write latency for a workload that does not need a relational model, and it would not improve the conflict semantics because the application has already declared that the most recent write wins.
Step 3 — Design the traffic layer. Configure a Global Accelerator with two endpoint groups, one per region, each pointing at a regional Application Load Balancer. Set the traffic dials to a 50/50 split and configure health checks against a deep health endpoint on each ALB that exercises the application's dependency on DynamoDB. Document why the health check must be deep: a shallow check against the ALB's own health endpoint will not detect a regional application failure.
Step 4 — Reason through conflict resolution. For each write path in the application, decide whether last-writer-wins is acceptable. For the user profile update, it is. For any counter or aggregate, decide whether approximate values are acceptable; if not, redesign that path to route writes for a given user to a single home region. Write down the routing rule that implements the home-region assignment.
Step 5 — Define the monitoring. List the alarms you would create: DynamoDB replication latency per region pair, Global Accelerator traffic split deviation from the expected range, per-region ALB error rate, and a synthetic canary that exercises the write path from both regions. For each alarm, write the first diagnostic step the on-call engineer should take.
Step 6 — Write the degraded-mode runbook. Describe what happens when one region is lost. The traffic layer should shift automatically via Global Accelerator health checks; the surviving region now carries double the read load, so the runbook should include a step to verify that the surviving region's capacity can absorb it. Note that no data promotion is required, which is the operational advantage of active-active over every other pattern.
Step 7 — Cost the design. Estimate the steady-state cost as roughly double the single-region compute and database spend, plus cross-region replication traffic. Compare that against a warm standby design that meets the same RTO and write a one-paragraph justification for why active-active is or is not warranted given the stated RTO of 30 seconds.
Scenario Question Drills (20 min)
Q1. An active-active application must fail over between regions in under 30 seconds, and the team has observed that some clients cache DNS answers for several minutes despite short TTLs. Which traffic-steering mechanism meets the requirement?
Q2. A global application stores user profile data in DynamoDB and needs every region to accept writes with single-digit-millisecond latency. Two users update the same profile from different regions within the same second. What does DynamoDB Global Tables do?
Q3. A financial application requires that a transaction is never applied twice and that writes complete in single-digit milliseconds in every region. Which data architecture fits?
Q4. A team enables Aurora Global Database multi-writer so that both regions can accept writes. Users in the secondary region report that write operations take noticeably longer than reads. Why?
Q5. A regional outage occurs, but Global Accelerator continues sending traffic to the failed region. The load balancer in that region is still returning HTTP 200 on its health check endpoint. What is the most likely cause?
Q6. An application uses DynamoDB Global Tables and users in region B report that their own updates briefly disappear and then reappear. What is the most likely explanation?
Q7. Which metric is the leading indicator of data-consistency problems in an active-active deployment using DynamoDB Global Tables?
Q8. A workload has an RTO of 5 minutes and an RPO of 1 minute, and the business has stated that cost is the primary constraint. Which DR pattern should the architect recommend?
Q9. A team wants to gradually drain traffic away from a region during a planned maintenance window without changing DNS records. Which Global Accelerator feature supports this?
Q10. An active-active application is deployed in two regions. A bad configuration change is pushed to both regions simultaneously and corrupts data that is then replicated everywhere. What should the architect have in place?
Q11. Which statement correctly distinguishes Global Accelerator from CloudFront for active-active traffic steering?
Q12. A scenario describes bidirectional data replication with sub-second lag between two regions, but states that only one region receives production traffic while the other runs at reduced capacity and scales up on failover. Which pattern is this?
Q13. An application uses a partitioned active-active design where each region owns a shard of the key space. What is the primary risk this design introduces?
Q14. A team is designing an active-active system and wants to know the maximum size of a single replicated record in DynamoDB. What is the limit?
Q15. An active-active deployment uses Global Accelerator with two endpoint groups. The on-call engineer receives an alarm that the traffic split has shifted from 50/50 to 95/5 without any change in total traffic. What is the first diagnostic step?
Peek into Tomorrow
Active-active is the most expensive point on the DR spectrum, and today's design exercise quietly assumed that the cost was acceptable. The open question it leaves is what you do when it is not. Most workloads do not have an RTO of thirty seconds, and the exam is full of scenarios where the stated RTO is minutes or tens of minutes and the budget is explicitly constrained. In those cases the architect's job is to find the cheapest pattern that still satisfies the requirement, and that means understanding the cost/RTO spectrum as a continuous tradeoff rather than a menu of four discrete options.
Tomorrow's material maps that spectrum explicitly: Backup & Restore at the high-RTO end, Pilot Light with core data replicated and minimal standing infrastructure, Warm Standby with a scaled-down replica already running, and Multi-Site Active-Active at the low-RTO, high-cost end. The interesting design work is in the middle of that range, where the difference between pilot light and warm standby is a matter of how much you are willing to pay to shave minutes off the recovery time. The question to carry forward is where your own workload actually sits on that spectrum once you stop assuming that the most resilient answer is the right one.
Sources
- AWS Global Accelerator Developer Guide — What Is AWS Global Accelerator?
- AWS Global Accelerator — Endpoint Groups, Traffic Dials and Health Checks
- Amazon DynamoDB Developer Guide — Global Tables
- Amazon DynamoDB — Global Tables: How It Works (Conflict Resolution)
- Amazon Aurora User Guide — Using Amazon Aurora Global Database
- Amazon Aurora — Write Forwarding in Aurora Global Database
- Amazon Route 53 Developer Guide — Choosing a Routing Policy
- AWS Whitepaper — Disaster Recovery Options in the Cloud
- AWS Well-Architected Framework — Reliability Pillar