Day 23 of 70 · Week 4
Day 23 / 70 Week 4 of 14 Phase 2: Compute, Containers & Global Databases

Amazon RDS Multi-AZ vs Read Replicas & Blue/Green Deployments

🕑 ~58 min read · 3 services covered
RDS Multi-AZ Read Replicas RDS Blue/Green Deployments

Recap: What Aurora Changed About the Availability Conversation

Day 22 established that Aurora does not behave like a conventional relational database with a bolt-on replication feature. Its storage layer is the replication mechanism: six copies of every write spread across three Availability Zones, with the compute layer reading and writing through that shared volume rather than shipping redo logs between independent servers. Aurora Global Database then extends the same idea across regions, replicating at the storage layer rather than through binlog shipping, which is why the documented cross-region lag sits under a second and a secondary region can be promoted in under a minute.

That model is the exception, not the rule, and today's material is deliberately the contrast case. Standard RDS engines — MySQL, PostgreSQL, SQL Server, Oracle, MariaDB — still give you three separate, non-overlapping capabilities that people routinely collapse into one mental bucket: Multi-AZ for availability, read replicas for read scale, and Blue/Green Deployments for change safety. None of them substitutes for another, and the exam leans hard on scenarios where a candidate picks the wrong one because the words "replica" and "standby" sound interchangeable. Where Aurora's storage-layer replication blurred the line between durability and read scaling, classic RDS keeps those concerns strictly separate — and that separation is the whole point of this day.

Foundations You'll Need Today

Today's material is about three features that all create a second copy of your database, and the entire day is really about telling those copies apart. Before the distinctions make sense, four ideas need to be on the table: what an Availability Zone actually is, what it means for a copy to be synchronous or asynchronous, what a database endpoint is and why it matters during a failure, and what replication lag measures. None of these are complicated, but the exam assumes you have them cold, and the prose in this day leans on all four without stopping to define them.

Availability Zones and Regions

AWS groups its datacenters into Regions, and each Region is a geographic area like Northern Virginia or Ireland. Inside a Region, AWS operates several Availability Zones, usually three or more. An Availability Zone is one or more physically separate datacenters with their own power, cooling, and network connections, placed far enough apart that a fire, flood, or power failure in one is unlikely to affect another, but close enough that data can move between them quickly. This two-level structure is the reason AWS talks about failure at two different scales. An Availability Zone failure is a localized event — one datacenter cluster goes dark — and the rest of the Region keeps running. A regional failure takes out every Availability Zone in that Region at once, which is rare but is the scenario that separates real disaster recovery from ordinary high availability. When today's material says Multi-AZ protects you against an AZ failure but not a regional outage, that is the distinction being drawn.

Synchronous vs. Asynchronous Replication

Replication means keeping a second copy of your data somewhere else, and the question that determines everything is when the original is allowed to say "your write succeeded." With synchronous replication, the database does not confirm the write to your application until the second copy has also durably stored it. That makes the copy guaranteed current, at the cost of every write waiting on a second machine — which is why synchronous replication is used for a standby sitting in a nearby Availability Zone, where the network round trip is short. With asynchronous replication, the original confirms the write immediately and sends it to the copy afterward, in the background. Writes stay fast, but the copy is always slightly behind, and if the original dies before the copy receives the latest changes, those changes are gone. This single difference — whether the confirmation waits for the copy — is what makes Multi-AZ lossless and read replicas potentially lossy, and it is the most-tested idea in today's material.

Database Endpoints and DNS Failover

When you create a database in AWS, you do not get a fixed IP address to connect to. You get an endpoint, which is a DNS name like mydb.abc123.us-east-1.rds.amazonaws.com. Your application connects to that name, and DNS resolves it to whichever machine is currently serving. This indirection is what makes automatic failover possible without reconfiguring anything. When the primary database fails, AWS promotes the standby and updates the DNS record behind the same endpoint name to point at the newly promoted machine. Your application reconnects to the identical hostname it has always used and reaches a different physical server. That is why today's material keeps saying the endpoint name does not move during a Blue/Green switchover — the whole design depends on applications connecting by name rather than by address. It is also why a failover shows up as a brief burst of connection errors rather than a permanent outage: the name is briefly unresolvable or pointing at a machine that is no longer primary.

Replication Lag

Replication lag is simply how far behind the copy is — the gap between the last change applied on the original and the last change applied on the replica. It is measured in seconds, and it is not a fixed property of the system: it grows when the original is receiving more writes than the replica can apply, when the replica is undersized relative to the primary, or when the network path between them is slow. A lag of zero means the copy is perfectly current; a lag of thirty seconds means anything written in the last half minute may not be visible on the replica yet. This is why today's material treats lag as an operational metric you alarm on rather than an implementation detail, and why a read replica can never be used to satisfy a requirement that demands the copy be exactly current. Lag is the price you pay for the speed that asynchronous replication buys you.

With that grounding, here is why RDS offers three separate mechanisms for three separate problems, and why the exam will not let you substitute one for another.

1. Why This Is on the Exam

Almost every SAP-C02 scenario that involves a relational database eventually forces a choice between three things that sound similar and behave nothing alike. The exam is not testing whether you can recite the definition of Multi-AZ; it is testing whether you can read a requirement like "the database must survive an AZ failure with no data loss" or "reporting queries are saturating the primary" or "we need to move from MySQL 5.7 to 8.0 without a long outage" and map it to exactly one of these three features without over-building. That mapping problem is the reason this day exists, and it sits squarely in Domain 2 (Design Resilient Architectures) with a secondary pull into Domain 4 (Design Cost-Optimized Architectures), because the wrong answer is usually the more expensive one.

The trap is linguistic. Multi-AZ gives you a standby that is a full copy of your data, and read replicas give you copies of your data too, so candidates assume they are variations on the same theme. They are not. A Multi-AZ standby is invisible to your application, cannot serve reads, and exists solely so that a failure can be absorbed without a human being paged. A read replica is visible, addressable, serves reads, and exists to take load off the primary. A Blue/Green Deployment is not a runtime topology at all — it is a change-management mechanism that happens to create a second environment. Three different problems, three different tools, and the exam will hand you a scenario where the requirement sentence contains the deciding keyword if you know what to look for.

There is also a cost dimension that the exam likes to probe. Multi-AZ roughly doubles the instance cost of a database because you are paying for a second instance you never query. Read replicas also cost money but they earn their keep by absorbing read traffic, so the justification is different. Blue/Green Deployments are transient — you pay for the green environment only during the upgrade window — which makes them cheap relative to the risk they retire. A scenario that says "minimize cost while meeting a 15-minute RTO" is often testing whether you understand that Multi-AZ is not a disaster recovery strategy at all, since it protects against an AZ failure inside one region and does nothing for a regional outage.

2. How Each Mechanism Actually Works

Multi-AZ is synchronous block-level replication to a standby in a different Availability Zone. When your application commits a transaction, the primary writes to its storage and the write is acknowledged only after it has been durably recorded on the standby's storage as well. That synchronous acknowledgement is the entire point: it means the standby is never behind, so a failover cannot lose committed transactions. The standby runs the same engine version on the same instance class, sits in a different AZ, and is not reachable by your application — there is no endpoint for it, no connection string, no way to send it a query. When the primary fails a health check, RDS flips the DNS record behind your database's single endpoint to point at the standby, which is promoted to primary. Your application reconnects to the same hostname and, after a brief connection interruption, is talking to a different physical instance.

Read replicas work on a completely different principle. They are independent database instances that receive changes asynchronously from the primary, using the source engine's native replication mechanism — binlog-based for MySQL and MariaDB, write-ahead log streaming for PostgreSQL, and engine-specific log shipping for SQL Server and Oracle. Because the replication is asynchronous, a read replica is always some amount of time behind the primary, and that lag is a first-class operational metric rather than an implementation detail. Each replica has its own endpoint, so your application can route read-only queries to it explicitly. A replica can be promoted to a standalone primary, which is how you do a controlled cutover or recover from a primary that is unrecoverable — but promotion is a one-way door, and any transactions the replica had not yet received at the moment of promotion are simply gone.

Blue/Green Deployments are a different category of thing entirely. When you create one, RDS provisions a complete green environment — a primary and, if the source is Multi-AZ, a standby — running the target engine version, and sets up logical replication from the blue (current production) environment to the green one. The green environment is a real, queryable database that you can point a staging application at, run your migration scripts against, and load-test. Because replication is logical rather than physical, the two environments can run different engine versions, which is what makes major-version upgrades possible. When you are satisfied, you trigger a switchover: RDS briefly blocks writes on blue, waits for green to catch up, renames the environments so that green takes over the production endpoint, and then blue becomes the old environment you can delete. The switchover is measured in seconds to a couple of minutes, not hours.

3. The Core Decision Boundary

Every scenario question on this topic reduces to one question: what is the requirement actually asking for? If the requirement is about surviving failure without losing data, the answer is Multi-AZ. If it is about handling more read traffic than a single instance can serve, the answer is read replicas. If it is about changing the database — engine version, schema, parameters — with a rollback path, the answer is Blue/Green. The exam rarely states the requirement in those words, so the skill being tested is translation: "the application must remain available if an Availability Zone fails" is Multi-AZ, "the analytics team's queries are degrading checkout latency" is read replicas, and "we must upgrade the engine with less than five minutes of downtime and be able to revert" is Blue/Green.

The boundary gets interesting when a scenario asks for more than one of these at once, which is common in the harder questions. A production database that must survive an AZ failure, serve a reporting workload, and be upgradeable is a Multi-AZ primary with one or more read replicas hanging off it, and the Blue/Green Deployment is created from that Multi-AZ primary so the green environment inherits the same topology. These compose cleanly because they operate at different layers — Multi-AZ is a storage-level property of the primary, read replicas are separate instances consuming its replication stream, and Blue/Green is a temporary logical-replication overlay. What does not compose is trying to use one to do another's job, and that is where the distractors live.

Requirement in the scenarioCorrect mechanismWhy the others fail
Survive an AZ failure with zero data lossMulti-AZRead replicas are async, so failover loses recent writes; Blue/Green is not a runtime topology
Offload reporting queries from the primaryRead replicaMulti-AZ standby accepts no reads; Blue/Green is temporary and not for production traffic
Upgrade engine version with fast rollbackBlue/Green DeploymentMulti-AZ failover does not change engine version; replica promotion is one-way and lossy
Recover from a regional outageCross-region read replica or snapshot copyMulti-AZ is single-region by definition
Serve reads in another region with low latencyCross-region read replicaMulti-AZ standby is in a different AZ, not a different region
Test a schema migration against production dataBlue/Green DeploymentRead replicas are read-only for the migration's purposes and share the schema

4. Configuration Modes and Their Tradeoffs

Multi-AZ has a small number of knobs and each one has a clear cost. The instance class of the standby mirrors the primary, so enabling Multi-AZ roughly doubles your compute bill for that database — there is no option to run a smaller standby. You choose the standby's Availability Zone or let RDS pick one, and the choice matters only for capacity planning and for avoiding a correlated failure with the primary's AZ. For engines that support it, there is a Multi-AZ cluster deployment option that adds two readable standbys behind a reader endpoint, which is the one configuration where Multi-AZ and read scaling overlap; it is worth knowing it exists because a scenario that asks for both high availability and read offload on a single managed cluster is pointing at it. Failover priority is configurable in the cluster form, letting you control which standby takes over first.

Read replicas are where the configuration space opens up, and the tradeoffs are mostly about lag and cost. You can create replicas in the same AZ, a different AZ, or a different region; cross-region replicas are the standard answer for regional read latency and for a coarse disaster recovery posture, at the cost of higher replication lag because the data crosses region boundaries. Replicas can be chained — a replica of a replica — which reduces load on the primary but compounds lag down the chain. You can promote a replica to a standalone instance, which breaks the replication link permanently and is the mechanism behind both planned cutovers and emergency recovery. The critical tradeoff to internalize is that asynchronous replication means a replica is a lagging copy, and any scenario that requires the replica to be exactly current is a scenario that is actually asking for Multi-AZ.

Blue/Green Deployments trade a temporary doubling of infrastructure for a dramatically reduced change window. The green environment runs the target engine version and, by default, the same instance class as blue, so you pay for a second full database for the duration of the upgrade — typically hours to a few days, not months. You control the switchover timing explicitly, which means you can schedule it for a low-traffic window and you can abandon the whole exercise by deleting green if validation fails. The main constraint is that logical replication does not cover every object type identically across engines, so you must validate that your schema, users, and stored procedures replicated correctly before switching. There is also a switchover timeout: if green cannot catch up within the configured window, the switchover aborts and blue continues serving, which is a safety property rather than a failure.

5. Sizing, Limits and Quotas

The numbers that matter here are mostly about how far each mechanism scales and what it costs you in latency. Multi-AZ failover is documented as typically completing in one to two minutes, driven by DNS propagation and instance promotion rather than by data volume, since the standby is already current. That is the number to quote when a scenario asks how long an AZ failure costs you. Read replica lag is not a fixed number — it depends on write volume, network path, and replica capacity — but the operational rule is that a replica running the same instance class as the primary under a comparable read load will generally keep up, while an undersized replica will fall progressively further behind. Cross-region replicas add the inter-region network round trip to every replication batch, so their lag is structurally higher than same-region replicas.

Blue/Green switchover is documented as typically under a minute of write downtime, which is the figure that makes it the answer to "minimize downtime during a major version upgrade" scenarios. The green environment must be able to keep up with the blue write rate during the replication phase, so an undersized green instance is a common cause of a switchover that never converges. There are also engine-version constraints: Blue/Green Deployments support specific source and target version pairs, and not every engine supports them at all, so a scenario involving an engine that lacks the feature has to fall back to snapshot-restore or replica-promotion strategies. Read replicas have a per-primary limit that varies by engine, and Multi-AZ cluster deployments have their own instance-count ceilings.

PropertyMulti-AZRead replicaBlue/Green
Replication typeSynchronous, block-levelAsynchronous, engine-native logsLogical replication
Data loss on failoverNone for committed transactionsPossible — whatever had not replicatedNone if switchover completes cleanly
Typical failover/switchover time1-2 minutesPromotion takes minutes; no automatic failoverUnder a minute of write downtime
Serves readsNo (except Multi-AZ cluster reader endpoint)Yes, via its own endpointYes, but only for validation, not production
Cross-region capableNoYesNo — single region
Cost profile~2x instance cost, permanentPer-replica cost, permanent~2x cost, temporary

6. Failure Modes and What They Look Like in Production

The most common production failure with read replicas is lag that nobody is watching until it becomes a correctness bug. The symptom is subtle: a user writes a record, the application immediately reads it back from a replica, and the record is not there. This is not a database bug; it is the application being written as though replication were synchronous. The first diagnostic move is to check the replica lag metric — ReplicaLag in CloudWatch for MySQL and MariaDB, and the equivalent for other engines — and correlate it with write volume. If lag spikes during batch jobs or bulk imports, the fix is usually to route those reads to the primary or to size the replica up. If lag grows monotonically and never recovers, the replica has fallen too far behind and may need to be recreated from a snapshot.

Multi-AZ failures are rarer but more disruptive when they happen, because the failover is automatic and the application experiences it as a connection drop. The symptom is a burst of connection errors lasting roughly the failover window, followed by recovery with no operator action. The diagnostic move is to check the RDS event log for the failover event and confirm the new AZ, then verify that the application's connection pool actually reconnected rather than holding dead connections. A surprising number of "the database is down" incidents after a Multi-AZ failover are actually connection-pool misconfiguration: the database recovered in ninety seconds but the application kept handing out stale connections for twenty minutes. Setting a short connection lifetime and enabling pool validation is the standard mitigation.

Blue/Green failures cluster around the switchover. If logical replication cannot keep up, the switchover times out and aborts, leaving blue in place — which is safe but means your maintenance window was wasted. If replication silently skipped objects that logical replication does not support, you discover it after switchover when the application hits a missing stored procedure or a user that did not carry over. The diagnostic move before switching is to compare object counts and definitions between blue and green, and to run the application's own test suite against green. After a switchover, the old blue environment still exists and can be used to investigate what differed, which is a genuine advantage over an in-place upgrade where the previous state is gone.

7. The Operational and SRE Angle

From an SRE perspective these three features map onto three different reliability concerns, and each needs its own monitoring. Multi-AZ is about availability, so the metric that matters is whether failover actually works — and the only way to know is to test it, which is what the RDS reboot-with-failover operation is for. A database that has been running Multi-AZ for two years without a tested failover is an untested assumption, not a guarantee. The alarm shape is straightforward: alert on the RDS event category for failover events so that an automatic failover is visible to the team even though it required no action, because an unexplained failover is a signal that something upstream is wrong.

Read replicas are about capacity and correctness, so the alarms are lag-based and the SLO implication is about read freshness. A reasonable pattern is to alarm when replica lag exceeds a threshold that your application can tolerate — if the application can accept five seconds of staleness, alarm at ten — and to treat sustained lag as a capacity signal rather than a transient. The runbook for a lagging replica is short: identify whether the cause is write volume, replica size, or a long-running query on the replica blocking replication, then either scale the replica, move the offending reads, or kill the blocking query. The subtle SRE point is that read replicas introduce a consistency boundary into your architecture, and that boundary should be documented as an explicit SLO rather than discovered by a user.

Blue/Green Deployments are a change-management control, so they belong in the deployment pipeline rather than in the on-call rotation. The operational discipline is to treat the switchover as a deployable event with a rollback plan, to validate green with production-shaped traffic before switching, and to keep blue around long enough to investigate if something goes wrong. The metric worth tracking is switchover duration over time, because a switchover that used to take thirty seconds and now takes four minutes is telling you that green is struggling to keep up with blue's write rate — a leading indicator that the next upgrade will be harder.

8. Edge Cases and Exam Gotchas

The single most-tested gotcha is that Multi-AZ does not scale reads. Candidates see a standby that holds a full copy of the data and assume it can serve queries; it cannot, and any scenario that pairs "high availability" with "offload reporting" is testing whether you know that. The exception is the Multi-AZ cluster deployment, which does expose a reader endpoint — but that is a distinct configuration, not the default, and the exam will usually signal it explicitly if it is in play. The second most-tested gotcha is that Multi-AZ is not disaster recovery. It protects against an AZ failure within one region; a regional outage takes both the primary and the standby with it, and the answer for that is a cross-region read replica, a cross-region snapshot copy, or an Aurora Global Database if the engine allows it.

On the replica side, the gotcha is promotion. Promoting a read replica is irreversible in the sense that the replica becomes an independent primary and stops following the original; there is no "demote back" operation. It is also lossy, because the replica was behind at the moment of promotion. A scenario that says "we need to fail over to the replica with no data loss" is internally contradictory for standard RDS and the correct answer is usually to point out that Multi-AZ is the mechanism that provides that property. Cross-region replicas add a second gotcha: they cannot be promoted into a Multi-AZ configuration in one step, and the promotion process for a cross-region replica is a multi-step operation rather than a single API call.

For Blue/Green, the gotchas are about scope and support. Not every engine supports Blue/Green Deployments, and among those that do, not every version pair is supported, so a scenario that specifies an unusual engine or version may be steering you toward a different strategy. Logical replication does not carry every object type, so a database with heavy use of engine-specific features may need manual remediation before switchover. And the switchover is not instantaneous — it involves a brief write block — so a scenario demanding literally zero downtime is not satisfiable by Blue/Green either; the honest answer is "under a minute," and the exam expects you to know the difference between near-zero and zero.

9. This vs. the Services It Gets Confused With

The confusion set here is small but the exam exploits it relentlessly. Multi-AZ gets confused with read replicas because both create a second copy of the data. Read replicas get confused with Multi-AZ for the same reason, in the opposite direction. Blue/Green Deployments get confused with both because they also create a second environment, and with snapshot-restore because both are upgrade strategies. The distinguishing question is always what the second copy is for: absorbing failure, absorbing read load, or absorbing change risk. Once you can answer that in one word, the scenario resolves.

There is also a cross-service confusion worth naming. Aurora has its own versions of all three ideas — Aurora Replicas, Aurora Global Database, and Aurora Blue/Green Deployments — and they behave differently from their standard-RDS counterparts in ways that matter. Aurora Replicas share the same storage volume as the writer, so they have far lower lag than a binlog-based read replica and can be promoted in seconds. Aurora Global Database replicates at the storage layer, as Day 22 covered, which is why its cross-region lag is measured in milliseconds rather than seconds. If a scenario names Aurora, the standard-RDS intuitions about lag and promotion time are the wrong ones to apply.

Pick this…When the requirement is…
Multi-AZSurvive an AZ failure with no committed-transaction loss and no operator action
Multi-AZ clusterSame as above, plus readable standbys behind a reader endpoint
Same-region read replicaOffload read traffic without cross-region latency or cost
Cross-region read replicaServe reads near users in another region, or provide a coarse regional DR target
Blue/Green DeploymentChange engine version, schema, or parameters with a validated rollback path
Snapshot restoreRecover to a point in time, or move data across regions when no replica exists
Aurora equivalentsThe engine is Aurora — lag and promotion characteristics differ substantially

Hands-On Lab: Validate a Major-Version Upgrade with Blue/Green

The goal of this lab is to run a full Blue/Green Deployment cycle against a database that is already Multi-AZ, so you experience how the three mechanisms from this day compose in a realistic topology. Budget about 45 minutes, most of which is waiting for the green environment to provision and catch up. You will need an AWS account where you can create an RDS instance, and you should work in a non-production account or at minimum a non-production VPC.

1. Build the baseline. Create an RDS for MySQL instance on a version that supports Blue/Green Deployments, with Multi-AZ enabled and a small instance class. Note the endpoint and the Availability Zone of the primary. While it provisions, create a database, a table with a few thousand rows, and at least one stored procedure or view so that you have a non-trivial schema to validate later.

2. Add a read replica. Create a read replica of the instance and wait for it to reach the available state. Connect to the replica's endpoint and confirm you can read the data you inserted. Then insert a new row on the primary and immediately query the replica for it — you will likely see it missing for a moment, which is the asynchronous replication lag from section 2 made concrete. Record how long it takes to appear.

3. Generate write load. Run a simple loop that inserts rows into your table continuously, or use a load-testing tool if you have one. The purpose is to give the green environment something to keep up with during the replication phase, which is what makes the switchover timing meaningful rather than trivial.

4. Create the Blue/Green Deployment. From the RDS console, create a Blue/Green Deployment targeting a newer engine version. Watch the green environment provision. Note that it inherits the Multi-AZ topology from blue, so you are now paying for four instances: blue primary, blue standby, green primary, green standby.

5. Validate green. Connect to the green endpoint and compare object counts and definitions against blue. Run your stored procedure. Confirm that the rows your write loop is generating are appearing in green, and check the replication lag between blue and green. This is the step that catches logical-replication gaps before they become production incidents.

6. Switch over. Trigger the switchover and time it. Observe that writes are briefly blocked, that the production endpoint now resolves to what was the green environment, and that your application (or a test client) reconnects without a configuration change because the endpoint name did not move.

7. Inspect the aftermath. The old blue environment still exists. Connect to it and confirm it is now stale — it is no longer receiving writes. This is your rollback path and your investigation surface. Delete it when you are satisfied, and delete the read replica, to stop the meter.

8. Write down what you observed. Record the switchover duration, the maximum replication lag you saw, and whether any schema object failed to carry over. Those three numbers are the ones you would report after a real upgrade, and having seen them once makes the exam's Blue/Green scenarios much easier to reason about.

Scenario Question Drills (20 min)

Q1. A team needs to upgrade RDS MySQL from 5.7 to 8.0 with minimal risk and a fast rollback path. What should they use?

A. In-place major version upgrade during a maintenance window
B. RDS Blue/Green Deployments to validate on a synced green environment before switchover
C. Read replica promotion
D. Multi-AZ failover
Correct answer: B. Blue/Green Deployments create a fully replicated green environment on the new version, validated before a fast, low-risk switchover — reducing the blast radius of major upgrades.

Q2. A production RDS instance must survive the loss of an Availability Zone with no loss of committed transactions and no operator intervention. What should be enabled?

A. A read replica in a second AZ
B. Multi-AZ deployment
C. A Blue/Green Deployment
D. Automated snapshots with a short retention window
Correct answer: B. Multi-AZ replicates synchronously to a standby in another AZ and fails over automatically, so no committed transaction is lost. Read replicas are asynchronous and require manual promotion.

Q3. A reporting workload is saturating the primary database and degrading checkout latency. The reports can tolerate a few seconds of staleness. What is the correct fix?

A. Enable Multi-AZ so the standby can serve the reports
B. Create a read replica and point the reporting queries at its endpoint
C. Create a Blue/Green Deployment and run reports against green
D. Increase the primary instance class
Correct answer: B. Read replicas exist to absorb read traffic and have their own endpoints. A Multi-AZ standby accepts no reads, and a Blue/Green environment is a temporary change-management construct, not a reporting target.

Q4. An application writes a record and immediately reads it back, but the read is routed to a read replica and intermittently returns nothing. What is the most likely cause?

A. The replica is in a different Availability Zone
B. Asynchronous replication lag — the write had not yet reached the replica
C. The replica's security group is blocking the read
D. Multi-AZ is not enabled on the primary
Correct answer: B. Read replicas replicate asynchronously, so a read immediately after a write can miss it. The fix is to route read-after-write queries to the primary or accept a documented staleness window.

Q5. A company's primary region suffers a full regional outage. They have Multi-AZ enabled on their RDS instance. What is their recovery position?

A. Multi-AZ fails over to the standby in the other region automatically
B. Multi-AZ provides no protection — both primary and standby are in the failed region
C. The standby is promoted in a surviving region after a manual step
D. Automated snapshots are restored to the other region with zero data loss
Correct answer: B. Multi-AZ is a single-region construct; the standby lives in a different Availability Zone, not a different region. Regional DR requires a cross-region read replica, cross-region snapshot copy, or an Aurora Global Database.

Q6. A team wants to test a schema migration against a full copy of production data, with the ability to abandon the change and keep serving from the original database. What should they use?

A. A read replica, migrated in place
B. A Blue/Green Deployment, validating green before switchover
C. A Multi-AZ failover to the standby
D. A manual snapshot restore into a new instance
Correct answer: B. Blue/Green gives you a synchronized, queryable copy that you can validate against and then either switch to or discard, which is exactly the abandon-the-change requirement.

Q7. Which statement about promoting a read replica is correct?

A. Promotion is reversible — the replica can be demoted back to following the original
B. Promotion is one-way and can lose transactions that had not yet replicated
C. Promotion preserves all data because replication is synchronous
D. Promotion requires Multi-AZ to be enabled first
Correct answer: B. A promoted replica becomes an independent primary and stops following the source; there is no demote operation, and any writes not yet replicated at the moment of promotion are lost.

Q8. After an automatic Multi-AZ failover, the database is healthy but the application keeps throwing connection errors for twenty minutes. What is the most likely explanation?

A. The failover did not actually complete
B. The application's connection pool is holding stale connections to the old primary
C. The standby is in a different region
D. Read replicas are lagging behind the new primary
Correct answer: B. Failover completes in roughly one to two minutes, but pools that do not validate connections keep handing out dead ones. Short connection lifetimes and pool validation are the standard mitigation.

Q9. A Blue/Green Deployment switchover aborts because the green environment cannot catch up with the blue write rate. What is the state of the system?

A. Blue has been taken offline and the application is down
B. Blue continues serving production; the switchover simply did not happen
C. Green has been promoted with partial data
D. Both environments are now read-only
Correct answer: B. The switchover timeout is a safety property: if green cannot converge, the switchover aborts and blue keeps serving. The usual cause is an undersized green instance.

Q10. Which configuration provides both high availability and readable standbys behind a single reader endpoint?

A. Standard Multi-AZ with a single standby
B. Multi-AZ cluster deployment
C. A read replica in the same AZ as the primary
D. A Blue/Green Deployment left running permanently
Correct answer: B. Multi-AZ cluster deployment adds readable standbys exposed through a reader endpoint, which is the one Multi-AZ form that also serves reads. Standard Multi-AZ standbys accept no queries.

Q11. A team needs low-latency reads for users in a second region and a coarse disaster recovery target, and can accept higher replication lag. What should they create?

A. A second Multi-AZ standby in the other region
B. A cross-region read replica
C. A Blue/Green Deployment in the other region
D. An additional Availability Zone in the primary region
Correct answer: B. Cross-region read replicas serve regional reads and can be promoted for regional recovery, at the cost of structurally higher lag because replication crosses region boundaries.

Q12. Which of these is NOT a valid reason to choose a read replica over Multi-AZ?

A. Offloading analytics queries from the primary
B. Serving reads with low latency in another region
C. Guaranteeing zero data loss on an Availability Zone failure
D. Providing a promotion target for a planned cutover
Correct answer: C. Zero data loss on an AZ failure is precisely what asynchronous replication cannot promise. That requirement belongs to Multi-AZ, whose replication is synchronous.

Q13. A scenario specifies an RDS engine that does not support Blue/Green Deployments, but the team still needs a low-risk major-version upgrade. What is the most reasonable approach?

A. Enable Multi-AZ and rely on failover to perform the upgrade
B. Restore a snapshot into a new instance on the target version, validate it, then cut over
C. Promote a read replica and upgrade it in place
D. Blue/Green Deployments work on every RDS engine, so the premise is wrong
Correct answer: B. When Blue/Green is unavailable, the snapshot-restore-and-validate pattern is the fallback: build the target version separately, test it, then cut over. Multi-AZ failover does not change engine versions.

Q14. During a Blue/Green Deployment, which objects are most likely to require manual remediation before switchover?

A. Tables and their primary keys
B. Engine-specific objects that logical replication does not carry, such as certain stored procedures or users
C. Indexes on the primary key
D. Nothing — logical replication is exhaustive across all object types
Correct answer: B. Logical replication does not cover every object type identically across engines, so stored procedures, users, and similar objects must be validated and often recreated manually before switchover.

Q15. A workload runs on Aurora rather than standard RDS. Which statement about applying today's intuitions is correct?

A. Aurora Replicas behave identically to binlog-based read replicas in lag and promotion time
B. Aurora Replicas share the writer's storage volume, so lag is far lower and promotion is much faster
C. Aurora does not support read scaling at all
D. Aurora Global Database is a Multi-AZ construct within one region
Correct answer: B. Aurora Replicas read from the same distributed storage volume as the writer, so they lag far less than log-shipping replicas and can be promoted in seconds. Aurora Global Database is a cross-region construct, not Multi-AZ.

Peek into Tomorrow: When the Database Itself Is the Bottleneck

Everything in this day assumed a relational engine whose scaling story is vertical: you make the primary bigger, you add replicas to spread reads, and you accept that writes funnel through one instance. That assumption holds until it does not, and the failure is not graceful. A relational primary that is saturated on writes cannot be fixed by adding replicas, because replicas do not accept writes at all — they only multiply your read capacity while the write path stays exactly as narrow as it was. The question this day leaves open is what you do when the write volume itself is the constraint, and when the schema is changing fast enough that migrations are a constant tax rather than an occasional project.

Tomorrow's material turns to DynamoDB, and the first thing it does is move the bottleneck somewhere less obvious. DynamoDB has no primary instance to resize, so the naive assumption is that it scales without limit — but it partitions data by the hash of the partition key, and each partition has its own throughput ceiling. That means a table can throttle even when the table-level capacity looks generous, because a low-cardinality or hot partition key concentrates traffic onto a handful of partitions. The open question worth carrying into tomorrow is how you would even detect that condition, since the table-level metrics look healthy while individual requests fail. The answer involves per-partition visibility and a key-design discipline that has no real analogue in the relational world you have been living in for the last two days.

Sources