Week Synthesis — Database Selection Matrix & Scenario Drills
Week 4 in Review
Week 4 moved through the data tier in the order the exam tends to test it: Aurora's storage-layer replication, RDS Multi-AZ versus read replicas, DynamoDB partitioning and capacity modes, Global Tables and DAX, S3 storage classes and lifecycle, and finally ElastiCache's Redis-versus-Memcached split. The through-line is that every one of these services is a different answer to the same question — where does the authoritative copy of the data live, and how does a second copy get made? Aurora answers it by replicating at the storage layer six ways across three Availability Zones, which is why a Global Database secondary can be promoted in under a minute without replaying binlogs. DynamoDB answers it by hashing the partition key and letting Global Tables replicate multi-active with last-writer-wins conflict resolution. S3 answers it with storage classes that trade retrieval latency for cost, plus CRR and SRR for asynchronous object copies. ElastiCache answers it by not being authoritative at all — Redis replicates and persists, Memcached does neither, and that distinction is the whole reason a session store scenario points at Redis.
The week also built a vocabulary of failure modes that recur in scenario questions: hot partitions from low-cardinality keys, read replicas that scale reads but lose data on promotion, and cache nodes that vanish without replication. This synthesis day pulls those threads into decision matrices rather than re-teaching any single service.
Foundations You'll Need Today
Today is a decision day, and the decisions all rest on a handful of ideas that the rest of this curriculum has been using without stopping to define. None of them are complicated, but if any one of them is fuzzy, the selection matrix later in this page will read as a list of service names rather than a set of constraints you can reason about. Here are the four that matter most for this particular day.
Regions and Availability Zones
A region is a geographic area — Northern Virginia, Ireland, Tokyo — and each region is a fully independent copy of AWS. When you put a database in a region, it lives there and nowhere else unless you deliberately copy it somewhere. Inside each region are Availability Zones, usually three or more. An Availability Zone is effectively a separate data center with its own power, cooling, and network connections, close enough to the other zones in the region that data can move between them in a couple of milliseconds, but far enough apart that a fire, flood, or power failure in one is unlikely to take out another.
This matters because almost every durability claim in this curriculum is really a claim about zones. When a service says it "survives the loss of an Availability Zone," it means AWS keeps a copy of your data in at least two physically separate buildings, so one building failing does not lose the data. When a service says it is "single-AZ," it means exactly one copy exists, and losing that building loses the data. The phrase "multi-AZ" is shorthand for the first case. Regions matter for a different reason: copying data between regions is slower and more expensive than copying between zones, which is why the exam treats "cross-region" as a meaningfully harder requirement than "multi-AZ."
Replication Lag and Eventual Consistency
When you keep a second copy of data, that copy has to be updated after the original changes. That update takes time — usually milliseconds, sometimes seconds, occasionally longer if the network is congested or the second copy is far away. The gap between "the original changed" and "the copy caught up" is called replication lag. A system where the copy is always perfectly in step with the original is called strongly consistent; a system where the copy is usually in step but might briefly be behind is called eventually consistent.
The practical consequence is the one that trips people up in real life and on the exam: if you write a value and then immediately read it back from the copy, you might get the old value, because the copy has not caught up yet. This is not a bug — it is the tradeoff you accepted in exchange for being able to spread reads across multiple machines. The fix is not to make the copy faster; it is to send reads that must see the latest value back to the original. Keep this in mind, because several questions today hinge on exactly that distinction.
Primary Copies, Read Replicas, and the System of Record
Most databases in this curriculum have one machine that accepts writes — the primary — and optionally one or more machines that hold copies and only answer reads, called read replicas. The primary is the system of record: the authoritative copy that everyone else is derived from. A read replica is a convenience, not a source of truth. If you promote a read replica to become the new primary, it becomes the system of record at that moment, but anything it had not yet received from the old primary is simply gone.
This distinction shows up constantly in scenario questions. A cache is never the system of record — if the cache is lost, the data must still exist somewhere else, or the design is broken. A read replica is not the system of record either, which is why a scenario that says "the user must immediately see their own update" cannot be answered by pointing them at a replica. Whenever a question asks where the authoritative copy lives, it is asking you to identify the system of record.
Partition Keys and Why They Cause Hot Spots
DynamoDB, the key-value database featured today, does not store all your data in one place. It splits the data across many storage units called partitions, and it decides which partition an item belongs to by running the item's partition key through a hash function. The partition key is one of the attributes you designate when you create the table — for example, a customer ID or a device ID. The hash spreads items across partitions so that no single partition has to hold everything.
The catch is that each partition has its own throughput ceiling, independent of the table's overall capacity. If your partition key has only a handful of distinct values — say, a status field with five possible values — then all your writes land on five partitions no matter how much total capacity the table has, and those five partitions throttle while the rest of the table sits idle. This is called a hot partition. The fix is to choose a partition key with many distinct values that also matches how you query the data, which is why the exam asks about key design so often.
With that grounding, here is the selection matrix — the point of today is to read a scenario's constraints and recognize which of these properties it is actually asking about.
The Core Selection Matrix
Almost every data-tier scenario on SAP-C02 reduces to a small number of stated constraints: access pattern, consistency requirement, scale shape, and geographic spread. The services in this week's set are not interchangeable, and the exam is designed so that exactly one of them satisfies all four constraints simultaneously while the others fail on at least one. The fastest way to answer these questions is to read the scenario for the constraint that eliminates the most options, then verify the survivor against the remaining constraints rather than evaluating every service against every requirement.
The most common eliminator is the access pattern. If the scenario describes joins, multi-row transactions, or an existing relational schema that the team does not want to rewrite, the answer is in the Aurora/RDS family and DynamoDB is out regardless of how the scale language reads. If the scenario describes key-value lookups at very high request rates with a flexible or evolving schema, DynamoDB is the answer and the relational options are distractors. If the scenario describes repeated reads of the same small set of items and explicitly mentions latency in microseconds or offloading read pressure from a database, the answer is a cache — and the cache is never the system of record, which is the trap most of these questions set.
The second eliminator is geographic write scope. A scenario that says multiple regions must accept writes simultaneously rules out Aurora Global Database, because Aurora Global DB has a single writer region even in its multi-writer configuration, which is scoped within one region. DynamoDB Global Tables is the only service in this week's set that accepts writes in every replica region by default. A scenario that says one region writes and other regions need low-latency reads is the opposite case, and there Aurora Global Database's sub-second storage-layer replication is the intended answer.
| Constraint in the scenario | Service that satisfies it | Why the others fail |
|---|---|---|
| Relational schema, joins, ACID transactions | Aurora or RDS | DynamoDB has no joins; ElastiCache is not durable; S3 is object storage |
| Key-value access at very high, unpredictable request rates | DynamoDB (on-demand) | Relational engines need capacity planning; caches are not authoritative |
| Microsecond reads of a hot, small item set | DAX or ElastiCache | DynamoDB itself is single-digit milliseconds; S3 is tens of milliseconds |
| Multi-region simultaneous writes | DynamoDB Global Tables | Aurora Global DB has one writer region; RDS replicas are read-only |
| Cross-region reads with sub-second lag, single writer | Aurora Global Database | RDS cross-region replicas lag more and are harder to promote |
| Cheap retention of rarely read objects for years | S3 Glacier Deep Archive | Database storage classes do not exist; ElastiCache is volatile |
| Session state that must survive a node failure | ElastiCache for Redis with replicas | Memcached has no replication; DynamoDB works but is usually overkill |
Consistency Models Across the Week
Consistency is the constraint candidates most often skip past, and it is where the week's services diverge most sharply. Aurora and RDS present a single-writer, strongly consistent primary; reads from a read replica are eventually consistent and can lag, which is why a scenario that says "a user updates a record and must immediately see the change" cannot be answered with a read replica. Aurora Global Database keeps the primary strongly consistent and makes the secondary region eventually consistent with typical sub-second lag, so a scenario that requires read-your-writes in the secondary region is a trap — the correct answer is to route that user to the primary region.
DynamoDB offers eventually consistent reads by default and strongly consistent reads as an option that costs twice the read capacity units. Global Tables replicate asynchronously and resolve conflicts with last-writer-wins, which means the service is eventually consistent across regions by construction. A scenario that demands a globally consistent counter or a globally unique constraint is not solvable with Global Tables alone, and the exam expects you to recognize that rather than reach for the service because it says "global."
S3 provides read-after-write consistency for new object PUTs and eventual consistency for overwrite PUTs and DELETEs in the general case, with strong consistency for list operations. ElastiCache inherits whatever the backing store provides and adds its own staleness window on top, which is why cache invalidation strategy is an application concern rather than a service feature. The practical rule for the exam is that any scenario using the word "immediately" or "must reflect" is testing whether you know which of these layers introduces a lag.
| Service | Default read consistency | Cross-region behavior | Scenario phrase that points here |
|---|---|---|---|
| Aurora / RDS primary | Strong | Not applicable (single region) | "must read its own writes" |
| RDS read replica | Eventually consistent | Cross-region replica lags further | "scale read throughput" |
| Aurora Global Database | Strong in primary, eventual in secondary | Sub-second storage-layer replication | "secondary region readable, fast promotion" |
| DynamoDB | Eventually consistent (strong optional, 2x RCU) | Global Tables: last-writer-wins | "multi-region writes" |
| DAX | Eventually consistent (item cache TTL) | Per-region cache | "microsecond reads of hot items" |
| S3 | Strong for new PUTs and LIST | CRR/SRR are asynchronous | "durable object retention" |
| ElastiCache | Whatever the app writes; TTL-bounded | Global Datastore for Redis only | "offload repeated reads" |
Global Replication: Aurora Global DB vs DynamoDB Global Tables
This is the single most-tested comparison in the week, and the two services are frequently offered as the only two plausible answers in the same question. The distinction is not about which is faster or newer; it is about write topology. Aurora Global Database is a single-writer, multi-reader topology. One region owns the primary cluster and accepts all writes; secondary regions hold read-only replicas that receive changes through the storage layer rather than through binlog shipping. That design is what produces the sub-second replication lag and the ability to promote a secondary to primary in under a minute, and it is also what makes the service unsuitable for a scenario where two regions must both accept writes.
DynamoDB Global Tables is a multi-active topology. Every replica region accepts writes, and the service reconciles concurrent updates to the same item using last-writer-wins semantics based on the timestamp of the write. That makes it the right answer for active-active applications where users in each region write locally, and the wrong answer for anything that requires a globally consistent invariant. A scenario describing a globally unique username registry or a financial ledger that must never double-count is not a Global Tables scenario, even though the words "global" and "multi-region" appear.
The cost and operational profiles differ in ways the exam also probes. Aurora Global Database requires you to manage a primary and at least one secondary cluster, and failover is a deliberate, managed operation — which is why Route 53 Application Recovery Controller and readiness checks pair naturally with it. Global Tables requires no failover operation at all, because there is nothing to fail over to; the tradeoff is that you have accepted eventual consistency and conflict resolution as permanent properties of the system. When a scenario asks for the lowest possible RTO with the least operational ceremony, Global Tables wins. When it asks for a relational schema with a fast, controlled regional failover, Aurora Global Database wins.
| Dimension | Aurora Global Database | DynamoDB Global Tables |
|---|---|---|
| Write topology | Single writer region | Multi-active, all regions write |
| Replication mechanism | Storage layer, not binlog | Service-managed item replication |
| Typical replication lag | Sub-second | Sub-second, region dependent |
| Conflict resolution | Not applicable (one writer) | Last-writer-wins |
| Failover action required | Yes — managed promotion | None — writes already local |
| Schema | Relational, joins, transactions | Key-value, flexible schema |
| Best-fit scenario phrase | "relational, fast regional failover" | "active-active, users write locally" |
Caching and Object Storage Boundaries
ElastiCache and S3 sit at opposite ends of the data-tier spectrum, and both appear in scenarios as supporting cast rather than as the primary answer. ElastiCache exists to absorb read pressure and reduce latency for repeated access to the same data, and the exam tests two things about it: whether you chose Redis or Memcached, and whether you correctly refused to treat the cache as durable storage. Redis is the answer whenever the scenario mentions replication, persistence, failover, pub/sub, or complex data structures such as sorted sets. Memcached is the answer when the scenario describes a simple, horizontally scaled, stateless cache where losing a node's contents is acceptable and multi-threading matters more than durability.
S3 appears in data-tier scenarios as the durable, cheap, effectively unbounded store for objects that do not need to be queried relationally. The exam tests storage class selection against retrieval-time and availability requirements, and lifecycle rules as the mechanism that automates transitions. The recurring trap is One Zone-IA: it is cheaper than Standard-IA, but it does not survive the loss of an Availability Zone, so any scenario that mentions durability across AZs or compliance retention rules it out. Glacier Deep Archive is the correct answer for the lowest-cost class with a retrieval time measured in hours, and it is replicated across multiple AZs, which is exactly why it beats One Zone-IA in those scenarios.
The boundary between these two and the databases is worth stating plainly. A cache is never the system of record, and an object store is never a query engine. Scenarios that describe a cache as the authoritative source, or that ask you to run relational queries directly against S3, are testing whether you will reject the premise rather than pick the closest-sounding service.
| Requirement | Pick | Reject |
|---|---|---|
| Session store surviving node failure | ElastiCache for Redis with replicas | Memcached (no replication) |
| Simple stateless cache, multi-threaded | ElastiCache for Memcached | Redis (unnecessary complexity) |
| Microsecond reads of hot DynamoDB items | DAX | ElastiCache (requires app-level cache logic) |
| Lowest cost, hours-long retrieval, multi-AZ durability | S3 Glacier Deep Archive | S3 One Zone-IA (single AZ) |
| Automatic cost optimization across access patterns | S3 Intelligent-Tiering | Manual lifecycle rules |
| Cross-region object copies for DR | S3 Cross-Region Replication | Lifecycle transitions (same region only) |
Hands-on Lab / Practical Action (45 min)
The lab for a synthesis day is a decision drill rather than a build. Take the seven-row selection matrix above and, for each row, write a two-sentence scenario that would force that row's answer and a one-sentence distractor that a careless reader would pick instead. The goal is to practice reading constraints rather than recalling service features, because that is the skill the exam actually measures.
Start with the multi-region write case, since it is the highest-yield. Write a scenario in which an application serves users in two regions and each region must accept writes locally with sub-second replication. Then write the near-miss version: the same application, but only one region accepts writes while the other serves reads with sub-second lag. The first scenario points at DynamoDB Global Tables; the second points at Aurora Global Database. If you can articulate why the two scenarios differ in one sentence each, you have the comparison locked in.
Next, build the cache-versus-database drill. Write a scenario describing a product catalog read thousands of times per second with occasional updates, and decide whether the answer is DAX, ElastiCache for Redis, or DynamoDB with provisioned capacity. Then change one constraint — say, the catalog must support sorted-set queries for a leaderboard — and observe how the answer shifts to Redis. The point is to feel how a single constraint change moves the answer across services.
Finally, do the storage-class drill. Write four scenarios that differ only in retrieval-time tolerance and AZ-durability requirements, and map each to a storage class: Standard, Standard-IA, One Zone-IA, and Glacier Deep Archive. Pay attention to the One Zone-IA case, because it is the one where a careless reader picks the cheaper option and misses the durability requirement. Finish by writing the lifecycle rule that would automate the transitions you chose, expressed as day thresholds.
Mixed Scenario Quiz (25 Questions)
Q1. An application needs single-digit-millisecond reads and writes at massive, unpredictable scale with a flexible schema, and users in three regions must all be able to write locally. What fits best?
Q2. A relational workload runs on Aurora in us-east-1 and needs a secondary region that serves reads with sub-second lag and can be promoted to primary in under a minute during a regional outage. What should be used?
Q3. A DynamoDB table uses a status attribute with five possible values as its partition key and is throttling under heavy write volume even though the table's provisioned capacity is far from exhausted. What is the cause?
Q4. A gaming leaderboard reads the same small set of top-ranked items thousands of times per second and needs microsecond read latency. What should sit in front of DynamoDB?
Q5. A session store must survive the failure of a cache node without losing session data and must fail over automatically. Which configuration fits?
Q6. An object is rarely accessed, must survive the loss of an entire Availability Zone, and a 12-hour retrieval time is acceptable. Which storage class is the lowest-cost fit?
Q7. A team needs to upgrade an RDS MySQL instance from 5.7 to 8.0 with minimal risk and a fast rollback path. What should they use?
Q8. A read-heavy reporting workload is saturating the primary Aurora instance. The reports tolerate data that is a few seconds stale. What is the appropriate change?
Q9. A globally distributed application requires a globally unique username registry where two users can never claim the same name, even if they register in different regions at the same moment. Which approach is correct?
Q10. A company stores application logs in S3 and wants to minimize storage cost automatically as access patterns change, without writing lifecycle rules for every possible pattern. What should they use?
Q11. A DynamoDB table serves a workload with highly predictable, steady traffic and the team wants the lowest cost per request. Which capacity mode fits?
Q12. An application writes to an Aurora cluster and immediately reads the same row back through a read replica, occasionally seeing stale data. What is the correct fix?
Q13. A compliance requirement states that archived objects must be retained for seven years and cannot be deleted by any user, including administrators, before that period ends. What should be configured?
Q14. A team is choosing between ElastiCache for Redis and ElastiCache for Memcached for a simple page-fragment cache where losing a node's contents is acceptable and the workload is CPU-bound on the cache tier. Which fits?
Q15. An application needs to store JSON documents that are queried by a small number of known keys at very high request rates, with no joins and no multi-item transactions. Which service is the best fit?
Q16. A workload requires a relational schema with foreign keys and multi-statement transactions, and must scale reads to several replicas. Which service fits?
Q17. A team wants to reduce DynamoDB read costs for a workload that repeatedly reads the same items within a short window, without changing application code to manage a cache. What should they use?
Q18. An S3 bucket holds objects that are accessed frequently for the first 30 days, occasionally for the next 60, and then never again but must be retained for compliance. What is the appropriate design?
Q19. A DynamoDB table is being designed for an IoT workload where each device writes frequently and queries are always by device ID. What partition key design is appropriate?
Q20. A company needs a durable, effectively unbounded store for backup files that are written once and read only during a restore. Cost is the primary concern and restores can take hours. What should they use?
Q21. An application uses Aurora Global Database with a primary in eu-west-1 and a secondary in us-east-1. Users in us-east-1 report that after updating their profile they sometimes see the old value on refresh. What is the cause?
Q22. A team needs to run analytical queries against data currently stored in DynamoDB, joining it with data in Aurora. What is the appropriate approach?
Q23. A workload requires a cache that supports sorted sets for a real-time leaderboard and must survive a node failure. Which service fits?
Q24. A company wants to replicate S3 objects to a bucket in another region for disaster recovery, with replication completing within minutes of each upload. What should they configure?
Q25. A scenario states that a workload needs a relational database, must accept writes in two regions simultaneously, and must never lose a committed transaction. Which statement is correct?
Looking Ahead
Everything in this week's matrix assumed you could observe the system well enough to know which constraint was actually binding. That assumption is doing a lot of work. A hot partition, a lagging read replica, and a cache that has silently stopped serving hits all present as "the application is slow," and the selection matrix is useless if you cannot tell which of the three you are looking at. The week's services emit different signals — DynamoDB throttling metrics, Aurora replica lag, ElastiCache hit rate — but a single alarm on any one of them will page you constantly without telling you whether the situation is an incident or normal variance.
Day 29 turns to that problem directly, starting with CloudWatch alarms and the difference between a static threshold and an anomaly detection band. The open question is how to combine several noisy signals into one page that only fires when the pattern is genuinely bad — high latency and elevated error rate together, rather than either one alone. Composite alarms are the mechanism, and the design question is which signals to correlate so that the alarm is sensitive to real incidents without becoming background noise.