Deep Review — Domain 2 Weak Areas (Resilient Architectures)
Recap
Day 59 was a Domain 1 pass — Organizations, Transit Gateway, Direct Connect, and the rest of the multi-account governance and hybrid networking material — driven entirely by whatever the Day 58 distractor analysis flagged as weak. That review had a clean shape to it: the wrong answers in Domain 1 tend to be wrong for structural reasons. A scenario asks for isolation between two VPCs and the tempting answer is a flat Transit Gateway route table; a scenario asks for centralized identity and the tempting answer is per-account IAM users. You can usually reason your way to the right answer from the topology alone, because networking and governance questions are about what is connected to what.
Domain 2 is the same lifecycle stage but a different concern. Compute, databases, and the SRE/DR cluster are not primarily about topology — they are about which failure mode the scenario is actually describing, and the distractors are built to sound like the failure mode you were expecting rather than the one on the page. A question about read scaling and a question about high availability use almost identical vocabulary, and the wrong answer is usually the correct service applied to the wrong problem. That is why this day is organized as pattern entries rather than as a topic sweep: the goal is not to re-learn Aurora, it is to recognize the specific sentence constructions that separate a Multi-AZ answer from a read-replica answer.
Foundations You'll Need Today
Today's material is about resilience — keeping a system working when part of it breaks. That word gets used loosely, so before the pattern entries make sense, it is worth pinning down five ideas that the rest of this page assumes you already have. None of them require hands-on experience; they are all about what problem each mechanism is trying to solve.
Availability Zones and what "Multi-AZ" actually buys you
AWS runs its data centers in groups called regions, and each region is subdivided into several physically separate clusters called Availability Zones, or AZs. They sit in different buildings with independent power and cooling, connected by fast private links, so a fire or flood in one AZ does not take down the others. When a database is described as "Multi-AZ," it means AWS keeps a second, continuously updated copy of your data in a different AZ and can switch over to it automatically if the first one fails. The important thing to internalize is that this copy exists purely as a spare. It is not there to share the workload — it is there so that a failure in one building does not become an outage. That single fact is the reason Multi-AZ shows up as a wrong answer in so many of today's patterns.
Synchronous versus asynchronous replication
When you keep a copy of data somewhere else, you have to decide when the write is considered "done." Synchronous replication means the original waits for the copy to confirm before telling the application the write succeeded. Nothing is ever lost, but every write is a little slower because it has to travel to two places. Asynchronous replication means the original confirms the write immediately and ships it to the copy in the background. Writes stay fast, but there is a window — usually milliseconds to seconds — where the copy is behind. If the original dies inside that window, the writes that had not yet shipped are gone. This trade-off is the entire reason the exam distinguishes Multi-AZ from read replicas, and it is why the phrase "no data loss" in a scenario is such a strong signal.
Read replicas versus a failover standby
A read replica is a second copy of a database that is deliberately made available to answer read queries — reports, dashboards, search — so that the primary database is left free to handle writes. It is a capacity tool, not a safety tool. A failover standby, by contrast, is the Multi-AZ copy described above: it holds the same data but is not meant to serve traffic, and its only job is to take over if the primary dies. The two are easy to confuse because both are "a second copy of the database," and the exam exploits that confusion constantly. The question to ask yourself is whether the scenario is complaining about too much load or about the risk of an outage. Load means read replicas. Outage means a standby.
RTO and RPO
These two acronyms describe how much pain a business is willing to accept during a disaster, and they are the numbers that decide which recovery strategy is correct. Recovery Time Objective, or RTO, is how long the system is allowed to be down — "we can survive fifteen minutes of downtime" is an RTO of fifteen minutes. Recovery Point Objective, or RPO, is how much data you are allowed to lose, measured in time — "we can afford to lose the last five minutes of transactions" is an RPO of five minutes. A tighter RTO or RPO always costs more, because it requires more infrastructure standing by and ready. When a scenario gives you these numbers, they are not decoration; they are the constraint that eliminates most of the answer choices before you even look at the services.
DNS, TTL, and why caching matters for failover
DNS is the system that translates a human-readable name like app.example.com into the numeric address of a server. When a client looks up a name, it does not ask the authoritative source every time — it caches the answer for a period specified by the record's Time To Live, or TTL. A TTL of sixty seconds means a client may keep using a stale address for up to a minute after you change it. This is normally a helpful optimization, but during a failover it becomes a liability: you can update DNS instantly, and still have users hitting the dead region because their resolver is holding onto the old answer. That is why some failover mechanisms deliberately avoid DNS altogether, and why the phrase "must not depend on DNS caching" in a scenario points at a specific class of answer.
With that grounding, here is how Domain 2 questions are actually built — and why the wrong answers are almost always the right service applied to the wrong problem.
Domain 2 Distractor Patterns
Pattern 1 — Multi-AZ offered as a read-scaling answer
The most common Domain 2 miss is a scenario that describes read pressure and offers Multi-AZ as the fix. The tell is in the verb: if the requirement is phrased as "handle more read traffic," "offload reporting queries," or "scale read throughput," the answer is read replicas, and Multi-AZ is a distractor that sounds responsible because it is the thing you are supposed to enable on a production database. Multi-AZ exists to survive an Availability Zone failure. It maintains a synchronous standby in a second AZ and fails over to it automatically, but that standby does not serve reads. Enabling Multi-AZ on a database that is read-bound changes nothing about its read capacity.
The reverse construction is equally common and equally missed. A scenario describes a requirement for automatic failover with minimal data loss and no application changes, and offers read replicas as the answer. Read replicas replicate asynchronously, so promoting one during a failure can lose the writes that had not yet shipped, and the application has to be repointed at a new endpoint. The distinguishing sentence is almost always about the endpoint: Multi-AZ keeps the same endpoint across failover, which is why it is the answer whenever the scenario says the application cannot be reconfigured. On Mock Exam 1 this pattern typically appeared as a two-part question where the first half was read scaling and the second half was failover, and the correct answer required both a read replica and Multi-AZ rather than either one alone.
Pattern 2 — Aurora Global Database versus DynamoDB Global Tables
Both services put data in multiple regions, and both are offered as answers to any question containing the word "global." The fork is whether the scenario needs writes in more than one region. Aurora Global Database has a single writer region and read-only secondary regions; it replicates at the storage layer with sub-second lag and can promote a secondary in under a minute. DynamoDB Global Tables are multi-active — every region accepts writes — with last-writer-wins conflict resolution. A scenario that says "users in both regions must be able to write" is a Global Tables scenario. A scenario that says "a secondary region must be readable with low lag and promotable during a regional outage" is an Aurora Global Database scenario.
The tempting wrong answer is usually the one that matches the database engine the scenario already mentioned. If the question establishes a relational schema with foreign keys and then asks for multi-region writes, the distractor set will include DynamoDB Global Tables, and it is wrong because it abandons the relational model the scenario specified. Read the constraint about the data model before you read the constraint about regions. On the mock, the tell was a phrase like "the application requires transactional integrity across related tables" — that sentence eliminates every key-value option in the list regardless of how well the replication story fits.
Pattern 3 — DAX and ElastiCache as interchangeable caches
A read-heavy DynamoDB scenario with a latency requirement invites two answers: DAX and ElastiCache for Redis. DAX is the correct one when the application already speaks the DynamoDB API and the goal is to shave read latency without adding cache-management code, because DAX is API-compatible and sits in front of the table transparently. ElastiCache is the correct one when the workload needs data structures DynamoDB does not model, needs pub/sub, or is caching something other than DynamoDB items. The distractor is tempting because ElastiCache is the more familiar service and because "add a cache" is directionally right in both cases.
The tell is whether the scenario mentions the DynamoDB API specifically. If it says the application makes GetItem calls and needs microsecond reads for hot items, DAX. If it says the application needs a session store with automatic failover, that is Redis with replicas, and DAX is not in the running because DAX does not store sessions. A related trap: Memcached offered for anything requiring persistence or failover. Memcached has no replication, so a node failure loses the keys on that node — any scenario containing "must survive a node failure" eliminates it immediately.
Pattern 4 — Reserved capacity offered for unpredictable workloads
DynamoDB capacity mode questions are usually decided by one adjective. On-demand capacity is the answer when traffic is unpredictable, spiky, or new — you pay per request and never plan capacity. Provisioned capacity with auto scaling is the answer when traffic is steady and predictable, because it is cheaper per unit of throughput. The distractor is provisioned capacity offered to a scenario describing a launch or a seasonal spike, and it is tempting because provisioned capacity is the "serious production" answer in most other contexts.
The same shape appears in the EC2 cost questions from Week 8, and it is worth noticing that the exam reuses it. Compute Savings Plans are the answer when the scenario wants flexibility across families and regions; EC2 Instance Savings Plans are the answer when the family and region are fixed and the discount matters more. Spot is the answer only when the workload is fault-tolerant and can absorb a two-minute interruption notice. Any scenario that says "cannot tolerate interruption" removes Spot from consideration no matter how large the discount is. On the mock, the Spot distractor appeared in a question about a stateful primary database, which is the clearest possible signal that the option is wrong.
Pattern 5 — Partition key redesign versus capacity increase
A DynamoDB throttling scenario will offer both "increase the table's provisioned capacity" and "redesign the partition key." The correct answer depends on whether the throttling is table-wide or partition-specific. If the scenario says the table is throttling despite ample provisioned capacity, or names a low-cardinality attribute as the key, the problem is a hot partition and more capacity will not help — each partition has its own throughput ceiling, and a key with a handful of distinct values concentrates all writes onto a few partitions. The tell is the word "despite": "throttling despite high table-level capacity" is a hot-partition sentence.
The distractor is tempting because increasing capacity is the obvious first move and it does resolve genuine table-wide throttling. The exam distinguishes the two by giving you the key attribute. If the scenario names the partition key and it is something like a status field or a date, that is the signal. If the scenario says the table is new and traffic grew faster than expected with no mention of the key, capacity is the answer. Read for whether the key is specified before you decide.
Pattern 6 — Composite alarms versus threshold tuning
An alert-fatigue scenario offers three plausible fixes: lower the threshold, lengthen the evaluation period, or build a composite alarm. The correct answer is almost always the composite alarm, because the scenario's complaint is that single-metric alarms fire on events that self-resolve. A composite alarm requires multiple alarm states to be true simultaneously — high latency and elevated error rate together — which is exactly the correlation the scenario is asking for. Lowering the threshold makes the noise worse, and lengthening the evaluation period delays detection of real incidents without removing the false positives.
The tell is the phrase "self-resolve" or "isolated spikes." If the scenario says the alarm fires and then clears on its own, the problem is not sensitivity, it is correlation. A related distractor is deleting the alarm, which is offered as an option in almost every alert-fatigue question and is never correct. On the mock, the composite alarm answer was paired with a distractor about anomaly detection bands; anomaly detection is a legitimate alternative for metrics with a stable baseline, but it does not address the specific complaint about correlated signals, so it was wrong in that context.
Pattern 7 — Synthetics offered for internal performance problems
When a scenario says internal metrics look healthy but users report failures, the answer is CloudWatch Synthetics canaries. Canaries run scripted browser flows from outside your infrastructure, so they catch DNS failures, CDN misconfiguration, and front-end regressions that server-side metrics cannot see. The distractor is X-Ray, which is tempting because X-Ray is the tracing answer and tracing sounds like the right tool for "we cannot see what is happening." X-Ray traces requests that reach your services; it does not tell you that requests are not reaching them.
The reverse is also tested. A scenario describing growing p99 latency across a dozen microservices with no indication of which hop is slow is an X-Ray question, and Synthetics is the distractor. The distinguishing sentence is whether the failure is at the edge or inside the call chain. "Customers cannot load the login page" is edge. "The checkout flow is slow but succeeds" is inside the chain. On the mock, the two appeared in adjacent questions, which is a deliberate design — the exam wants you to notice that the same vocabulary supports both answers and that only the failure location decides it.
Pattern 8 — FIS without stop conditions
Chaos engineering questions almost always include an option that runs the experiment without a safety mechanism. The correct answer pairs the experiment with stop conditions tied to CloudWatch alarms, which abort the experiment automatically if a linked alarm enters ALARM state. The distractor is the same experiment with manual monitoring, and it is tempting because it sounds operationally responsible — someone is watching. Manual monitoring is not a stop condition; it depends on a human noticing and reacting within the blast radius window.
The tell is the word "safely" or "without risking production." Any scenario that asks how to test failure in production safely is asking about stop conditions. A second distractor in this family is running the experiment in a staging environment instead, which is offered whenever the scenario mentions production risk. Staging does not validate production behavior — the whole point of a Game Day is that the failure mode is exercised against real infrastructure with real dependencies. If the scenario says the team wants to validate that automated recovery actually works, staging is wrong because the automation being tested is production automation.
Pattern 9 — Route 53 failover versus Global Accelerator
Both steer traffic away from an unhealthy region, and the fork is whether DNS caching is acceptable. Route 53 failover depends on clients re-resolving, which means the failover time is bounded below by TTL and by client-side resolver behavior you do not control. Global Accelerator uses static anycast IPs and reroutes at the AWS network edge, so failover does not depend on DNS at all. A scenario that says "sub-second failover" or "independent of DNS TTL" is a Global Accelerator scenario. A scenario that says "route users to the lowest-latency healthy region" is a Route 53 latency-based routing scenario.
The distractor is Route 53 offered for the sub-second requirement, and it is tempting because Route 53 is the service most people reach for first when the word "failover" appears. The tell is any explicit statement about DNS behavior or client caching. A related trap is CloudFront offered as a failover mechanism; CloudFront caches and accelerates content but does not perform regional health-based failover of an origin in the way the scenario is describing. On the mock, the Global Accelerator answer was distinguished by the phrase "without relying on client-side DNS caching," which is about as direct a signal as the exam gives.
Pattern 10 — Route 53 ARC versus plain health-check failover
Route 53 ARC is the answer when the scenario needs the failover mechanism itself to survive a regional failure, or needs proof that the standby region is actually ready. Readiness checks continuously validate that the standby has the capacity and configuration to take over; routing controls live on a control plane distributed across regions and partitions so that the act of failing over does not depend on the region that just failed. Plain Route 53 health-check failover has neither property. The distractor is plain failover offered to a scenario about a critical system, and it is tempting because it is simpler and it does work for the common case.
The tell is a phrase about the failover mechanism's own reliability, or about validating standby readiness before relying on it. "The failover control must not depend on a single region" is an ARC sentence. "We need to know the standby can actually handle production traffic" is a readiness-check sentence. A second distractor is a Lambda triggered by a CloudWatch alarm to perform the failover, which is offered because it sounds automated; it is wrong because the Lambda and the alarm both live in the region that may be failing.
Pattern 11 — Backup and Restore offered for a fifteen-minute RTO
The DR strategy spectrum is a cost-versus-RTO trade, and the exam tests whether you can place a stated RTO on it. Backup and Restore has an RTO measured in hours to days because recovery means rebuilding infrastructure and restoring data. Pilot Light keeps core data replicated and minimal infrastructure running, landing around tens of minutes. Warm Standby runs a scaled-down full replica, landing in minutes. Multi-Site Active-Active serves live traffic from multiple regions and is the only pattern that reaches near-zero RTO. A scenario stating a fifteen-minute RTO with a cost-minimization constraint is a Pilot Light scenario.
The distractor is Backup and Restore offered whenever the scenario emphasizes cost, and it is tempting because it is genuinely the cheapest option. The tell is the RTO number itself — if the scenario quantifies an RTO that Backup and Restore cannot meet, cost stops being the deciding constraint. The reverse distractor is Active-Active offered for a workload with a multi-hour RTO and a tight budget, which is wrong because it over-solves the problem at several times the cost. Read the RTO first, then the budget, and only then the pattern.
Pattern 12 — AWS Backup Vault Lock versus snapshot frequency
A ransomware-resilience scenario offers more frequent snapshots, S3 versioning, and AWS Backup with a locked vault in an isolated account. The correct answer is the last one, and the reason is that the first two do not survive a compromised administrative account. More frequent snapshots in the same account can be deleted by the same credentials that were compromised. Vault Lock enforces WORM immutability, so the retention policy cannot be shortened or the recovery points deleted before expiry, even by the root user. Cross-account copy into a separate account means the compromised account cannot reach the backups at all.
The tell is any phrase about the attacker having administrative access, or about backups that "cannot be deleted." Frequency and versioning are distractors because they address durability, not adversarial deletion. A related trap is offering cross-region copy alone; cross-region copy protects against regional loss but not against an attacker with credentials in the source account, because the copy job and its permissions live there too. On the mock, the correct option was the only one that mentioned both isolation and immutability.
Pattern 13 — Fargate offered where DaemonSets or host access are required
EKS compute questions turn on what the workload needs from the node. Fargate profiles run each pod in an isolated micro-VM with no shared node, which means no DaemonSets, no privileged containers, and no host-level agents. Managed node groups give you AWS-managed Auto Scaling Groups with full node access. A scenario that requires a log-collection DaemonSet on every node, or a security agent that needs host visibility, is a managed node group scenario, and Fargate is the distractor. It is tempting because Fargate is the lower-operational-overhead answer and the scenario usually mentions operational burden.
The tell is the word "DaemonSet," or any requirement that implies a persistent per-node process. A second tell is cost: Fargate has a higher per-pod cost, so a scenario that emphasizes cost at high steady-state utilization points at managed nodes even without a DaemonSet requirement. The same shape appears in ECS questions, where awsvpc mode gives each task its own ENI and security group — a scenario asking for per-task network isolation is an awsvpc scenario, and bridge mode is the distractor because it is the older default.
Pattern 14 — Warm pools versus lifecycle hooks
Both are Auto Scaling features that touch instance lifecycle, and the exam separates them by what the scenario is complaining about. A scenario describing slow scale-out because the application takes minutes to bootstrap is a warm pool scenario: pre-initialized stopped instances are promoted into service instead of booting cold. A scenario describing in-flight requests being dropped during scale-in, or logs not being flushed before termination, is a lifecycle hook scenario: the instance is held in Terminating:Wait while a drain script runs. The distractor is whichever feature the scenario did not ask for, and it is tempting because both are described as "managing the instance lifecycle."
The tell is the direction of the problem. Scale-out latency is a warm pool. Scale-in disruption is a lifecycle hook. A scenario that mentions both — slow to start and dropping connections on the way down — needs both features, and the exam will offer a single-feature answer as the distractor. On the mock, the correct option was the one that named the specific wait state, Pending:Wait or Terminating:Wait, rather than describing the feature generically.
Hands-on Lab / Practical Action (45 min)
Rebuild your Domain 2 error log from Mock Exam 1 and the Day 15-42 practice sets, then re-derive each miss against the pattern entries above. The goal is not to re-answer the questions — you already know the answers — but to write down, for each miss, which sentence in the scenario was the tell you failed to read. Work through the missed questions in the order they appear in the mock, and for each one record three things: the option you chose, the option that was correct, and the specific phrase in the question stem that distinguishes them. If you cannot find a distinguishing phrase, that is itself a finding — it usually means you were missing a piece of mechanism rather than a piece of pattern recognition, and the fix is to go back to the relevant day rather than to drill more questions.
Once the log is built, sort the misses by pattern entry. Most people find that their errors cluster into three or four patterns rather than spreading evenly, and the clustering is the useful output. A cluster in Pattern 1 means you are reading "high availability" and "read scaling" as the same requirement; a cluster in Pattern 11 means you are not anchoring on the RTO number before evaluating cost. For each cluster, write a one-line rule in your own words and add it to the cheat sheet you started on Day 58. Keep the rules short enough to scan — "RTO stated first, budget second" is a better rule than a paragraph explaining the DR spectrum.
Then re-do the ten questions you flagged as answered-by-elimination rather than by knowledge. These are the highest-value items in the whole review, because a lucky answer and a wrong answer represent the same gap. For each, cover the options and state the correct answer and its justification out loud before revealing the list. If you cannot produce the justification without the options in front of you, the question goes back into the cluster list regardless of whether you originally got it right.
Finish by writing ten new scenario questions of your own, one per pattern entry that appeared in your cluster list, using the same construction the exam uses: a requirement sentence, four options, and one distractor that is the correct service applied to the wrong problem. Writing the distractor is the exercise — it forces you to articulate the tell explicitly, which is the same skill the exam is measuring. Add these to the Day 58 cheat sheet and carry them into Mock Exam 2 tomorrow.
Budget the time as follows: twenty minutes on the error log and clustering, fifteen minutes on the elimination-flagged questions, and ten minutes writing the new scenarios. If the clustering step runs long, cut the new-scenario writing rather than the elimination review — the elimination review is where the actual gaps are.
Scenario Question Drills (20 min)
Q1. A production RDS for MySQL database is struggling under a reporting workload that runs heavy read queries during business hours. The database already runs Multi-AZ. What should be added to relieve the read pressure?
Q2. An application requires automatic database failover with no application reconfiguration and no data loss. Which RDS configuration satisfies this?
Q3. A global application needs users in both North America and Europe to write to the same dataset with single-digit-millisecond latency, using a flexible schema. Which service fits?
Q4. A relational application with foreign-key constraints and transactional integrity requirements must survive a full regional outage with a secondary region readable at low latency. Which option fits?
Q5. A DynamoDB-backed application makes GetItem calls for a small set of extremely hot items and needs microsecond read latency without adding cache-management code. What should be added?
Q6. A session store must survive a cache node failure without losing data and must fail over automatically. Which configuration fits?
Q7. A new DynamoDB table is being launched for a product with unknown and highly variable traffic. Which capacity mode should be selected?
Q8. A DynamoDB table is throttling under write load despite having provisioned capacity far above the observed request rate. The partition key is an order status attribute with five possible values. What is the correct fix?
Q9. On-call engineers are paged repeatedly by a latency alarm that fires during brief spikes and clears on its own. What change reduces the noise while preserving detection of real incidents?
Q10. Internal CloudWatch metrics show all servers healthy, but customers report that the login page fails to load. What monitoring should be added to catch this class of failure?
Q11. A microservices application has rising p99 latency, but it is unclear which of twelve services is responsible. What should be used to identify the bottleneck?
Q12. A team wants to test EC2 instance failure in production without risking an uncontrolled outage. What must the FIS experiment include?
Q13. A critical application must fail over between regions in under a second, and the failover must not depend on client-side DNS caching. What should be used?
Q14. A financial system needs failover control that itself survives a regional outage, plus continuous validation that the standby region has the capacity to take over. What should be configured?
Q15. A workload can tolerate a 15-minute RTO and a 5-minute RPO, and the business wants to minimize standing infrastructure cost. Which DR pattern fits?
Preview
The pattern entries above are a map of where Domain 2 tends to break, but a map is not the same as a measurement. Everything on this page was derived from a single mock exam and a set of practice questions you have now seen twice, which means some of the improvement you are about to observe may be recognition rather than understanding. The open question Mock Exam 2 has to answer is whether the clustering you found today was real — whether the three or four patterns you wrote rules for were genuinely your weak areas, or whether they were simply the questions you happened to miss on one particular draw of seventy-five.
There is a second question underneath that one, and it is harder to answer. The Day 59 and Day 60 reviews were both targeted, which means they covered the domains the first mock flagged and skipped the ones it did not. If Mock Exam 1 was unrepresentative in either direction, the gaps it missed are still there and will surface tomorrow as misses in domains you have not reviewed at all. Watch for that specifically: a score that improves in Domain 2 but drops in Domain 1 or Domain 3 is not noise, it is the review having shifted attention rather than closed gaps. Mock Exam 2 is a full 75-question, 180-minute sitting under real conditions, and the score delta against Mock Exam 1 is the only signal that matters.