Day 56 of 70 · Week 8
Day 56 / 70 Week 8 of 14 Phase 4: Migration, Hybrid & Cost Optimization

AWS Database Migration Service (DMS) — Homogeneous & Heterogeneous Migration with CDC

🕑 ~58 min read · 5 services covered · 15 scenario questions

DMS and SCT questions are among the most reliable topics on SAP-C02 — expect three or four direct or scenario-embedded questions. This lesson covers the architecture, the migration-type decision tree, load strategies, sizing, LOB handling, validation, and the edge cases the exam favors, plus 15 practice questions.

AWS DMS AWS SCT Change Data Capture (CDC) DMS Serverless Aurora / RDS

Recap of Day 55

Day 55 closed out the cost-optimization thread with the machinery that makes AWS spend legible: cost allocation tags activated for billing, AWS Budgets watching both actual and forecasted spend, and the Cost and Usage Report as the granular line-item source that feeds custom FinOps tooling. The enforcement detail worth carrying forward is that tag hygiene is a governance problem before it is a reporting problem — a mandatory tagging policy backed by an SCP or a Config rule is what makes per-business-unit cost reports meaningful, because a report built on inconsistently tagged resources is just a more expensive way to be uncertain.

That work is about knowing what you are spending on infrastructure you already run. Today's topic is the adjacent concern on the same service-lifecycle thread: physically moving that infrastructure without breaking it. A migration project is where cost discipline and architectural discipline collide, because the decisions that determine whether a cutover succeeds — which load mode, how much replication capacity, whether the schema needs conversion — are made weeks before the first row moves, and the bill for getting them wrong is measured in downtime rather than dollars. The tagging and budgeting habits from Day 55 do not disappear during a migration; they are what let you attribute the cost of the parallel-running source and target, and what tell you when the old environment is finally safe to decommission.

Foundations You'll Need Today

This lesson moves fast through a lot of database vocabulary, and most of it is assumed rather than explained. Before the DMS material proper, here are the five ideas the rest of the page quietly depends on. None of them are complicated on their own; the difficulty is that the exam expects you to hold all of them at once while reading a scenario.

Database engines are separate products, not interchangeable parts

When people say "a database," they usually mean two different things at once: the data itself, and the software that stores and serves that data. The software is the engine, and the major ones — Oracle, Microsoft SQL Server, PostgreSQL, MySQL, and MariaDB — are genuinely different products built by different companies over decades. They all speak a common language called SQL for basic queries, which is why they look similar from a distance, but each one has its own dialect, its own way of writing stored procedures, its own data types, and its own internal file formats. You cannot take a database file from Oracle and hand it to PostgreSQL and expect it to open, any more than you can run a Windows program on a Mac without some kind of translation layer. This is the single most important fact behind today's lesson: whether the source and target run the same engine or different engines changes the entire migration plan, because a same-engine move is just copying data while a cross-engine move means rewriting the schema and the code that sits on top of it.

RDS and Aurora are AWS running the database for you

Historically, running a database meant buying a server, installing the engine on it, patching the operating system, tuning the configuration, setting up backups, and handling failover yourself. Amazon RDS is AWS's answer to that: you pick an engine — RDS supports several, including SQL Server, Oracle, PostgreSQL, and MySQL — and AWS provisions the server, installs and patches the engine, takes automated backups, and can stand up a standby copy in a second data center automatically. You still connect to it and query it exactly as you would any database; you just do not administer the machine underneath. Amazon Aurora is a related but distinct service: it is a database engine that AWS built itself, designed to be compatible with either MySQL or PostgreSQL, so applications written for those engines can often talk to Aurora without changes. That is why this lesson keeps writing things like "Aurora PostgreSQL" and "RDS SQL Server" — the first word is the AWS service, the second is the engine dialect it speaks. When a migration scenario names a target like that, it is telling you which engine family you are landing in, which is exactly the information you need to decide whether schema conversion is required.

Databases keep a log of every change, and that log is how replication works

Every serious database engine writes down what it is doing before it does it. Before a row is updated, the engine records the intent to update it in a sequential file called a transaction log — Oracle calls it a redo log, MySQL calls it a binary log, SQL Server calls it a transaction log, but the idea is the same everywhere. This log exists primarily so the database can recover from a crash: if the power fails halfway through a write, the engine replays the log on restart and ends up in a consistent state. The useful side effect, and the one this entire lesson rests on, is that the log is a complete, ordered record of every change that has ever been made to the data. If you can read that log, you can reconstruct the current state of the database without ever querying the tables themselves. That is precisely what change data capture does: instead of repeatedly asking "what does this table look like now," DMS reads the log and says "here are the twelve changes that happened in the last second, apply them to the target." It is faster, it does not disturb the source, and it is why a migration can keep a target synchronized with a live production database indefinitely.

Primary keys, indexes, sequences, and triggers

Four pieces of ordinary relational database furniture come up repeatedly in the failure-mode and gotcha sections, so they are worth pinning down. A primary key is the column (or set of columns) that uniquely identifies each row in a table — an employee ID, an order number, a customer email. The database enforces that no two rows share one, and it builds an internal lookup structure around it so that finding a specific row is fast. An index is the same idea applied to other columns: a sorted lookup structure that lets the database find rows by, say, last name without reading the entire table. A sequence (or an auto-increment column, which is the same concept under a different name) is a counter the database maintains to hand out the next available ID number automatically when you insert a row. And a trigger is a small piece of logic the database runs automatically in response to an event — "whenever a row is inserted into the orders table, also write a row into the audit table." The reason all four matter today is that DMS moves rows, and none of these four things are rows. A missing primary key makes replication dramatically slower because the tool has to search for the row it wants to change. A sequence counter does not travel with the data, so the target's counter has to be manually advanced after migration or new inserts will collide with migrated ones. And triggers fire during migration just as they would during normal operation, which can produce duplicate side effects unless they are disabled first.

VPCs and Availability Zones

A VPC, or Virtual Private Cloud, is a private network you define inside AWS. It is the AWS equivalent of the network in an office building: you decide what address range it uses, you carve it into subnets, and you control what can talk to what. Nothing in AWS runs outside a VPC — every database, server, and load balancer lives inside one. An Availability Zone, or AZ, is a physically separate data center within an AWS region, with its own power, cooling, and network connections. A region like us-east-1 contains several AZs, and they are close enough together that traffic between them is fast, but far enough apart that a fire or flood in one does not take out the others. The practical consequence, and the reason this comes up in today's lab and in one of the quiz questions, is that moving data between two AZs is not free and is not instantaneous. It costs a small amount of money in data transfer charges and adds a small amount of latency to every operation. When you place a migration tool, you want it sitting next to the thing it writes to most heavily, because that is where the bulk of the traffic goes.

With that grounding, here is why DMS exists and what problem it actually solves.

1. Why DMS Is on the Exam

AWS Database Migration Service exists to solve one specific architectural problem: moving a live database from one engine, platform, or region to another while the source continues serving production traffic. That constraint — the source stays online — is the entire reason the service exists, and it is the detail most exam scenarios are quietly testing for. When a question describes a company that "cannot tolerate extended downtime," "has a maintenance window of only a few minutes," or "must keep the legacy system available until cutover," the answer is almost always a DMS task configured for continuous replication rather than a one-shot export and import.

On SAP-C02 this maps most directly to the migration and modernization domain, where DMS appears both as a standalone question and as one component inside a larger migration narrative. A typical scenario will describe a portfolio of workloads, mention that one of them is a database with a specific engine pairing, and expect you to select the correct toolchain — which means recognizing that DMS moves data but does not convert schema, that SCT converts schema but does not move data, and that neither of them is the right answer when the question is actually about rehosting a server or syncing a file share. The exam rarely asks "what is DMS" in isolation; it asks you to place DMS correctly inside a decision tree that also contains MGN, DataSync, SCT, and the Snow family.

The second reason this topic carries weight is that it is unusually rich in failure modes that map cleanly onto exam distractors. Missing primary keys, unsynchronized sequences, triggers firing during migration, and undersized replication instances are all real production problems, and each one has a plausible-sounding but wrong answer sitting next to it. Understanding the mechanism well enough to predict which of those problems a given configuration will produce is what separates a candidate who memorized "use Full Load + CDC for low downtime" from one who can reason about a scenario the study guide never covered.

2. How DMS Actually Works

DMS is built from three primitives, and almost every configuration question reduces to understanding how they interact. The first is the replication instance: an AWS-managed compute resource, running the DMS engine, that performs the actual work of reading from the source and writing to the target. It is not a control-plane abstraction — it is a real instance with real CPU, memory, and network limits, and those limits are what determine how fast your migration can go. Memory matters more than most people expect, because the instance buffers rows and cached change records in flight; when memory pressure forces it to spill to disk, throughput collapses in a way that looks like a mysterious slowdown rather than an obvious resource exhaustion.

The second primitive is the endpoint, which is nothing more than a stored connection definition for a source or a target. DMS supports a broad catalog on both sides: Oracle, SQL Server, PostgreSQL, MySQL and MariaDB, MongoDB, and SAP ASE as sources, with RDS, Aurora, Redshift, OpenSearch Service, Kinesis Data Streams, S3, and DynamoDB among the targets. The important architectural point is that endpoints are directional and independent — a source endpoint is read from, a target endpoint is written to, and the same physical database can appear as both in different tasks. This is what makes patterns like ongoing replication into a data lake possible without any notion of a "cutover" at all.

The third primitive is the task, which is where the actual decisions live. A task binds a source endpoint to a target endpoint and then specifies three things: table mappings that select which schemas and tables participate, transformation rules that can rename or filter data in flight, and the migration type that determines whether the task takes a snapshot, streams changes, or does both. Tasks are also the unit of parallelism — a single task has a practical throughput ceiling, so high-volume migrations are split across multiple tasks, often one per large table or per partition, rather than scaled up on a single enormous replication instance.

Underneath all three, DMS reads from the source using whatever mechanism that engine exposes for change tracking — transaction logs, redo logs, binary logs, or change tables — and applies changes to the target as ordinary SQL. That detail matters because it explains most of the service's limitations: if the source engine cannot cheaply identify which row changed, DMS cannot either, and the cost of that identification shows up directly in replication latency.

3. The Core Decision Boundary: Homogeneous vs. Heterogeneous

The single most consequential branching decision in any DMS migration is whether the source and target run the same database engine family. When they do — on-premises SQL Server to RDS SQL Server, or self-managed MySQL to Aurora MySQL — the migration is homogeneous, and no schema translation is required. When they do not — Oracle to Aurora PostgreSQL being the canonical example — the migration is heterogeneous, and the schema itself has to be rewritten before any data can move. The exam tests this boundary constantly, usually by describing an engine pairing in the middle of a longer scenario and expecting you to notice whether a conversion step is implied.

The reason this distinction carries so much weight is that it changes the toolchain, the timeline, and the risk profile all at once. A homogeneous migration is fundamentally a data-movement problem: pre-create the target schema using native tooling, point DMS at both sides, and let it run. A heterogeneous migration is a software-porting problem wearing a database costume, because stored procedures, views, functions, and triggers written against one engine's dialect have to be rewritten against another's, and that work is done by developers rather than by infrastructure.

CharacteristicHomogeneous (same engine family)Heterogeneous (cross-engine)
ExampleOn-prem SQL Server → RDS SQL ServerOracle → Aurora PostgreSQL
Schema conversionNot required, but the target schema must be pre-created with native tools (backup/restore, pg_dump, mysqldump --no-data)Required — AWS SCT converts tables, views, stored procedures, and functions
Secondary objects (indexes, FKs, triggers)Not migrated automatically; script them separatelySCT attempts conversion; complex procedural code usually needs manual rework
Transformation rulesLimitedFull support — renaming, data-type mapping, filtering
Relative cost and complexityLowestModerate to highest

The trap here is the conversion percentage. SCT reports how much of the schema it converted automatically, and a figure in the low nineties sounds like the migration is essentially finished. It is not. The objects SCT cannot convert are disproportionately the complex ones — stored procedures with proprietary extensions, intricate triggers, engine-specific functions — and they are exactly the objects that carry business logic. A 93% automatic conversion rate means the remaining 7% is the part that needs a developer who understands both dialects, and no exam answer that implies a heterogeneous migration completes unattended is correct.

4. Configuration Modes and Their Tradeoffs

Once you have decided whether the migration is homogeneous or heterogeneous, the next set of knobs determines how much downtime the cutover actually costs. DMS offers three migration types, and the difference between them is entirely about what happens to writes that occur while the initial data is being copied. Full Load only takes a point-in-time snapshot of the existing data and writes it to the target. It is the simplest configuration and the fastest to reason about, but it has no mechanism for capturing changes that happen during the copy, which means the source has to be effectively frozen for the duration or you accept that the target is already stale by the time the load finishes. In practice this implies a real maintenance window proportional to the size of the dataset.

CDC only is the mirror image: it replicates ongoing changes from a defined starting point, assuming the base data already exists on the target. This is the right choice when you have loaded the initial dataset through some other means — a native backup and restore, a snapshot, or a previous full load — and now need to keep the target synchronized. It is also the configuration behind the pattern where DMS is not a migration tool at all but a permanent replication feed, streaming database changes into a data lake or analytics target indefinitely with no cutover ever planned.

Full Load plus CDC is the combination that most exam scenarios are pointing at, and it is worth understanding why it works rather than just memorizing it. DMS takes the full-load snapshot while simultaneously caching the changes that occur during the load. When the snapshot completes, it applies the cached changes to bring the target current, then continues streaming CDC from that point forward. The result is that the target stays within seconds of the source indefinitely, and the actual cutover window shrinks to the time it takes to stop writes on the source, let the final changes apply, and repoint the application — seconds to a few minutes regardless of how large the dataset is. Any scenario that stresses minimal downtime, an inability to take the database offline, or a near-zero-downtime cutover is describing this configuration.

Large object handling is a separate axis with its own tradeoff, and it is the one most often gotten wrong. Full LOB mode migrates complete objects of any size but processes rows sequentially, which is slow. Limited LOB mode sets a size threshold and truncates anything above it, which enables parallel processing and is fast — but the truncation is silent, so a threshold set too low loses data without failing loudly. Inline LOB mode handles objects under the threshold inline with the row and migrates larger ones in a second pass, which gets both speed and completeness when the distribution of object sizes is favorable. The exam pattern is a scenario that tells you most LOBs are under some size like 64KB; the correct answer is inline mode at that threshold, not limited mode, because limited mode would discard the occasional larger object.

5. Sizing, Limits and Quotas

Sizing a DMS migration is a matter of matching replication capacity to the volume and change rate of the data, and the failure mode of getting it wrong is a migration that runs but never converges. The replication instance is the unit of capacity in the classic model, and the metrics that tell you whether it is adequate are CPU utilization, freeable memory, and swap usage. Memory is the one people underestimate: the instance buffers rows and cached change records in flight, and when it runs out it spills to disk, which shows up as a throughput cliff rather than a gradual degradation. CPU saturation tends to appear during heavy transformation rules or high-volume CDC apply, where the instance is doing real work per row rather than just moving bytes.

There is also a practical per-task ceiling to plan around. A single DMS task rarely exceeds roughly 100,000 rows per second, which means that very high-volume migrations are parallelized across multiple tasks rather than scaled up on one larger instance. The common patterns are one task per large table, or partitioned table mappings that split a single large table across several tasks. This is worth internalizing because it inverts the intuition that bigger instance equals faster migration; past a certain point the constraint is the task, not the hardware.

DMS Serverless removes the sizing decision entirely by auto-provisioning and auto-scaling DMS Capacity Units based on the workload. The published starting guidance is roughly 2 to 4 DCUs for datasets under 100GB, 4 to 8 DCUs for 100GB to 1TB, and 8 to 16 DCUs above 1TB, scaling upward from there as the workload demands. The exam signal for serverless is a scenario that emphasizes unpredictable or variable migration volume, or that explicitly wants to avoid capacity planning — the same reasoning that makes serverless attractive for any spiky workload.

Signal in the scenarioConfiguration to choose
Dataset under 100GB, steady volumeProvisioned replication instance, small class
100GB to 1TBProvisioned instance, mid class, or DMS Serverless starting around 4-8 DCU
Over 1TBProvisioned instance sized for peak, or DMS Serverless starting around 8-16 DCU
Unpredictable or highly variable volumeDMS Serverless — auto-scales DCUs, no capacity planning
Aggregate throughput above ~100,000 rows/secParallelize across multiple tasks (per table or per partition)

6. Failure Modes and What They Look Like in Production

The most instructive DMS failures are the ones that do not announce themselves. A migration task that is running but falling further behind looks identical in the console to one that is keeping up, unless you are watching the right metric. The two that matter most are CDCLatencySource and CDCLatencyTarget, which split replication lag into its two halves: source-side latency measures how long it takes DMS to read changes out of the source transaction log, and target-side latency measures how long it takes to apply them. When a migration is falling behind, the first diagnostic move is always to check which of those two is growing, because the remediation is completely different depending on the answer.

Rising source-side latency usually means DMS cannot read the change stream fast enough. The classic cause is a table without a primary key or unique index, which forces DMS to identify changed rows by scanning rather than by direct lookup; the cost of that identification grows with table size, so the symptom is a migration that starts fine and degrades as the dataset grows. Rising target-side latency points at the apply side instead, and the usual culprits are an undersized target instance, missing indexes on the target that make each apply expensive, or triggers on the target firing for every replicated row. Both of these are diagnosable from the metrics alone, which is why the exam likes them: the scenario gives you a symptom and expects you to name the metric that would distinguish the two causes.

A third class of failure is silent data loss rather than lag. Limited LOB mode with a threshold set below the actual maximum object size will truncate data without failing the task, and a partitioned table that is not explicitly configured in the table mappings can be partially or entirely skipped. These are the failures that surface after cutover, when the source has already been decommissioned and the discrepancy is discovered by an application rather than by monitoring. The defense is validation, which is covered in the next section, and the reason it is worth calling out here is that these failures are invisible to every metric DMS publishes — the task reports success while having moved the wrong data.

7. Operational and SRE Angle

Treating a DMS migration as an operational event rather than a one-off task is what separates a cutover that goes smoothly from one that becomes an incident. The monitoring baseline is straightforward: replication instance CPU, freeable memory, and swap usage for capacity; CDCLatencySource and CDCLatencyTarget for replication health; and the task-level error and warning counts for anything that has failed outright. The alarms worth configuring before the migration starts are on the two latency metrics, because they are the leading indicator that the target is drifting away from the source, and on freeable memory, because the spill-to-disk cliff is much cheaper to prevent than to recover from mid-migration.

Validation deserves to be treated as a first-class part of the runbook rather than an afterthought. DMS data validation compares row counts and checksums between source and target, and it can run continuously during the CDC phase rather than only at the end. For anything business-critical, that automated comparison should be supplemented with application-level queries that exercise real business logic — a checksum match on a table does not prove that a stored procedure behaves the same way on the new engine. Running validation continuously during CDC means discrepancies surface within minutes, while the source is still authoritative and the migration can still be corrected, rather than after cutover when the only remaining option is a rollback.

The cutover itself should be a written, rehearsed sequence with a defined rollback point. The shape is consistent across migrations: confirm CDC lag is near zero, stop writes to the source, wait for the final changes to apply, verify validation is clean, repoint the application connection string, resume writes, and only then begin the post-migration cleanup. That cleanup is easy to forget and expensive to skip — auto-increment and sequence values on the target need to be advanced past the highest migrated identifier or new inserts will collide, and any triggers disabled before the migration need to be re-enabled. A dry run against a non-production copy of the target is the cheapest way to find the steps the runbook is missing, and it is the same discipline that Day 34's Game Day material applies to failure scenarios generally.

8. Edge Cases and Exam Gotchas

These are the details that turn a correct-sounding answer into a wrong one, and they are worth committing to memory because the exam returns to them repeatedly. The unifying theme is that DMS moves rows, and anything that depends on state outside the rows themselves — identity counters, triggers, secondary schema objects — is not carried along automatically. Every one of the items below is a real production failure that has a plausible distractor sitting next to it on the exam.

  • Tables without a primary key. CDC needs a way to identify which row changed. Without a primary key or at least a unique index, DMS falls back to full-table scans per change, which degrades CDC performance by orders of magnitude as the table grows. Add a key before migrating, or expect this to be the answer to "why is CDC so slow."
  • Sequences and identity columns. Values do not synchronize between source and target. After cutover, a post-migration script must advance the target's sequence above the highest migrated identifier, or new inserts collide with existing rows.
  • Triggers and computed columns. These can fire during migration and produce duplicate writes or constraint violations. The standard practice is to disable triggers before the migration and re-enable them after cutover.
  • Partitioned tables. These require explicit configuration in the table mappings. Mishandling them silently skips data or fails in ways that are hard to diagnose from the task status alone.
  • Homogeneous does not mean everything migrates. Even with no engine conversion, secondary objects — indexes, foreign keys, triggers, stored procedures — are not brought over automatically and need separate handling.
  • DMS does not create the target schema. Homogeneous migrations require the schema to be pre-created with native tools; heterogeneous migrations require SCT's converted schema to be applied to the target before DMS moves any data.

9. This vs. the Services It Gets Confused With

The migration toolchain contains four services that are easy to conflate because they all move something from on-premises to AWS, and the exam exploits that overlap deliberately. The clean way to keep them apart is to ask what unit of work each one operates on: DMS operates on database rows, SCT operates on schema and code, MGN operates on whole servers, and DataSync operates on files and objects. Once you have identified the unit, the correct tool follows almost mechanically, and the distractors stop being tempting.

The pairing that matters most is DMS with SCT, because they are complementary rather than alternatives. A heterogeneous migration needs both, in sequence: SCT converts the schema and reports what it could not convert, a developer remediates the remainder, the converted schema is applied to the target, and then DMS moves the data with CDC to keep downtime minimal. Scenarios that describe an engine change and then offer DMS alone as an answer are testing whether you noticed that the schema conversion step is missing.

ServiceUnit of workPick it when…
AWS DMSDatabase rows, optionally with CDCYou need to move data between databases, or stream database changes continuously, with minimal downtime. It does not convert schema.
AWS SCTSchema and code objectsThe source and target engines differ and stored procedures, views, or functions need conversion. It does not move data.
AWS MGNWhole servers and applicationsYou are rehosting a server or application to EC2 with continuous block-level replication and no application changes.
AWS DataSyncFiles and objectsYou are moving NFS, SMB, or HDFS data to S3, EFS, or FSx on a schedule. It is not database-aware.

The exam's favorite disguise is a scenario that buries the engine pairing in a paragraph of unrelated detail and then offers MGN or DataSync as plausible answers. If the question is about a database and the engines differ, the answer is SCT plus DMS. If the question is about a database and the engines match, the answer is DMS alone with a pre-created schema. If the question is about a server or a file share, neither DMS nor SCT belongs in the answer at all.

10. Hands-on Lab (45 min)

Architect and rehearse a zero-downtime migration from on-premises MySQL to Aurora MySQL.

This lab is deliberately homogeneous, because the point is to isolate the operational mechanics of a cutover from the schema-conversion work that a heterogeneous migration would add. Work through it end to end, including the dry run, and write down every step you had to discover rather than look up — those are the steps your runbook is missing.

  1. Provision a DMS replication instance inside the target VPC, in the same Availability Zone as the Aurora cluster, sized for the source data volume. If you want to compare approaches, provision a second configuration using DMS Serverless and note the DCU range it selects.
  2. Because this is homogeneous, pre-create the target schema with native tooling — mysqldump --no-data against the source, applied to Aurora. Confirm that indexes, foreign keys, and triggers exist on the target before any data moves, since DMS will not create them.
  3. Create the source and target endpoints and test each connection. A failing endpoint test is far cheaper to diagnose now than mid-migration.
  4. Create a migration task set to Full Load + CDC, with table mappings that include the schemas you intend to migrate and exclude anything you do not. Enable data validation on the task.
  5. Start the task and watch the full load complete. While it runs, generate write traffic against the source so there are real changes for CDC to capture.
  6. Once the full load finishes, confirm that CDC has caught up by watching CDCLatencySource and CDCLatencyTarget settle near zero. If either climbs, diagnose which side is lagging before proceeding.
  7. Run a dry cutover against a non-production copy of the application: stop writes, wait for the final CDC apply, confirm validation is clean, repoint the connection string, and resume writes. Time the window and record it.
  8. Perform the post-migration cleanup: advance every auto-increment and sequence value on the target above the highest migrated identifier, and re-enable any triggers that were disabled. Then run an application-level query that exercises real business logic, not just a row count.

11. Scenario Question Drills — 15 Questions

Q1. A company is migrating an on-premises Oracle database to Amazon Aurora PostgreSQL. Stored procedures make heavy use of Oracle-specific PL/SQL extensions. What is the correct migration approach?

A. Use AWS DMS alone with transformation rules
B. Use AWS SCT to convert the schema and code first, remediate any objects SCT cannot auto-convert, then use AWS DMS with Full Load + CDC to migrate the data
C. Use AWS MGN to rehost the database server as-is
D. Manually export and import the data using SQL scripts only
Correct answer: B. This is a heterogeneous migration, so SCT must convert schema and code before DMS moves data. Option A fails because DMS does not convert schema; option C rehosts the server without modernizing the engine, which is not what was asked; option D ignores the schema conversion problem entirely and offers no low-downtime path.

Q2. A migration task shows healthy full-load throughput, but once CDC begins, apply latency on the target climbs steadily. The source table has no primary key. What is the most likely cause?

A. The replication instance is in the wrong region
B. Without a primary key or unique index, DMS must scan the target table to locate each changed row during CDC apply
C. CDC does not support MySQL as a source
D. The target endpoint SSL certificate expired
Correct answer: B. Missing primary keys or unique indexes force DMS into full-table scans to apply each change, and the cost grows with table size. Option A would affect both full load and CDC, not just CDC; option C is false; option D would fail the endpoint test outright rather than degrade apply latency.

Q3. A company must migrate a 4TB production database with a maintenance window of no more than five minutes. Which migration type satisfies that constraint?

A. Full Load only
B. CDC only, with no prior full load
C. Full Load + CDC
D. Schema conversion only
Correct answer: C. Full Load + CDC captures changes during the initial snapshot and keeps streaming them, so the cutover window is just the time to stop writes and let the final changes apply. Option A requires the source to be frozen for the entire 4TB load; option B assumes the base data already exists on the target; option D is not a migration type at all.

Q4. After a successful DMS migration and cutover, new rows inserted into the target Aurora MySQL database are failing with duplicate key errors on the auto-increment primary key. What was missed?

A. Data validation was never enabled
B. The target's AUTO_INCREMENT counter was never advanced above the highest migrated ID
C. The replication instance was undersized
D. LOB mode was set incorrectly
Correct answer: B. Sequence and identity values do not synchronize with DMS, so a post-migration step must advance the counter past the highest migrated value. Option A would have caught data discrepancies but not this; option C affects throughput, not key collisions; option D is unrelated to primary key behavior.

Q5. A profiling step shows 99% of LOB columns in the source table are under 64KB, with a rare few reaching 10MB. Which LOB setting is most efficient without losing data?

A. Full LOB mode
B. Limited LOB mode with a 64KB max size
C. Inline LOB mode with a 64KB inline threshold
D. Disable LOB migration entirely
Correct answer: C. Inline mode migrates small objects fast while still fully migrating the rare large ones in a second pass. Option A works but is needlessly slow because it processes rows sequentially; option B would silently truncate the 10MB objects; option D loses data outright.

Q6. A company migrating on-premises SQL Server to RDS SQL Server is surprised that after the DMS migration, none of their indexes, foreign keys, or stored procedures exist on the target. Why?

A. DMS failed silently
B. Homogeneous migrations do not automatically migrate secondary schema objects — only data
C. SCT should have been used instead
D. CDC mode must be enabled to migrate indexes
Correct answer: B. Even in homogeneous migrations, DMS moves data rather than the full schema, so secondary objects need separate handling with native tools. Option A is wrong because the task reported success; option C is wrong because no engine conversion is needed here; option D is false — CDC does not migrate indexes.

Q7. A CDC task is falling behind and you need to determine whether the bottleneck is reading from the source or applying to the target. Which metrics distinguish the two?

A. CPUUtilization and FreeableMemory only
B. CDCLatencySource and CDCLatencyTarget
C. NetworkIn and NetworkOut
D. FreeStorageSpace
Correct answer: B. CDCLatencySource isolates read-side delay from the source transaction log, and CDCLatencyTarget isolates apply-side delay on the target. Option A tells you about instance capacity but not which side is lagging; option C is too coarse to attribute the delay; option D relates to disk pressure rather than replication position.

Q8. A migration workload's data volume is highly unpredictable week to week, and the team wants to avoid manually sizing and re-sizing a replication instance. What should they use?

A. A dms.r6i.24xlarge provisioned instance, sized for peak
B. DMS Serverless, which auto-scales DMS Capacity Units to match workload
C. Multiple small provisioned instances running in parallel
D. AWS DataSync instead of DMS
Correct answer: B. DMS Serverless is designed for variable workloads and removes manual capacity planning by auto-scaling DCUs. Option A over-provisions for the average case and still requires manual sizing; option C adds operational complexity without removing the sizing decision; option D is the wrong tool for database migration.

Q9. During a heterogeneous migration, SCT reports a 93% automatic schema conversion rate. What is the correct interpretation?

A. The migration is essentially finished; proceed straight to cutover
B. The remaining 7% of objects, often the most complex, still require manual remediation before the migration is production-ready
C. 7% of the data will be permanently lost
D. SCT should be re-run with different settings until it reaches 100%
Correct answer: B. The unconverted objects are disproportionately the complex ones — stored procedures with proprietary extensions, intricate triggers — and they carry business logic. Option A skips the hardest work; option C confuses schema objects with data; option D is not how SCT works, since some objects are simply not auto-convertible.

Q10. A migration must sustain well above 100,000 rows per second in aggregate. What is the correct approach?

A. Provision the largest available replication instance and run a single task
B. Parallelize across multiple tasks, for example one per large table or per partition
C. Switch to CDC-only mode
D. Use AWS DataSync to move the rows instead
Correct answer: B. A single task rarely exceeds roughly 100,000 rows per second, so aggregate throughput above that ceiling requires multiple tasks. Option A hits the per-task limit regardless of instance size; option C changes what is replicated, not how fast; option D is the wrong tool for database rows.

Q11. A migration project includes both a database migration and a large on-premises NFS file share that must land in Amazon EFS. Which tool handles the file share?

A. AWS DMS, configured for file targets
B. AWS DataSync
C. AWS SCT
D. AWS MGN
Correct answer: B. DataSync transfers files and objects between on-premises storage and AWS storage services such as EFS, FSx, and S3. Option A is database-specific and has no file-share target; option C converts schema; option D rehosts servers.

Q12. A table has triggers that fire on INSERT to enforce business logic. During a Full Load + CDC migration, duplicate side effects are observed on the target. What is the standard remediation?

A. Switch to CDC-only mode
B. Disable the triggers before migration and re-enable them after cutover
C. Increase the replication instance size
D. Enable Limited LOB mode
Correct answer: B. Triggers fire during migration unless explicitly disabled, producing duplicate side effects or constraint violations as DMS writes rows. Option A does not stop triggers from firing; option C addresses throughput, not correctness; option D is unrelated to trigger behavior.

Q13. Which pairing correctly matches tool to purpose for a heterogeneous database migration?

A. DMS converts schema; SCT moves data
B. SCT converts schema and code; DMS moves the data
C. MGN converts schema; DataSync moves data
D. A single DMS task handles both automatically with no other tools
Correct answer: B. SCT converts schema and code, and DMS moves data with optional CDC. Option A reverses the two; option C names tools that do not perform either job; option D ignores the schema conversion step that a heterogeneous migration requires.

Q14. A company wants ongoing, continuous replication from an on-premises Oracle database into an S3 data lake in Parquet format for analytics, with no plan to ever cut over the source. What DMS configuration fits?

A. Full Load only, re-run daily
B. Full Load + CDC targeting S3, used as an ongoing CDC pipeline rather than a one-time cutover
C. AWS SCT alone
D. AWS Snowball for a one-time export
Correct answer: B. DMS to S3 is a common pattern for streaming database changes into a data lake continuously, with no cutover ever planned. Option A re-runs a full snapshot daily and cannot keep up with continuous change; option C converts schema but moves no data; option D is a one-time offline transfer with no ongoing replication.

Q15. Where should the DMS replication instance be placed relative to the target database to minimize latency and cross-AZ data transfer cost?

A. In a separate region from both source and target
B. In the same VPC and Availability Zone as the target database
C. In the same VPC and Availability Zone as the source database, regardless of target location
D. Placement does not affect performance or cost
Correct answer: B. Co-locating the replication instance with the target minimizes apply-side latency and avoids cross-AZ transfer charges on the larger volume of data being written. Option A adds cross-region latency and cost; option C optimizes the read side at the expense of the write side; option D is false.

12. Peek into Tomorrow

Day 56 is the last new-topic day of Phase 4, and it closes the migration and cost-optimization arc that began with the 7 Rs on Day 43. Everything from here forward is practice, review, and exam technique — which sounds like a relief until you consider what it actually tests. The decision tree in this lesson is not difficult to recall when you have unlimited time, a page of notes, and a single question in front of you. You can reason from the engine pairing to the toolchain, from the downtime requirement to the load mode, from the LOB distribution to the threshold, and arrive at the right answer deliberately.

The mock exam removes all of that scaffolding. Seventy-five questions in 180 minutes is roughly 2.4 minutes per question on average, and the DMS question will not announce itself as a DMS question — it will arrive as a paragraph about a company with a legacy Oracle system, a compliance deadline, and a constraint buried in the third sentence, sitting between a Transit Gateway routing question and a Savings Plans question. The retrieval has to happen in seconds, under time pressure, while three unrelated services from the previous 55 days are also competing for the same mental bandwidth. That is a different skill from understanding, and it is the skill Phase 5 exists to build.

13. Sources