Day 11 of 70 · Week 2
Day 11 / 70 Week 2 of 14 Phase 1: Multi-Account Governance & Networking

Direct Connect (DX) & DX Gateway

🕑 ~58 min read · 4 services covered
Direct Connect Private VIF Transit VIF BGP

Recap: From Inspection VPC to Private Wire

Day 10 put a centralized inspection VPC in the middle of the network, with AWS Network Firewall applying Suricata-compatible rules to traffic that the Transit Gateway steered into it. That design assumed the traffic was already inside AWS, or arriving over the public internet through a VPN. The appliance VPC pattern and its egress routing via TGW solved the question of where inspection happens, but it left the question of how packets get into AWS in the first place largely untouched.

Today extends that architecture one layer outward. Direct Connect replaces the internet path with a dedicated private circuit from a customer router into an AWS Direct Connect location, and the VIF types determine whether that circuit lands on a single VPC, on a Transit Gateway, or on AWS public service endpoints. The inspection VPC from yesterday does not go away — it becomes the enforcement point for traffic that arrives over the new private path. The same TGW route tables that steered traffic into the firewall now also carry the DX attachment, which means the segmentation decisions from Day 8 and the inspection decisions from Day 10 both apply to the circuit we are about to design.

Foundations You'll Need Today

Today's topic sits on top of several networking ideas that the rest of this curriculum assumes you already have. If you have never built a VPC by hand or configured a router, the terms below are the ones that will otherwise make the rest of this page read like a foreign language. None of them are complicated once the underlying problem is clear, so let's build them up from what each one is actually for.

VPCs and virtual private gateways

A VPC (Virtual Private Cloud) is your own isolated slice of the AWS network — a private address space where you launch servers, databases, and load balancers, and where nothing from the outside can reach them unless you explicitly allow it. Think of it as a private data center floor that AWS has carved out for you, with its own internal IP addressing scheme. A virtual private gateway (VGW) is the component that sits on the edge of a VPC and acts as the doorway to the outside world for private traffic. When you connect your on-premises network to a VPC, the VGW is the AWS-side endpoint of that connection — the thing your traffic arrives at once it crosses into AWS. This matters today because a Private VIF, one of the three Direct Connect interface types, terminates on a VGW, and understanding that a VGW is per-VPC is what makes the "one VPC versus many VPCs" decision in today's material make sense.

CIDR blocks and IP prefixes

Every network needs a way to describe "which addresses belong to me." That description is a CIDR block — a compact notation like 10.0.0.0/16 that means "all addresses starting with 10.0.0." The number after the slash says how many of the leading bits are fixed, so a smaller number means a bigger range. When this page says AWS "advertises the prefixes it wants you to reach," it means AWS is telling your router "send me anything destined for these address ranges." Routing, at its core, is just the process of matching a destination address against a list of prefixes and picking the best match. You do not need to do subnet math today, but you do need to recognize that a CIDR block is how networks name themselves, and that two networks can only talk if each one knows the other's prefix.

BGP: the protocol that decides which path traffic takes

When two networks connect, they need to agree on how to exchange routes — the lists of prefixes each side can reach. BGP (Border Gateway Protocol) is the standard protocol for that job, and it is what runs over a Direct Connect circuit once the physical link is up. The two routers on either end of the connection form a "BGP session," which is just an ongoing conversation where each side announces its prefixes and listens for the other's. The reason BGP matters so much today is that it also carries preference information: each route comes with attributes that say how attractive it is, and the router picks the most attractive one. That is how you make one of two Direct Connect circuits the preferred path and the other the standby, and it is why the resilience discussion later on this page keeps coming back to BGP attributes rather than to hardware.

ASNs: the identity numbers of networks

For BGP to work, each network needs a name. That name is an Autonomous System Number (ASN) — a number that uniquely identifies a network on the internet, the way a phone number identifies a line. AWS has its own ASN, and your organization has one too (either a public ASN, which is globally registered, or a private ASN, which is only meaningful between you and your peers). When this page talks about "AS path prepending," it means deliberately repeating your ASN in the route announcement so the path looks longer and therefore less attractive to the other side. That trick only makes sense once you know that an ASN is how networks identify themselves in BGP, which is why it is worth having straight before you read the resilience section.

VLANs and virtual interfaces (VIFs)

A single physical cable can carry traffic for several logically separate networks at once, and the mechanism that keeps them separate is a VLAN (Virtual Local Area Network). Each VLAN is tagged with a number, and the equipment on both ends uses that tag to sort incoming traffic into the right logical network. On a Direct Connect connection, each virtual interface (VIF) is exactly this: a VLAN on the shared physical link, with its own BGP session and its own set of prefixes. That is how one physical circuit can simultaneously serve a connection to a VPC, a connection to a Transit Gateway, and a connection to AWS public services without the three interfering with each other. The VIF is the unit you actually configure, and the physical connection is just the pipe underneath it.

Transit Gateway, in one paragraph

A Transit Gateway (TGW) is a central router inside AWS that many VPCs and on-premises networks can attach to, so that instead of every network needing a direct connection to every other network, they all connect to one hub. It solves the problem of connection sprawl: with ten VPCs, connecting each pair directly would require dozens of links, but with a TGW each VPC needs only one attachment. Today's material assumes you have a TGW in the picture, because the Transit VIF — the second of the three Direct Connect interface types — exists specifically to connect a Direct Connect circuit to a TGW rather than to a single VPC.

With that grounding, here's why Direct Connect exists and what problem it actually solves.

1. Why Direct Connect Is on the Exam

Direct Connect shows up on SAP-C02 for one reason: it is the only AWS connectivity option that removes the public internet from the path entirely, and almost every hybrid scenario in the exam is really a question about whether the public internet is acceptable. When a scenario says "regulatory requirement that traffic never traverse the public internet," or "consistent latency and jitter for a latency-sensitive trading system," or "predictable bandwidth for nightly bulk transfers that saturate a VPN," the answer is Direct Connect. When a scenario says "fastest to deploy," "lowest cost for occasional use," or "we need connectivity this week," the answer is a Site-to-Site VPN, and the exam is testing whether you can tell those two apart.

The service maps most directly to Domain 1 (Design Solutions for Organizational Complexity) in the SAP-C02 blueprint, specifically the hybrid connectivity and multi-account networking objectives. It also bleeds into Domain 2 (Design for New Solutions) whenever a scenario asks you to design a network that spans on-premises and multiple AWS accounts, and into Domain 4 (Accelerate Workload Migration and Modernization) because large migrations frequently depend on DX capacity to move data. The exam rarely asks "what is a VIF" in isolation. It asks you to choose between Private VIF, Transit VIF, and Public VIF given a described topology, and it asks you to evaluate whether a proposed DX design is actually resilient or merely redundant-looking.

The trap that catches most candidates is treating Direct Connect as a single product with a single behavior. It is really three products sharing a physical circuit: a layer-2 private connection to a VPC, a layer-3 connection to a Transit Gateway, and a path to AWS public endpoints. Each has different routing behavior, different limits, and different failure modes. A design that is correct for a Private VIF can be wrong for a Transit VIF, and the exam knows it.

2. How Direct Connect Actually Works

A Direct Connect connection is a physical cross-connect between a customer-owned router in a colocation facility and an AWS-owned router in the same facility. That physical layer is the part people forget: DX is not a service you configure in the console and forget. You either provision a dedicated connection (a port you lease from AWS, available in 1 Gbps, 10 Gbps, and 100 Gbps capacities) or you buy a hosted connection through an AWS Direct Connect Partner, who resells capacity on their own port in sub-1 Gbps increments. In both cases, the customer is responsible for the router on their side of the cross-connect, and AWS is responsible for the router on theirs.

Once the physical link is up, the two routers establish a BGP session. This is the control plane that makes DX interesting. AWS advertises the prefixes it wants you to reach — the VPC CIDRs attached to your Private VIF, the TGW-attached routes for a Transit VIF, or the AWS public service prefixes for a Public VIF — and you advertise the on-premises prefixes you want AWS to reach. The BGP session is where path selection happens, which is why every resilience question about DX eventually becomes a BGP question. If you have two circuits and both advertise the same prefixes with the same AS path length, you get equal-cost multipath and traffic splits across both. If you want active/passive, you manipulate the BGP attributes so one path is preferred.

The VIF is the logical interface that sits on top of the physical connection and the BGP session. A single dedicated connection can host multiple VIFs, and each VIF is a separate VLAN with its own BGP session and its own set of advertised prefixes. This is the mechanism that lets one physical circuit serve a Private VIF to a VPC, a Transit VIF to a TGW, and a Public VIF to S3 simultaneously. The VLAN ID is how AWS demultiplexes the traffic on the shared physical link. From the customer router's perspective, each VIF looks like a separate point-to-point link with its own neighbor IP.

The last piece is the Direct Connect Gateway, which is the construct that lets a single VIF reach resources in multiple AWS accounts and multiple regions. Without a DX Gateway, a Private VIF can only attach to a VPC in the same region and the same account as the connection. With a DX Gateway, you associate the gateway with VPCs in other accounts via a VPC association, and you associate the gateway with a Transit Gateway via a Transit Gateway association. The DX Gateway itself is a global resource — it is not regional — which is why it can bridge a us-east-1 connection to a us-west-2 VPC. This is the mechanism that makes DX viable for a multi-account landing zone, and it is the piece most candidates under-specify on the exam.

3. The Core Decision Boundary: Which VIF Type

Every Direct Connect scenario question reduces to a single fork: what does the customer need to reach, and how many things do they need to reach? If the answer is "one VPC in one account," a Private VIF is sufficient and simplest. If the answer is "many VPCs across many accounts, or a Transit Gateway that already aggregates them," a Transit VIF is the right choice. If the answer is "AWS public service endpoints like S3, DynamoDB, or the AWS APIs themselves, without going over the internet," a Public VIF is the answer. The exam will describe a topology and expect you to pick the VIF type that matches, and it will often include a distractor that is technically possible but operationally wrong.

The subtlety is that these are not mutually exclusive. A single dedicated connection can carry all three VIF types at once, and a mature enterprise design often does exactly that: a Transit VIF for VPC-to-VPC and VPC-to-on-prem traffic, a Public VIF for S3 and other public endpoints, and occasionally a Private VIF for a legacy VPC that predates the TGW. The exam question is usually not "which one" but "which combination," and the answer depends on whether the customer has already standardized on a Transit Gateway.

VIF typeReachesRequiresTypical use
Private VIFA single VPC (or multiple VPCs via a DX Gateway)Virtual private gateway or DX Gateway on the VPC sideSingle-VPC hybrid access, legacy designs
Transit VIFA Transit Gateway, and through it many VPCs and on-prem networksDX Gateway associated with a TGWHub-and-spoke hybrid, multi-account landing zones
Public VIFAWS public service endpoints (S3, DynamoDB, AWS APIs)Public ASN and public IP space on the customer sidePrivate access to S3 without internet or NAT

The decision boundary that matters most on the exam is Private VIF versus Transit VIF. A Private VIF attaches to a virtual private gateway (VGW) on a single VPC, or to a DX Gateway that can associate with up to 20 VPCs. A Transit VIF attaches to a DX Gateway that is associated with a Transit Gateway, and the TGW can then route to any number of VPCs and to other attachments. If the scenario mentions a Transit Gateway anywhere in the existing architecture, the answer is almost always a Transit VIF, because adding a Private VIF alongside an existing TGW creates a parallel routing path that is harder to reason about and harder to segment.

4. Configuration Modes and Their Tradeoffs

The first configuration decision is dedicated versus hosted. A dedicated connection gives you a port that only you use, with a fixed capacity of 1, 10, or 100 Gbps, and it takes weeks to provision because it involves a physical cross-connect at the DX location. A hosted connection is provisioned through a partner and can be as small as 50 Mbps, which makes it the right choice for a branch office or a proof of concept. The tradeoff is control: with a hosted connection you are sharing the partner's port, and your ability to troubleshoot at layer 1 is limited because you do not own the router on the far side of the cross-connect.

The second decision is the resiliency model, and this is where the exam spends most of its time. AWS publishes a set of resiliency models that range from a single connection (no redundancy, not recommended for production) through dual connections at the same location (redundant hardware but a shared facility, so a location-level failure takes both down) to dual connections at two different locations (the recommended production baseline). The highest tier adds a VPN backup over the internet, so that if both DX paths fail simultaneously — which is rare but has happened — traffic can still reach AWS. The exam will describe a workload's criticality and expect you to match it to the right tier.

The third decision is how to make one path preferred over the other when both are up. BGP gives you two standard tools. AS path prepending makes a path look longer by repeating your ASN in the AS path, which causes the remote router to prefer the other path. Local preference is a higher-priority attribute that is exchanged within an AS, so it is the right tool when you control both ends of the decision. In practice, most designs use AS path prepending on the standby path because it is simpler to reason about and does not require coordination with AWS. The exam will sometimes describe a design where both paths are active and traffic is splitting, and ask how to make one preferred — the answer is one of these two attributes.

The fourth decision is whether to use a DX Gateway. If the customer has a single VPC and a single account, a VGW is sufficient and a DX Gateway adds unnecessary complexity. If the customer has multiple accounts or multiple regions, the DX Gateway is what makes the design scale, because it decouples the connection from the VPC and lets you associate new VPCs without touching the physical circuit. The exam will often describe a multi-account landing zone and expect you to name the DX Gateway as the mechanism that lets one connection serve all of them.

5. Sizing, Limits, and Quotas

The numbers that matter for DX are the port capacities, the VIF limits, and the DX Gateway association limits. Dedicated connection capacities are 1 Gbps, 10 Gbps, and 100 Gbps. Hosted connection capacities range from 50 Mbps up to 10 Gbps depending on the partner. A single dedicated connection can host up to 50 VIFs, though in practice most designs use far fewer. A DX Gateway can be associated with up to 10 Transit Gateways and up to 20 virtual private gateways, and a Transit Gateway can be associated with up to 20 DX Gateways. These numbers are the ones the exam is most likely to test, because they determine whether a proposed design actually scales.

Link Aggregation Groups (LAGs) are the mechanism for bundling multiple connections into a single logical link. A LAG can contain up to four connections, and all connections in a LAG must be the same speed and terminate at the same DX location. The LAG presents a single BGP session to the customer, which simplifies routing but also means that a LAG is not a resiliency mechanism — it is a capacity mechanism. If the DX location fails, every connection in the LAG fails together. This is a common exam trap: a scenario describes a LAG and asks whether it provides redundancy, and the correct answer is no, because all members share a location.

The MTU on a Direct Connect connection is 1500 bytes by default, and jumbo frames (9001 bytes) are supported on dedicated connections and on some hosted connections. This matters for workloads that transfer large files or that use protocols sensitive to fragmentation. The exam occasionally includes an MTU detail in a scenario about performance, and the correct answer is usually to enable jumbo frames on both ends rather than to change the instance type.

LimitValueWhy it matters
Dedicated port speeds1, 10, 100 GbpsDetermines maximum throughput per circuit
Hosted port speeds50 Mbps – 10 GbpsSub-1 Gbps options for branches and POCs
VIFs per connection50How many logical interfaces one circuit can carry
LAG members4Capacity aggregation, not redundancy
DX Gateway → TGW associations10How many TGWs one gateway can serve
DX Gateway → VGW associations20How many VPCs one gateway can serve directly
Default MTU1500 bytesJumbo frames (9001) available on dedicated

6. Failure Modes and What They Look Like in Production

The most common DX failure is a BGP session flap, and it usually presents as intermittent packet loss rather than a clean outage. The symptom is that some traffic reaches AWS and some does not, because the BGP session is oscillating and the routing table is changing underneath the application. The first diagnostic move is to check the BGP session state on the customer router and on the AWS side via the Direct Connect console, then look at the CloudWatch metrics for the connection. If the session is flapping, the cause is usually a physical layer problem (a failing optic, a dirty fiber connector) or a BGP timer mismatch, not a configuration error.

The second failure mode is a single-connection outage, which is what happens when a customer has only one DX circuit and that circuit goes down. The symptom is a total loss of hybrid connectivity, and the recovery time depends entirely on whether a VPN backup exists. Without a backup, the customer is waiting for the physical circuit to be repaired, which can take hours or days. With a backup, the VPN takes over within minutes, but the bandwidth and latency characteristics change dramatically, which can break applications that were tuned for the DX path. This is why the exam emphasizes that a VPN backup is not a substitute for a second DX connection — it is a last-resort fallback.

The third failure mode is a DX location outage, which takes down every connection terminating at that location. This is the failure that a LAG does not protect against, and it is the reason the recommended production design uses two connections at two different locations. The symptom is that both circuits go down simultaneously, which is confusing if the operator assumed the two circuits were independent. The first diagnostic move is to check whether both connections share a location, which is visible in the Direct Connect console and in the connection ARNs.

The fourth failure mode is a routing misconfiguration, which is the most insidious because the physical layer is healthy and the BGP session is up. The symptom is that traffic reaches AWS but takes the wrong path — for example, on-prem traffic to a VPC goes over the VPN instead of the DX circuit because the BGP attributes are wrong. The first diagnostic move is to look at the effective routes on the customer router and on the AWS side, and to compare the AS path lengths. This is where AS path prepending mistakes show up: a prepend applied to the wrong path makes the backup path preferred, which is the opposite of the intent.

7. The Operational and SRE Angle

Direct Connect is one of the few AWS services where the customer owns part of the physical infrastructure, which means the operational model is different from a fully managed service. The customer is responsible for the router, the cross-connect, and the BGP configuration on their side. AWS is responsible for the port on their side and for the DX location. This split means that a DX incident often requires a support case with both AWS and the colocation provider, and the runbook needs to reflect that. The first step in any DX runbook is to determine which side of the cross-connect the problem is on, because that determines who gets paged.

The metrics that matter for DX monitoring are the connection state, the BGP session state, the light levels on the optic, and the throughput on each VIF. CloudWatch exposes connection-level metrics, and the customer router exposes the BGP and interface metrics. The SLO implication is that DX availability is a shared responsibility: AWS publishes a service level agreement for the connection itself, but the end-to-end availability of the hybrid path depends on the customer's router and the colocation facility. A realistic SLO for a hybrid path is therefore lower than the SLO for a purely in-AWS path, and the exam sometimes tests whether you understand that distinction.

The runbook shape for a DX incident is: confirm the scope (one VIF, one connection, or one location), check the BGP session state on both sides, check the physical layer metrics, and then decide whether to fail over to the backup path. The failover decision is usually automated via BGP, but the runbook should still document the manual steps in case the automation fails. The post-incident review should capture whether the failure was within the expected failure domain — a single connection failing is expected and should be transparent; a location failing is expected and should be covered by the second location; a BGP misconfiguration is not expected and should drive a configuration change.

8. Edge Cases and Exam Gotchas

The first gotcha is that a LAG is not redundant. The exam will describe a customer with four connections in a LAG and ask whether the design is resilient, and the correct answer is no, because all four connections terminate at the same DX location. The second gotcha is that a VPN backup is not a substitute for a second DX connection. The exam will describe a customer with one DX connection and a VPN backup and ask whether the design meets a high-availability requirement, and the correct answer depends on the stated RTO — a VPN backup is acceptable for a moderate RTO but not for a mission-critical one.

The third gotcha is that a Private VIF cannot reach a VPC in another region without a DX Gateway. The exam will describe a customer with a DX connection in us-east-1 and a VPC in us-west-2 and ask how to connect them, and the correct answer is a DX Gateway, not a second connection. The fourth gotcha is that a Public VIF requires public IP space and a public ASN on the customer side, which many customers do not have. The exam will describe a customer that wants private access to S3 and ask what they need, and the correct answer includes the public IP and ASN requirement, not just the VIF.

The fifth gotcha is that the DX Gateway is a global resource, not a regional one. The exam will describe a customer with connections in two regions and ask how many DX Gateways they need, and the correct answer is one, because a single DX Gateway can be associated with connections in multiple regions. The sixth gotcha is that a Transit VIF requires a DX Gateway associated with a Transit Gateway, and the TGW must have the appropriate route tables configured. The exam will describe a customer with a Transit VIF and a TGW and ask why traffic is not reaching a specific VPC, and the correct answer is usually a missing TGW route table association or propagation.

The seventh gotcha is that the BGP ASN on the customer side must be a public ASN if the customer wants to use a Public VIF, but can be a private ASN for Private and Transit VIFs. The exam will describe a customer with a private ASN and ask whether they can use a Public VIF, and the correct answer is no. The eighth gotcha is that the DX connection itself does not encrypt traffic. The exam will describe a customer with a regulatory requirement for encryption in transit and ask whether DX satisfies it, and the correct answer is no — DX is private but not encrypted, and the customer must add MACsec or a VPN on top of DX if encryption is required.

9. Direct Connect vs. the Alternatives

The comparison the exam cares about most is Direct Connect versus Site-to-Site VPN. Both provide hybrid connectivity, but they differ on every axis that matters: DX is private, consistent, and expensive with a long lead time; VPN is public, variable, cheap, and fast to deploy. The exam will describe a scenario and expect you to pick the one that matches the constraints, and the most common correct answer is a combination — DX for the primary path and VPN for the backup. The second comparison is DX versus Transit Gateway, which is not really a comparison because they solve different problems: DX is the physical path into AWS, and TGW is the routing fabric inside AWS. A mature design uses both, with a Transit VIF connecting the DX circuit to the TGW.

The third comparison is DX versus PrivateLink, which is also not a direct comparison but shows up in exam scenarios about cross-account access. DX connects a customer network to AWS; PrivateLink exposes a specific service to a specific consumer. If the scenario is about a customer's data center reaching AWS, the answer is DX. If the scenario is about one AWS account consuming a service in another AWS account, the answer is PrivateLink. The fourth comparison is DX versus AWS VPN CloudHub, which is a hub-and-spoke VPN topology for connecting multiple branch offices. CloudHub is the right answer when the customer has many small sites and no dedicated circuits; DX is the right answer when the customer has a small number of large sites and needs consistent performance.

OptionPathLead timePick it when…
Direct ConnectPrivate, dedicatedWeeksConsistent latency, high bandwidth, regulatory privacy
Site-to-Site VPNPublic internet, encryptedMinutesFast deployment, low cost, backup path
Transit GatewayInside AWSMinutesRouting between many VPCs and on-prem
PrivateLinkInside AWS, service-levelMinutesExposing one service to one consumer, CIDR overlap
VPN CloudHubPublic internet, hub-and-spokeMinutesMany small branch sites, no dedicated circuits

The rule of thumb to carry into the exam: if the scenario mentions a physical data center, a regulatory requirement, or a latency-sensitive workload, start with Direct Connect. If it mentions a deadline measured in days, a small branch office, or a proof of concept, start with VPN. If it mentions many VPCs or many accounts, add a Transit Gateway and use a Transit VIF. If it mentions a single service being consumed across accounts, use PrivateLink instead. The exam rarely asks for a single service in isolation; it asks for the combination that satisfies all the stated constraints, and the combination is usually DX plus TGW plus a VPN backup.

Hands-on Lab: Designing a Redundant DX Architecture (45 min)

This lab walks through the design of a redundant Direct Connect architecture for a fictional enterprise, using two connections at two different DX locations with active/passive BGP path preference. The goal is not to provision real hardware — that would take weeks — but to produce the design artifacts that a real DX deployment requires: the topology diagram, the BGP configuration, the failover test plan, and the runbook. Work through the steps in order and document each decision.

Step 1 — Define the requirements. Assume the enterprise has a primary data center in Ashburn, Virginia, and a secondary data center in Columbus, Ohio. The workload is a financial trading application that requires consistent sub-10ms latency to us-east-1 and cannot tolerate more than 5 minutes of downtime. The compliance team requires that traffic never traverse the public internet. Write down the RTO (5 minutes), the RPO (near-zero, since the application is stateless), and the compliance constraint. These three numbers drive every subsequent decision.

Step 2 — Choose the DX locations. AWS publishes a list of Direct Connect locations, and the two connections must terminate at different locations to survive a location-level failure. For this design, choose two DX locations in the us-east-1 region that are geographically separate — for example, one in Ashburn and one in a different metro area within the region's DX footprint. Document why the two locations are independent: they have separate power, separate cooling, and separate network paths to the AWS backbone. If the two locations share a single fiber conduit, the design is not actually redundant, and the exam will test whether you noticed.

Step 3 — Choose the port speeds. The trading application needs consistent bandwidth, so choose dedicated 10 Gbps connections rather than hosted connections. Document the capacity calculation: peak throughput, headroom for burst, and the cost difference between 1 Gbps and 10 Gbps. If the application's peak is 2 Gbps, a 10 Gbps port provides 5x headroom, which is reasonable for a latency-sensitive workload. A 1 Gbps port would be cheaper but would leave no room for growth.

Step 4 — Design the VIF topology. The enterprise has a Transit Gateway that already aggregates its VPCs, so the correct choice is a Transit VIF on each connection, both associated with a single DX Gateway. Document the DX Gateway associations: one DX Gateway, associated with the Transit Gateway, which is associated with the VPCs. If the enterprise also needs private access to S3, add a Public VIF on one of the connections and document the public IP and ASN requirements.

Step 5 — Configure BGP for active/passive. Both connections will advertise the same on-premises prefixes, so without intervention BGP will treat them as equal-cost and split traffic. To make one path preferred, apply AS path prepending on the standby connection: prepend the customer ASN three times so the AS path looks longer and the primary path is preferred. Document the exact BGP configuration on the customer router, including the neighbor IPs, the ASNs, and the prepend policy. Explain why AS path prepending was chosen over local preference: it is simpler and does not require coordination with AWS.

Step 6 — Add a VPN backup. Even with two DX connections, a simultaneous failure of both locations is possible, so add a Site-to-Site VPN over the internet as a last-resort backup. Document the VPN's BGP configuration and the failover behavior: the VPN should only be preferred when both DX paths are down, which means its AS path must be longer than the DX paths even after prepending. This is a subtle configuration detail that the exam sometimes tests.

Step 7 — Write the failover test plan. A redundant design is only as good as its tested failover. Write a test plan that covers three scenarios: a single connection failure (shut down one BGP session and verify traffic shifts to the other connection within the RTO), a single location failure (simulate a location outage and verify the same), and a simultaneous failure (shut down both DX paths and verify the VPN takes over). For each scenario, document the expected behavior, the actual behavior, and the time to recover.

Step 8 — Write the runbook. The runbook should cover the first three diagnostic steps for a DX incident: confirm the scope, check the BGP session state on both sides, and check the physical layer metrics. It should also document the escalation path: which issues go to AWS Support, which go to the colocation provider, and which are internal. Finally, it should document the manual failover procedure in case the BGP automation fails.

Step 9 — Review against the exam. Go back through the design and identify every decision that maps to an exam concept: the VIF type, the DX Gateway, the BGP attributes, the resiliency tier, and the VPN backup. For each, write one sentence explaining why the alternative would be wrong. This is the step that turns a lab exercise into exam preparation.

Scenario Question Drills (20 min)

Q1. A company needs the highest resiliency Direct Connect design for a mission-critical workload. What should they provision?

A. A single 10 Gbps DX connection
B. Two DX connections at two different DX locations, each terminating on separate devices, with BGP failover and a VPN backup
C. One DX connection plus a second identical connection at the same location
D. Public VIF only, since it's cheaper
Correct answer: B. AWS's resiliency model for DX requires diversity across DX locations and devices, plus BGP-based failover; a VPN backup adds an additional layer for the rare case both DX paths fail.

Q2. A company has a single VPC in a single AWS account and wants private connectivity from its data center. Which VIF type is the simplest correct choice?

A. Transit VIF
B. Private VIF attached to a virtual private gateway
C. Public VIF
D. A DX Gateway is always required
Correct answer: B. A Private VIF attached to a VGW is the simplest design for a single VPC in a single account. A DX Gateway is only needed when the connection must reach multiple VPCs, accounts, or regions.

Q3. A company has four DX connections bundled into a Link Aggregation Group at a single DX location. Is this design resilient to a location-level failure?

A. Yes, because four connections provide 4x redundancy
B. No, because all four connections terminate at the same DX location and fail together
C. Yes, because a LAG automatically spans multiple locations
D. Only if the LAG uses BGP
Correct answer: B. A LAG is a capacity mechanism, not a resiliency mechanism. All members must terminate at the same DX location, so a location-level failure takes down the entire LAG.

Q4. A company has a DX connection in us-east-1 and needs to reach a VPC in us-west-2 over the same connection. What is required?

A. A second DX connection in us-west-2
B. A Direct Connect Gateway associated with the us-west-2 VPC
C. A VPC peering connection between the two regions
D. Nothing — a Private VIF can reach any region
Correct answer: B. A DX Gateway is a global resource that can associate with VPCs in other regions, letting a single connection reach them. A Private VIF alone is limited to the connection's region.

Q5. A company wants private access to Amazon S3 from its data center without traversing the public internet. Which VIF type is required?

A. Private VIF
B. Transit VIF
C. Public VIF
D. A VPC endpoint is the only option
Correct answer: C. A Public VIF provides private connectivity to AWS public service endpoints such as S3 and DynamoDB. It requires public IP space and a public ASN on the customer side.

Q6. A company has two DX connections at two different locations, both advertising the same prefixes. Traffic is splitting across both paths, but the company wants one path preferred. What should they configure?

A. Increase the MTU on the preferred path
B. Apply AS path prepending on the standby path
C. Disable BGP on the standby path
D. Use a LAG to combine the two connections
Correct answer: B. AS path prepending makes the standby path look longer, so BGP prefers the primary path. Local preference is an alternative but requires coordination within the AS.

Q7. A company has a DX connection and a Site-to-Site VPN backup. The DX connection fails. What is the most likely operational impact?

A. No impact — the VPN is a full substitute
B. Traffic fails over to the VPN, but bandwidth and latency characteristics change and may break applications tuned for DX
C. The VPN cannot be used as a backup for DX
D. Traffic fails over only if the VPN uses a Public VIF
Correct answer: B. A VPN backup restores connectivity but with different bandwidth and latency, which can break applications that were tuned for the DX path. It is a last-resort fallback, not a substitute for a second DX connection.

Q8. A company has a Transit Gateway aggregating its VPCs and wants to connect its data center over Direct Connect. Which VIF type should they use?

A. Private VIF attached to a VGW
B. Transit VIF associated with a DX Gateway that is associated with the TGW
C. Public VIF
D. Two Private VIFs, one per VPC
Correct answer: B. A Transit VIF connects the DX circuit to a Transit Gateway via a DX Gateway, letting the TGW route to all attached VPCs. Adding Private VIFs alongside an existing TGW creates a parallel path that is harder to segment.

Q9. A company has a private ASN and wants to use a Public VIF to reach S3. What is the problem?

A. Public VIFs do not support S3
B. A Public VIF requires a public ASN and public IP space on the customer side
C. Public VIFs require a DX Gateway
D. There is no problem — private ASNs work with Public VIFs
Correct answer: B. A Public VIF requires the customer to use a public ASN and public IP space, because it peers with AWS's public routing domain. Private ASNs are only valid for Private and Transit VIFs.

Q10. A company has a regulatory requirement that all traffic in transit must be encrypted. Does a Direct Connect connection satisfy this requirement on its own?

A. Yes, because DX is a private connection
B. No, because DX is private but not encrypted; the customer must add MACsec or a VPN on top
C. Yes, because DX uses TLS by default
D. Only if the connection uses a Public VIF
Correct answer: B. Direct Connect provides a private path but does not encrypt traffic. Customers with encryption requirements must add MACsec (on supported connections) or run a VPN over the DX circuit.

Q11. A company has DX connections in two different AWS regions and wants to reach VPCs in both. How many DX Gateways do they need?

A. One per region
B. One per connection
C. One, because a DX Gateway is a global resource
D. DX Gateways cannot span regions
Correct answer: C. A DX Gateway is a global resource that can be associated with connections in multiple regions, so a single gateway can serve both.

Q12. A company's DX connection is up and the BGP session is established, but on-prem traffic to a VPC is going over the VPN instead of the DX circuit. What is the most likely cause?

A. The DX connection is down
B. A BGP attribute misconfiguration, such as AS path prepending applied to the wrong path
C. The VPC has no route table
D. The VPN is always preferred over DX
Correct answer: B. If the physical layer is healthy and BGP is up, the problem is routing. AS path prepending applied to the wrong path makes the backup path preferred, which is the opposite of the intent.

Q13. A company needs to move 300 TB of historical data to AWS and has a 500 Mbps DX connection. What is the recommended approach?

A. Transfer everything over the DX connection
B. Use Snowball Edge for the bulk historical transfer, then use the DX connection for ongoing deltas
C. Upgrade the DX connection to 10 Gbps first
D. Use a VPN instead of DX
Correct answer: B. The hybrid pattern of offline bulk transfer for the large historical dataset plus network-based incremental sync for the small ongoing delta is the standard approach when the network link is the bottleneck.

Q14. A company wants to connect 40 small branch offices to AWS with minimal cost and no dedicated circuits. What is the best fit?

A. A Direct Connect connection per branch
B. AWS VPN CloudHub
C. A single Transit VIF shared by all branches
D. PrivateLink to each branch
Correct answer: B. VPN CloudHub is a hub-and-spoke VPN topology designed for many small sites without dedicated circuits. DX per branch would be prohibitively expensive and slow to deploy.

Q15. A company has a Transit VIF and a Transit Gateway, but traffic from on-prem is not reaching a specific VPC. What is the most likely cause?

A. The DX connection is down
B. A missing TGW route table association or propagation for that VPC
C. The VIF type is wrong
D. The VPC needs a Public VIF
Correct answer: B. A Transit VIF depends on the TGW's route tables to reach VPCs. If the VPC's attachment is not associated with the correct route table or its routes are not propagated, traffic will not reach it.

Peek into Tomorrow: Hybrid DNS

Today's design gets packets from the data center to AWS over a private circuit, but it leaves a question unanswered: once the packets arrive, how do the two sides resolve each other's names? A server in the data center that needs to reach a database in a private Route 53 hosted zone has to resolve that name somehow, and the default behavior — querying the public DNS — will not work, because the hosted zone is private. Conversely, an EC2 instance in a VPC that needs to reach an on-premises service by its internal hostname has to resolve a name that AWS's default resolver has never heard of. The DX circuit carries the packets, but it does not carry the DNS answers.

Tomorrow's topic, Route 53 Resolver, is the piece that closes this gap. It introduces Inbound endpoints, which let on-premises resolvers query AWS private hosted zones, and Outbound endpoints with forwarding rules, which let AWS resources resolve on-premises names over the same DX or VPN path we designed today. The open question is how to configure the forwarding rules so that only the queries that need to go on-prem actually go on-prem, and how to make the resolver endpoints resilient across Availability Zones. The answers depend on the same redundancy thinking we applied to the DX connections, but at the DNS layer instead of the packet layer.

Sources