Direct Connect (DX) & DX Gateway
Recap: From Inspection VPC to Private Wire
Day 10 put a centralized inspection VPC in the middle of the network, with AWS Network Firewall applying Suricata-compatible rules to traffic that the Transit Gateway steered into it. That design assumed the traffic was already inside AWS, or arriving over the public internet through a VPN. The appliance VPC pattern and its egress routing via TGW solved the question of where inspection happens, but it left the question of how packets get into AWS in the first place largely untouched.
Today extends that architecture one layer outward. Direct Connect replaces the internet path with a dedicated private circuit from a customer router into an AWS Direct Connect location, and the VIF types determine whether that circuit lands on a single VPC, on a Transit Gateway, or on AWS public service endpoints. The inspection VPC from yesterday does not go away — it becomes the enforcement point for traffic that arrives over the new private path. The same TGW route tables that steered traffic into the firewall now also carry the DX attachment, which means the segmentation decisions from Day 8 and the inspection decisions from Day 10 both apply to the circuit we are about to design.
Foundations You'll Need Today
Today's topic sits on top of several networking ideas that the rest of this curriculum assumes you already have. If you have never built a VPC by hand or configured a router, the terms below are the ones that will otherwise make the rest of this page read like a foreign language. None of them are complicated once the underlying problem is clear, so let's build them up from what each one is actually for.
VPCs and virtual private gateways
A VPC (Virtual Private Cloud) is your own isolated slice of the AWS network — a private address space where you launch servers, databases, and load balancers, and where nothing from the outside can reach them unless you explicitly allow it. Think of it as a private data center floor that AWS has carved out for you, with its own internal IP addressing scheme. A virtual private gateway (VGW) is the component that sits on the edge of a VPC and acts as the doorway to the outside world for private traffic. When you connect your on-premises network to a VPC, the VGW is the AWS-side endpoint of that connection — the thing your traffic arrives at once it crosses into AWS. This matters today because a Private VIF, one of the three Direct Connect interface types, terminates on a VGW, and understanding that a VGW is per-VPC is what makes the "one VPC versus many VPCs" decision in today's material make sense.
CIDR blocks and IP prefixes
Every network needs a way to describe "which addresses belong to me." That description is a CIDR block — a compact notation like 10.0.0.0/16 that means "all addresses starting with 10.0.0." The number after the slash says how many of the leading bits are fixed, so a smaller number means a bigger range. When this page says AWS "advertises the prefixes it wants you to reach," it means AWS is telling your router "send me anything destined for these address ranges." Routing, at its core, is just the process of matching a destination address against a list of prefixes and picking the best match. You do not need to do subnet math today, but you do need to recognize that a CIDR block is how networks name themselves, and that two networks can only talk if each one knows the other's prefix.
BGP: the protocol that decides which path traffic takes
When two networks connect, they need to agree on how to exchange routes — the lists of prefixes each side can reach. BGP (Border Gateway Protocol) is the standard protocol for that job, and it is what runs over a Direct Connect circuit once the physical link is up. The two routers on either end of the connection form a "BGP session," which is just an ongoing conversation where each side announces its prefixes and listens for the other's. The reason BGP matters so much today is that it also carries preference information: each route comes with attributes that say how attractive it is, and the router picks the most attractive one. That is how you make one of two Direct Connect circuits the preferred path and the other the standby, and it is why the resilience discussion later on this page keeps coming back to BGP attributes rather than to hardware.
ASNs: the identity numbers of networks
For BGP to work, each network needs a name. That name is an Autonomous System Number (ASN) — a number that uniquely identifies a network on the internet, the way a phone number identifies a line. AWS has its own ASN, and your organization has one too (either a public ASN, which is globally registered, or a private ASN, which is only meaningful between you and your peers). When this page talks about "AS path prepending," it means deliberately repeating your ASN in the route announcement so the path looks longer and therefore less attractive to the other side. That trick only makes sense once you know that an ASN is how networks identify themselves in BGP, which is why it is worth having straight before you read the resilience section.
VLANs and virtual interfaces (VIFs)
A single physical cable can carry traffic for several logically separate networks at once, and the mechanism that keeps them separate is a VLAN (Virtual Local Area Network). Each VLAN is tagged with a number, and the equipment on both ends uses that tag to sort incoming traffic into the right logical network. On a Direct Connect connection, each virtual interface (VIF) is exactly this: a VLAN on the shared physical link, with its own BGP session and its own set of prefixes. That is how one physical circuit can simultaneously serve a connection to a VPC, a connection to a Transit Gateway, and a connection to AWS public services without the three interfering with each other. The VIF is the unit you actually configure, and the physical connection is just the pipe underneath it.
Transit Gateway, in one paragraph
A Transit Gateway (TGW) is a central router inside AWS that many VPCs and on-premises networks can attach to, so that instead of every network needing a direct connection to every other network, they all connect to one hub. It solves the problem of connection sprawl: with ten VPCs, connecting each pair directly would require dozens of links, but with a TGW each VPC needs only one attachment. Today's material assumes you have a TGW in the picture, because the Transit VIF — the second of the three Direct Connect interface types — exists specifically to connect a Direct Connect circuit to a TGW rather than to a single VPC.
With that grounding, here's why Direct Connect exists and what problem it actually solves.
1. Why Direct Connect Is on the Exam
Direct Connect shows up on SAP-C02 for one reason: it is the only AWS connectivity option that removes the public internet from the path entirely, and almost every hybrid scenario in the exam is really a question about whether the public internet is acceptable. When a scenario says "regulatory requirement that traffic never traverse the public internet," or "consistent latency and jitter for a latency-sensitive trading system," or "predictable bandwidth for nightly bulk transfers that saturate a VPN," the answer is Direct Connect. When a scenario says "fastest to deploy," "lowest cost for occasional use," or "we need connectivity this week," the answer is a Site-to-Site VPN, and the exam is testing whether you can tell those two apart.
The service maps most directly to Domain 1 (Design Solutions for Organizational Complexity) in the SAP-C02 blueprint, specifically the hybrid connectivity and multi-account networking objectives. It also bleeds into Domain 2 (Design for New Solutions) whenever a scenario asks you to design a network that spans on-premises and multiple AWS accounts, and into Domain 4 (Accelerate Workload Migration and Modernization) because large migrations frequently depend on DX capacity to move data. The exam rarely asks "what is a VIF" in isolation. It asks you to choose between Private VIF, Transit VIF, and Public VIF given a described topology, and it asks you to evaluate whether a proposed DX design is actually resilient or merely redundant-looking.
The trap that catches most candidates is treating Direct Connect as a single product with a single behavior. It is really three products sharing a physical circuit: a layer-2 private connection to a VPC, a layer-3 connection to a Transit Gateway, and a path to AWS public endpoints. Each has different routing behavior, different limits, and different failure modes. A design that is correct for a Private VIF can be wrong for a Transit VIF, and the exam knows it.
2. How Direct Connect Actually Works
A Direct Connect connection is a physical cross-connect between a customer-owned router in a colocation facility and an AWS-owned router in the same facility. That physical layer is the part people forget: DX is not a service you configure in the console and forget. You either provision a dedicated connection (a port you lease from AWS, available in 1 Gbps, 10 Gbps, and 100 Gbps capacities) or you buy a hosted connection through an AWS Direct Connect Partner, who resells capacity on their own port in sub-1 Gbps increments. In both cases, the customer is responsible for the router on their side of the cross-connect, and AWS is responsible for the router on theirs.
Once the physical link is up, the two routers establish a BGP session. This is the control plane that makes DX interesting. AWS advertises the prefixes it wants you to reach — the VPC CIDRs attached to your Private VIF, the TGW-attached routes for a Transit VIF, or the AWS public service prefixes for a Public VIF — and you advertise the on-premises prefixes you want AWS to reach. The BGP session is where path selection happens, which is why every resilience question about DX eventually becomes a BGP question. If you have two circuits and both advertise the same prefixes with the same AS path length, you get equal-cost multipath and traffic splits across both. If you want active/passive, you manipulate the BGP attributes so one path is preferred.
The VIF is the logical interface that sits on top of the physical connection and the BGP session. A single dedicated connection can host multiple VIFs, and each VIF is a separate VLAN with its own BGP session and its own set of advertised prefixes. This is the mechanism that lets one physical circuit serve a Private VIF to a VPC, a Transit VIF to a TGW, and a Public VIF to S3 simultaneously. The VLAN ID is how AWS demultiplexes the traffic on the shared physical link. From the customer router's perspective, each VIF looks like a separate point-to-point link with its own neighbor IP.
The last piece is the Direct Connect Gateway, which is the construct that lets a single VIF reach resources in multiple AWS accounts and multiple regions. Without a DX Gateway, a Private VIF can only attach to a VPC in the same region and the same account as the connection. With a DX Gateway, you associate the gateway with VPCs in other accounts via a VPC association, and you associate the gateway with a Transit Gateway via a Transit Gateway association. The DX Gateway itself is a global resource — it is not regional — which is why it can bridge a us-east-1 connection to a us-west-2 VPC. This is the mechanism that makes DX viable for a multi-account landing zone, and it is the piece most candidates under-specify on the exam.
3. The Core Decision Boundary: Which VIF Type
Every Direct Connect scenario question reduces to a single fork: what does the customer need to reach, and how many things do they need to reach? If the answer is "one VPC in one account," a Private VIF is sufficient and simplest. If the answer is "many VPCs across many accounts, or a Transit Gateway that already aggregates them," a Transit VIF is the right choice. If the answer is "AWS public service endpoints like S3, DynamoDB, or the AWS APIs themselves, without going over the internet," a Public VIF is the answer. The exam will describe a topology and expect you to pick the VIF type that matches, and it will often include a distractor that is technically possible but operationally wrong.
The subtlety is that these are not mutually exclusive. A single dedicated connection can carry all three VIF types at once, and a mature enterprise design often does exactly that: a Transit VIF for VPC-to-VPC and VPC-to-on-prem traffic, a Public VIF for S3 and other public endpoints, and occasionally a Private VIF for a legacy VPC that predates the TGW. The exam question is usually not "which one" but "which combination," and the answer depends on whether the customer has already standardized on a Transit Gateway.
| VIF type | Reaches | Requires | Typical use |
|---|---|---|---|
| Private VIF | A single VPC (or multiple VPCs via a DX Gateway) | Virtual private gateway or DX Gateway on the VPC side | Single-VPC hybrid access, legacy designs |
| Transit VIF | A Transit Gateway, and through it many VPCs and on-prem networks | DX Gateway associated with a TGW | Hub-and-spoke hybrid, multi-account landing zones |
| Public VIF | AWS public service endpoints (S3, DynamoDB, AWS APIs) | Public ASN and public IP space on the customer side | Private access to S3 without internet or NAT |
The decision boundary that matters most on the exam is Private VIF versus Transit VIF. A Private VIF attaches to a virtual private gateway (VGW) on a single VPC, or to a DX Gateway that can associate with up to 20 VPCs. A Transit VIF attaches to a DX Gateway that is associated with a Transit Gateway, and the TGW can then route to any number of VPCs and to other attachments. If the scenario mentions a Transit Gateway anywhere in the existing architecture, the answer is almost always a Transit VIF, because adding a Private VIF alongside an existing TGW creates a parallel routing path that is harder to reason about and harder to segment.
4. Configuration Modes and Their Tradeoffs
The first configuration decision is dedicated versus hosted. A dedicated connection gives you a port that only you use, with a fixed capacity of 1, 10, or 100 Gbps, and it takes weeks to provision because it involves a physical cross-connect at the DX location. A hosted connection is provisioned through a partner and can be as small as 50 Mbps, which makes it the right choice for a branch office or a proof of concept. The tradeoff is control: with a hosted connection you are sharing the partner's port, and your ability to troubleshoot at layer 1 is limited because you do not own the router on the far side of the cross-connect.
The second decision is the resiliency model, and this is where the exam spends most of its time. AWS publishes a set of resiliency models that range from a single connection (no redundancy, not recommended for production) through dual connections at the same location (redundant hardware but a shared facility, so a location-level failure takes both down) to dual connections at two different locations (the recommended production baseline). The highest tier adds a VPN backup over the internet, so that if both DX paths fail simultaneously — which is rare but has happened — traffic can still reach AWS. The exam will describe a workload's criticality and expect you to match it to the right tier.
The third decision is how to make one path preferred over the other when both are up. BGP gives you two standard tools. AS path prepending makes a path look longer by repeating your ASN in the AS path, which causes the remote router to prefer the other path. Local preference is a higher-priority attribute that is exchanged within an AS, so it is the right tool when you control both ends of the decision. In practice, most designs use AS path prepending on the standby path because it is simpler to reason about and does not require coordination with AWS. The exam will sometimes describe a design where both paths are active and traffic is splitting, and ask how to make one preferred — the answer is one of these two attributes.
The fourth decision is whether to use a DX Gateway. If the customer has a single VPC and a single account, a VGW is sufficient and a DX Gateway adds unnecessary complexity. If the customer has multiple accounts or multiple regions, the DX Gateway is what makes the design scale, because it decouples the connection from the VPC and lets you associate new VPCs without touching the physical circuit. The exam will often describe a multi-account landing zone and expect you to name the DX Gateway as the mechanism that lets one connection serve all of them.
5. Sizing, Limits, and Quotas
The numbers that matter for DX are the port capacities, the VIF limits, and the DX Gateway association limits. Dedicated connection capacities are 1 Gbps, 10 Gbps, and 100 Gbps. Hosted connection capacities range from 50 Mbps up to 10 Gbps depending on the partner. A single dedicated connection can host up to 50 VIFs, though in practice most designs use far fewer. A DX Gateway can be associated with up to 10 Transit Gateways and up to 20 virtual private gateways, and a Transit Gateway can be associated with up to 20 DX Gateways. These numbers are the ones the exam is most likely to test, because they determine whether a proposed design actually scales.
Link Aggregation Groups (LAGs) are the mechanism for bundling multiple connections into a single logical link. A LAG can contain up to four connections, and all connections in a LAG must be the same speed and terminate at the same DX location. The LAG presents a single BGP session to the customer, which simplifies routing but also means that a LAG is not a resiliency mechanism — it is a capacity mechanism. If the DX location fails, every connection in the LAG fails together. This is a common exam trap: a scenario describes a LAG and asks whether it provides redundancy, and the correct answer is no, because all members share a location.
The MTU on a Direct Connect connection is 1500 bytes by default, and jumbo frames (9001 bytes) are supported on dedicated connections and on some hosted connections. This matters for workloads that transfer large files or that use protocols sensitive to fragmentation. The exam occasionally includes an MTU detail in a scenario about performance, and the correct answer is usually to enable jumbo frames on both ends rather than to change the instance type.
| Limit | Value | Why it matters |
|---|---|---|
| Dedicated port speeds | 1, 10, 100 Gbps | Determines maximum throughput per circuit |
| Hosted port speeds | 50 Mbps – 10 Gbps | Sub-1 Gbps options for branches and POCs |
| VIFs per connection | 50 | How many logical interfaces one circuit can carry |
| LAG members | 4 | Capacity aggregation, not redundancy |
| DX Gateway → TGW associations | 10 | How many TGWs one gateway can serve |
| DX Gateway → VGW associations | 20 | How many VPCs one gateway can serve directly |
| Default MTU | 1500 bytes | Jumbo frames (9001) available on dedicated |
6. Failure Modes and What They Look Like in Production
The most common DX failure is a BGP session flap, and it usually presents as intermittent packet loss rather than a clean outage. The symptom is that some traffic reaches AWS and some does not, because the BGP session is oscillating and the routing table is changing underneath the application. The first diagnostic move is to check the BGP session state on the customer router and on the AWS side via the Direct Connect console, then look at the CloudWatch metrics for the connection. If the session is flapping, the cause is usually a physical layer problem (a failing optic, a dirty fiber connector) or a BGP timer mismatch, not a configuration error.
The second failure mode is a single-connection outage, which is what happens when a customer has only one DX circuit and that circuit goes down. The symptom is a total loss of hybrid connectivity, and the recovery time depends entirely on whether a VPN backup exists. Without a backup, the customer is waiting for the physical circuit to be repaired, which can take hours or days. With a backup, the VPN takes over within minutes, but the bandwidth and latency characteristics change dramatically, which can break applications that were tuned for the DX path. This is why the exam emphasizes that a VPN backup is not a substitute for a second DX connection — it is a last-resort fallback.
The third failure mode is a DX location outage, which takes down every connection terminating at that location. This is the failure that a LAG does not protect against, and it is the reason the recommended production design uses two connections at two different locations. The symptom is that both circuits go down simultaneously, which is confusing if the operator assumed the two circuits were independent. The first diagnostic move is to check whether both connections share a location, which is visible in the Direct Connect console and in the connection ARNs.
The fourth failure mode is a routing misconfiguration, which is the most insidious because the physical layer is healthy and the BGP session is up. The symptom is that traffic reaches AWS but takes the wrong path — for example, on-prem traffic to a VPC goes over the VPN instead of the DX circuit because the BGP attributes are wrong. The first diagnostic move is to look at the effective routes on the customer router and on the AWS side, and to compare the AS path lengths. This is where AS path prepending mistakes show up: a prepend applied to the wrong path makes the backup path preferred, which is the opposite of the intent.
7. The Operational and SRE Angle
Direct Connect is one of the few AWS services where the customer owns part of the physical infrastructure, which means the operational model is different from a fully managed service. The customer is responsible for the router, the cross-connect, and the BGP configuration on their side. AWS is responsible for the port on their side and for the DX location. This split means that a DX incident often requires a support case with both AWS and the colocation provider, and the runbook needs to reflect that. The first step in any DX runbook is to determine which side of the cross-connect the problem is on, because that determines who gets paged.
The metrics that matter for DX monitoring are the connection state, the BGP session state, the light levels on the optic, and the throughput on each VIF. CloudWatch exposes connection-level metrics, and the customer router exposes the BGP and interface metrics. The SLO implication is that DX availability is a shared responsibility: AWS publishes a service level agreement for the connection itself, but the end-to-end availability of the hybrid path depends on the customer's router and the colocation facility. A realistic SLO for a hybrid path is therefore lower than the SLO for a purely in-AWS path, and the exam sometimes tests whether you understand that distinction.
The runbook shape for a DX incident is: confirm the scope (one VIF, one connection, or one location), check the BGP session state on both sides, check the physical layer metrics, and then decide whether to fail over to the backup path. The failover decision is usually automated via BGP, but the runbook should still document the manual steps in case the automation fails. The post-incident review should capture whether the failure was within the expected failure domain — a single connection failing is expected and should be transparent; a location failing is expected and should be covered by the second location; a BGP misconfiguration is not expected and should drive a configuration change.
8. Edge Cases and Exam Gotchas
The first gotcha is that a LAG is not redundant. The exam will describe a customer with four connections in a LAG and ask whether the design is resilient, and the correct answer is no, because all four connections terminate at the same DX location. The second gotcha is that a VPN backup is not a substitute for a second DX connection. The exam will describe a customer with one DX connection and a VPN backup and ask whether the design meets a high-availability requirement, and the correct answer depends on the stated RTO — a VPN backup is acceptable for a moderate RTO but not for a mission-critical one.
The third gotcha is that a Private VIF cannot reach a VPC in another region without a DX Gateway. The exam will describe a customer with a DX connection in us-east-1 and a VPC in us-west-2 and ask how to connect them, and the correct answer is a DX Gateway, not a second connection. The fourth gotcha is that a Public VIF requires public IP space and a public ASN on the customer side, which many customers do not have. The exam will describe a customer that wants private access to S3 and ask what they need, and the correct answer includes the public IP and ASN requirement, not just the VIF.
The fifth gotcha is that the DX Gateway is a global resource, not a regional one. The exam will describe a customer with connections in two regions and ask how many DX Gateways they need, and the correct answer is one, because a single DX Gateway can be associated with connections in multiple regions. The sixth gotcha is that a Transit VIF requires a DX Gateway associated with a Transit Gateway, and the TGW must have the appropriate route tables configured. The exam will describe a customer with a Transit VIF and a TGW and ask why traffic is not reaching a specific VPC, and the correct answer is usually a missing TGW route table association or propagation.
The seventh gotcha is that the BGP ASN on the customer side must be a public ASN if the customer wants to use a Public VIF, but can be a private ASN for Private and Transit VIFs. The exam will describe a customer with a private ASN and ask whether they can use a Public VIF, and the correct answer is no. The eighth gotcha is that the DX connection itself does not encrypt traffic. The exam will describe a customer with a regulatory requirement for encryption in transit and ask whether DX satisfies it, and the correct answer is no — DX is private but not encrypted, and the customer must add MACsec or a VPN on top of DX if encryption is required.
9. Direct Connect vs. the Alternatives
The comparison the exam cares about most is Direct Connect versus Site-to-Site VPN. Both provide hybrid connectivity, but they differ on every axis that matters: DX is private, consistent, and expensive with a long lead time; VPN is public, variable, cheap, and fast to deploy. The exam will describe a scenario and expect you to pick the one that matches the constraints, and the most common correct answer is a combination — DX for the primary path and VPN for the backup. The second comparison is DX versus Transit Gateway, which is not really a comparison because they solve different problems: DX is the physical path into AWS, and TGW is the routing fabric inside AWS. A mature design uses both, with a Transit VIF connecting the DX circuit to the TGW.
The third comparison is DX versus PrivateLink, which is also not a direct comparison but shows up in exam scenarios about cross-account access. DX connects a customer network to AWS; PrivateLink exposes a specific service to a specific consumer. If the scenario is about a customer's data center reaching AWS, the answer is DX. If the scenario is about one AWS account consuming a service in another AWS account, the answer is PrivateLink. The fourth comparison is DX versus AWS VPN CloudHub, which is a hub-and-spoke VPN topology for connecting multiple branch offices. CloudHub is the right answer when the customer has many small sites and no dedicated circuits; DX is the right answer when the customer has a small number of large sites and needs consistent performance.
| Option | Path | Lead time | Pick it when… |
|---|---|---|---|
| Direct Connect | Private, dedicated | Weeks | Consistent latency, high bandwidth, regulatory privacy |
| Site-to-Site VPN | Public internet, encrypted | Minutes | Fast deployment, low cost, backup path |
| Transit Gateway | Inside AWS | Minutes | Routing between many VPCs and on-prem |
| PrivateLink | Inside AWS, service-level | Minutes | Exposing one service to one consumer, CIDR overlap |
| VPN CloudHub | Public internet, hub-and-spoke | Minutes | Many small branch sites, no dedicated circuits |
The rule of thumb to carry into the exam: if the scenario mentions a physical data center, a regulatory requirement, or a latency-sensitive workload, start with Direct Connect. If it mentions a deadline measured in days, a small branch office, or a proof of concept, start with VPN. If it mentions many VPCs or many accounts, add a Transit Gateway and use a Transit VIF. If it mentions a single service being consumed across accounts, use PrivateLink instead. The exam rarely asks for a single service in isolation; it asks for the combination that satisfies all the stated constraints, and the combination is usually DX plus TGW plus a VPN backup.
Hands-on Lab: Designing a Redundant DX Architecture (45 min)
This lab walks through the design of a redundant Direct Connect architecture for a fictional enterprise, using two connections at two different DX locations with active/passive BGP path preference. The goal is not to provision real hardware — that would take weeks — but to produce the design artifacts that a real DX deployment requires: the topology diagram, the BGP configuration, the failover test plan, and the runbook. Work through the steps in order and document each decision.
Step 1 — Define the requirements. Assume the enterprise has a primary data center in Ashburn, Virginia, and a secondary data center in Columbus, Ohio. The workload is a financial trading application that requires consistent sub-10ms latency to us-east-1 and cannot tolerate more than 5 minutes of downtime. The compliance team requires that traffic never traverse the public internet. Write down the RTO (5 minutes), the RPO (near-zero, since the application is stateless), and the compliance constraint. These three numbers drive every subsequent decision.
Step 2 — Choose the DX locations. AWS publishes a list of Direct Connect locations, and the two connections must terminate at different locations to survive a location-level failure. For this design, choose two DX locations in the us-east-1 region that are geographically separate — for example, one in Ashburn and one in a different metro area within the region's DX footprint. Document why the two locations are independent: they have separate power, separate cooling, and separate network paths to the AWS backbone. If the two locations share a single fiber conduit, the design is not actually redundant, and the exam will test whether you noticed.
Step 3 — Choose the port speeds. The trading application needs consistent bandwidth, so choose dedicated 10 Gbps connections rather than hosted connections. Document the capacity calculation: peak throughput, headroom for burst, and the cost difference between 1 Gbps and 10 Gbps. If the application's peak is 2 Gbps, a 10 Gbps port provides 5x headroom, which is reasonable for a latency-sensitive workload. A 1 Gbps port would be cheaper but would leave no room for growth.
Step 4 — Design the VIF topology. The enterprise has a Transit Gateway that already aggregates its VPCs, so the correct choice is a Transit VIF on each connection, both associated with a single DX Gateway. Document the DX Gateway associations: one DX Gateway, associated with the Transit Gateway, which is associated with the VPCs. If the enterprise also needs private access to S3, add a Public VIF on one of the connections and document the public IP and ASN requirements.
Step 5 — Configure BGP for active/passive. Both connections will advertise the same on-premises prefixes, so without intervention BGP will treat them as equal-cost and split traffic. To make one path preferred, apply AS path prepending on the standby connection: prepend the customer ASN three times so the AS path looks longer and the primary path is preferred. Document the exact BGP configuration on the customer router, including the neighbor IPs, the ASNs, and the prepend policy. Explain why AS path prepending was chosen over local preference: it is simpler and does not require coordination with AWS.
Step 6 — Add a VPN backup. Even with two DX connections, a simultaneous failure of both locations is possible, so add a Site-to-Site VPN over the internet as a last-resort backup. Document the VPN's BGP configuration and the failover behavior: the VPN should only be preferred when both DX paths are down, which means its AS path must be longer than the DX paths even after prepending. This is a subtle configuration detail that the exam sometimes tests.
Step 7 — Write the failover test plan. A redundant design is only as good as its tested failover. Write a test plan that covers three scenarios: a single connection failure (shut down one BGP session and verify traffic shifts to the other connection within the RTO), a single location failure (simulate a location outage and verify the same), and a simultaneous failure (shut down both DX paths and verify the VPN takes over). For each scenario, document the expected behavior, the actual behavior, and the time to recover.
Step 8 — Write the runbook. The runbook should cover the first three diagnostic steps for a DX incident: confirm the scope, check the BGP session state on both sides, and check the physical layer metrics. It should also document the escalation path: which issues go to AWS Support, which go to the colocation provider, and which are internal. Finally, it should document the manual failover procedure in case the BGP automation fails.
Step 9 — Review against the exam. Go back through the design and identify every decision that maps to an exam concept: the VIF type, the DX Gateway, the BGP attributes, the resiliency tier, and the VPN backup. For each, write one sentence explaining why the alternative would be wrong. This is the step that turns a lab exercise into exam preparation.
Scenario Question Drills (20 min)
Q1. A company needs the highest resiliency Direct Connect design for a mission-critical workload. What should they provision?
Q2. A company has a single VPC in a single AWS account and wants private connectivity from its data center. Which VIF type is the simplest correct choice?
Q3. A company has four DX connections bundled into a Link Aggregation Group at a single DX location. Is this design resilient to a location-level failure?
Q4. A company has a DX connection in us-east-1 and needs to reach a VPC in us-west-2 over the same connection. What is required?
Q5. A company wants private access to Amazon S3 from its data center without traversing the public internet. Which VIF type is required?
Q6. A company has two DX connections at two different locations, both advertising the same prefixes. Traffic is splitting across both paths, but the company wants one path preferred. What should they configure?
Q7. A company has a DX connection and a Site-to-Site VPN backup. The DX connection fails. What is the most likely operational impact?
Q8. A company has a Transit Gateway aggregating its VPCs and wants to connect its data center over Direct Connect. Which VIF type should they use?
Q9. A company has a private ASN and wants to use a Public VIF to reach S3. What is the problem?
Q10. A company has a regulatory requirement that all traffic in transit must be encrypted. Does a Direct Connect connection satisfy this requirement on its own?
Q11. A company has DX connections in two different AWS regions and wants to reach VPCs in both. How many DX Gateways do they need?
Q12. A company's DX connection is up and the BGP session is established, but on-prem traffic to a VPC is going over the VPN instead of the DX circuit. What is the most likely cause?
Q13. A company needs to move 300 TB of historical data to AWS and has a 500 Mbps DX connection. What is the recommended approach?
Q14. A company wants to connect 40 small branch offices to AWS with minimal cost and no dedicated circuits. What is the best fit?
Q15. A company has a Transit VIF and a Transit Gateway, but traffic from on-prem is not reaching a specific VPC. What is the most likely cause?
Peek into Tomorrow: Hybrid DNS
Today's design gets packets from the data center to AWS over a private circuit, but it leaves a question unanswered: once the packets arrive, how do the two sides resolve each other's names? A server in the data center that needs to reach a database in a private Route 53 hosted zone has to resolve that name somehow, and the default behavior — querying the public DNS — will not work, because the hosted zone is private. Conversely, an EC2 instance in a VPC that needs to reach an on-premises service by its internal hostname has to resolve a name that AWS's default resolver has never heard of. The DX circuit carries the packets, but it does not carry the DNS answers.
Tomorrow's topic, Route 53 Resolver, is the piece that closes this gap. It introduces Inbound endpoints, which let on-premises resolvers query AWS private hosted zones, and Outbound endpoints with forwarding rules, which let AWS resources resolve on-premises names over the same DX or VPN path we designed today. The open question is how to configure the forwarding rules so that only the queries that need to go on-prem actually go on-prem, and how to make the resolver endpoints resilient across Availability Zones. The answers depend on the same redundancy thinking we applied to the DX connections, but at the DNS layer instead of the packet layer.
Sources
- AWS Direct Connect User Guide — What Is AWS Direct Connect?
- AWS Direct Connect User Guide — Creating a Virtual Interface
- AWS Direct Connect User Guide — Direct Connect Gateways
- AWS Direct Connect User Guide — Resiliency Toolkit
- AWS Direct Connect User Guide — Link Aggregation Groups
- AWS Direct Connect User Guide — Routing Policies and BGP Communities
- AWS Whitepaper — Amazon Virtual Private Cloud Connectivity Options
- AWS Whitepaper — Building a Scalable and Secure Multi-VPC AWS Network Infrastructure
- AWS Site-to-Site VPN User Guide
- AWS Whitepaper — Control Planes and Data Planes