AWS Route 53 Resolver (Hybrid DNS)
Recap: From Private Links to Private Names
Day 11 established that Direct Connect gives you a dedicated private path into AWS, and that the VIF you choose determines what that path can reach: a Private VIF terminates on a single VPC, a Transit VIF terminates on a Transit Gateway and therefore reaches many VPCs, and a Public VIF reaches AWS public service endpoints. It also established that resilience is a property of the physical design, not the logical one — redundant connections at separate DX locations with BGP failover, not a single link with a backup configuration on paper.
Today extends that material in a direction the VIF discussion deliberately left open. A Transit VIF gives you IP reachability between on-premises networks and every VPC attached to the Transit Gateway, and BGP failover gives you a path that survives a link failure. Neither of those things gives you name resolution. An on-premises application that can ping a database's private IP address still cannot connect to it if the connection string uses a hostname, and an EC2 instance that can reach an on-premises server over the Transit VIF still cannot resolve corp.internal to an address. Route 53 Resolver is the piece that closes that gap, and it is the reason hybrid connectivity questions on the exam almost always have a DNS component hiding inside them.
Foundations You'll Need Today
DNS: Turning Names Into Addresses
Computers find each other by IP address, but humans remember names. DNS is the system that bridges the two: when an application is told to connect to db.prod.internal, something has to look that name up and hand back an IP address before a single packet can be sent. The component that performs the lookup is called a resolver, and the process of asking one resolver, which may in turn ask others, is called recursion. The answers themselves live in records, which are grouped into containers called hosted zones — think of a hosted zone as the file cabinet for one domain, and the records as the individual index cards inside it. Two failure responses are worth knowing by name because today's lesson distinguishes them: NXDOMAIN means "I asked and the name genuinely does not exist," while SERVFAIL or a timeout means "I could not get an answer at all." That difference is the difference between a missing record and a broken path, and it is the first thing you check when DNS misbehaves.
VPCs, Subnets, and the Address at .2
A VPC is your own private, isolated network inside AWS — a walled-off slice of the cloud where you decide the addressing. You carve that address space into subnets, each of which lives in one Availability Zone and holds a range of IP addresses written in CIDR notation, such as 10.0.0.0/16. CIDR notation is just a compact way of saying "this many addresses starting here": the number after the slash tells you how large the range is, so a smaller number means a bigger network. Every VPC automatically gets a built-in DNS resolver, and AWS reserves a fixed address for it at the base of the VPC's range plus two — which is why a VPC starting at 10.0.0.0 has its resolver at 10.0.0.2. Instances learn to use that address through the DHCP option set, which is the configuration AWS hands to each instance at boot telling it things like which DNS server to use. You never create this resolver; it is simply there, and today's endpoints exist to extend it beyond the VPC.
Elastic Network Interfaces (ENIs)
An elastic network interface is the virtual network card attached to a resource in a VPC. It holds one or more private IP addresses, it lives in exactly one subnet, and it is the thing that actually sends and receives packets — when you think "this server has IP 10.0.1.25," you are really thinking about its ENI. ENIs matter today because Resolver endpoints are not servers you log into; they are sets of ENIs that AWS places in the subnets you choose, each with its own IP address. That is why the lesson keeps talking about "the subnets containing the endpoint ENIs" — those subnets are where the DNS traffic physically originates, and therefore the subnets whose networking rules govern whether the query can leave.
Route Tables: How Packets Know Where to Go
A route table is a list of rules attached to a subnet that answers one question for every outgoing packet: where do I send this next? Each rule pairs a destination range with a target — for example, "traffic for 10.0.0.0/8 goes to the Transit Gateway." If no rule matches a packet's destination, the packet is dropped, and the sender experiences a timeout with no explanation. This is the single most important concept for today's troubleshooting, because a DNS query leaving an outbound endpoint is just an ordinary packet: it needs a matching route to the on-premises DNS server's address, or it dies silently inside the VPC. A perfectly configured Resolver rule with a missing route looks exactly like no rule at all.
Private vs. Public Hosted Zones
A public hosted zone holds records that anyone on the internet can look up — this is how amazon.com resolves for the whole world. A private hosted zone holds records that only resolve from inside the VPCs you explicitly associate with it, which is how you give internal resources names like db.prod.internal that the outside world cannot even see. The association is the whole mechanism: a private zone attached to a VPC answers queries from that VPC, and a private zone attached to nothing answers nobody. Today's lesson is largely about the gap this creates — a private zone is perfectly resolvable from inside AWS and completely invisible from on-premises until you build something to bridge the two.
With that grounding, here's why Route 53 Resolver exists and what problem it actually solves.
1. Why Hybrid DNS Is a Recurring Exam Problem
Hybrid DNS shows up on the SAP-C02 exam because it is the point where three otherwise independent design decisions collide: network topology, identity of names, and failure domain. A candidate who has memorized that Transit Gateway connects VPCs and that Direct Connect connects on-premises can still fail a question that asks why an application works from one subnet but not another, or why a migration cutover succeeded for IP-based connections and failed for hostname-based ones. The exam uses DNS as a discriminator precisely because it is invisible in a network diagram. Two architectures that look identical on a whiteboard behave completely differently depending on whether name resolution was designed or assumed.
The domain mapping is mostly Domain 1, the organizational complexity and hybrid connectivity domain, but the topic bleeds into Domain 2 as well. Any resilience question that involves a secondary region, a standby data center, or a failover runbook has a DNS dependency, because failover that changes an IP address without changing the name is not failover from the application's perspective. Route 53 Resolver also appears in migration scenarios, where the question is how to keep on-premises clients resolving the same names before, during, and after a cutover without editing application configuration. The recurring shape of the question is a constraint that rules out the obvious answer: you cannot use a public hosted zone because the names must not be publicly resolvable, you cannot edit every client's /etc/hosts, and you cannot put a NAT gateway in the path because the traffic must stay on private connectivity.
What makes the topic tractable is that the AWS-side mechanism is small. There are two endpoint types, one rule type, and a handful of quotas. The complexity is entirely in the direction of the query and in the routing that carries it. Once you can state, for any given query, which resolver asks which resolver and over which path, the exam questions become mechanical. The rest of this lesson is organized around building that directional intuition, because almost every wrong answer on a hybrid DNS question comes from reversing the direction of an endpoint.
2. How Resolver Actually Answers a Query
Every VPC has a Route 53 Resolver, and it is not something you create. It is reachable at the VPC CIDR base plus two — the classic 10.0.0.2 for a 10.0.0.0/16 VPC — and it is the resolver that every EC2 instance uses by default through the DHCP option set. This resolver answers queries for names in private hosted zones associated with the VPC, resolves public names by recursing against the internet, and, critically, is the component that Resolver endpoints extend. The endpoints do not replace the VPC resolver; they give it a way to talk to resolvers outside the VPC.
An inbound endpoint is a set of elastic network interfaces, one per subnet you specify, each with an IP address from that subnet. On-premises DNS servers are configured to forward queries for AWS-hosted names to those IP addresses. The query arrives at the ENI, is handed to the VPC resolver, and is answered from the private hosted zone associated with the VPC. The direction is on-premises to AWS. An outbound endpoint is also a set of ENIs, but it works the other way: the VPC resolver, when it encounters a name that matches a Resolver rule, forwards the query from those ENIs to the target IP addresses named in the rule. The direction is AWS to on-premises. Both endpoint types are regional constructs, both require at least two IP addresses for availability, and both are billed per endpoint-hour plus per query.
The rule type that matters for hybrid DNS is the forwarding rule, which pairs a domain name with a list of target IP addresses and optionally a port. Rules are associated with VPCs, and association is what makes the rule visible to that VPC's resolver. A rule for corp.internal with targets pointing at two on-premises DNS servers means that any query from an associated VPC for a name ending in corp.internal is forwarded out through the outbound endpoint rather than being recursed publicly. This is why the mechanism is described as conditional forwarding: the rule is a condition, and only matching queries take the outbound path. Everything else follows the default behavior.
The path the forwarded query takes is the part candidates most often get wrong. The outbound endpoint's ENIs live in your VPC subnets, so the query leaves the VPC as ordinary IP traffic from those ENI addresses. It reaches the on-premises resolver only if the route tables and the hybrid connectivity — the Transit VIF or the VPN attachment from Day 11 — actually carry traffic from those subnets to the on-premises DNS server addresses. A correctly configured Resolver rule with a broken route is indistinguishable from a rule that does not exist, and the failure looks like a timeout rather than a resolution error.
3. The Core Decision Boundary: Which Direction Is the Query Going?
Nearly every hybrid DNS scenario reduces to a single question: who is asking, and where does the answer live? If the asker is on-premises and the answer lives in a private hosted zone in AWS, you need an inbound endpoint. If the asker is in AWS and the answer lives on-premises, you need an outbound endpoint plus a forwarding rule. The two are not interchangeable, they are not alternatives to each other, and a design that needs both directions needs both endpoints. The exam exploits this by describing a scenario in terms of the application rather than the query direction, so the first move is always to restate the scenario as a query.
The second-order decision is whether the on-premises side is even capable of forwarding. An inbound endpoint only works if the on-premises DNS servers can be configured with a conditional forwarder pointing at the endpoint IPs. In an environment where the on-premises DNS is a managed appliance the team does not control, that may be impossible, and the design has to change — for example, by resolving the AWS names through a different mechanism entirely. The exam rarely states this constraint explicitly, but it appears as a phrase like "the on-premises DNS team will not modify their configuration," which is a signal that the inbound endpoint answer is wrong.
The third decision is scope. Resolver rules are associated with VPCs, and a rule associated with one VPC does not apply to another. In a multi-account landing zone, the standard pattern is to create the outbound endpoint and the rules in a central networking account, then share the rules with workload accounts using AWS RAM — the same sharing mechanism from Day 6. This is why hybrid DNS questions often have a multi-account flavor: the endpoint is expensive and should exist once, but the rules need to be visible everywhere.
| Scenario shape | Direction | Component | Where it lives |
|---|---|---|---|
On-prem app resolves db.prod.internal in a private hosted zone | On-prem → AWS | Inbound endpoint | VPC that owns the zone |
EC2 instance resolves erp.corp.internal on-prem | AWS → on-prem | Outbound endpoint + forwarding rule | Central networking VPC |
| Both directions needed | Bidirectional | Both endpoint types | Central networking VPC |
| On-prem DNS cannot be modified | On-prem → AWS | Not solvable with inbound endpoint | Redesign required |
| Many accounts need the same on-prem names | AWS → on-prem | Outbound endpoint + RAM-shared rules | Central networking account |
4. Endpoint Configuration and What Each Choice Costs
The first configuration decision is how many IP addresses to give an endpoint and in which subnets. AWS requires at least two IP addresses per endpoint for availability, and the guidance is to place them in different Availability Zones. This is not a formality: an endpoint with both addresses in one AZ is a single-AZ dependency for every DNS query in the environment, and DNS failure is unusually total because it breaks new connections rather than degrading existing ones. The cost of the second AZ is one more ENI and one more IP address, which is trivial next to the blast radius of a zonal DNS outage.
The second decision is the protocol and port. Resolver endpoints support DNS over UDP and TCP on port 53, and they also support DNS over HTTPS on port 443 for outbound rules. The DoH option matters in environments where the on-premises resolver is exposed through a service that only accepts HTTPS, or where the security team requires encrypted DNS on the wire. Choosing DoH changes the target port in the rule and the protocol the on-premises side must accept; it does not change the direction or the rule semantics.
The third decision is rule scope and priority. A forwarding rule can be associated with many VPCs, and rules are evaluated by most-specific match, so a rule for corp.internal and a rule for hr.corp.internal can coexist with the more specific one winning. This is how you split a single on-premises namespace across two different DNS servers — for example, sending HR queries to a dedicated resolver while everything else goes to the general one. The tradeoff is operational: every additional rule is another object to keep in sync with the on-premises DNS team's zone delegation, and a stale rule produces timeouts rather than clean failures.
The fourth decision is whether to use a single outbound endpoint for everything or to segment by environment. A single endpoint in a shared networking VPC is cheaper and simpler, but it means production and development DNS traffic share a path and a failure domain. Segmenting gives you isolation at the cost of more endpoints and more rules. The exam generally rewards the shared-endpoint answer unless the scenario explicitly calls for isolation, because the question is usually testing whether you know that endpoints are regional and shareable rather than whether you can build the most elaborate design.
5. Sizing, Limits, and the Numbers Worth Knowing
Resolver endpoints are sized by IP address count, not by a throughput tier you select. Each IP address on an endpoint can handle a certain number of queries per second, and the practical guidance is to add addresses when you approach that ceiling rather than to over-provision up front. The relevant published figures are that each Resolver endpoint IP address supports up to 10,000 queries per second, and that an endpoint can have up to six IP addresses. That gives a single endpoint a theoretical ceiling of 60,000 queries per second, which is far beyond most enterprise DNS loads but is worth knowing because it tells you the scaling lever is address count.
The quota that catches people is the number of Resolver rules per Region, which is 1,000 by default, and the number of VPC associations per rule, which is also 1,000. Neither is likely to bind in a normal environment, but the association quota matters in a large landing zone where a single shared rule is associated with every VPC in the organization. The other quota worth remembering is that you can have up to 600 inbound and outbound endpoints per Region combined, which is effectively unlimited for design purposes but confirms that endpoints are a per-Region resource.
On the pricing side, the model is endpoint-hour plus query volume. Each endpoint IP address is billed per hour whether or not it serves traffic, and each query through an endpoint is billed per million queries. This is the reason the shared-endpoint pattern is the standard recommendation: an endpoint per VPC multiplies the hourly charge by the number of VPCs for no functional benefit, since a single endpoint in a shared VPC can serve rules associated with many VPCs. The exam does not usually ask for exact prices, but it does ask for the design that avoids unnecessary per-VPC endpoints, and the cost model is the justification.
| Item | Default / published value | Why it matters |
|---|---|---|
| Minimum IP addresses per endpoint | 2 | Availability; place in separate AZs |
| Maximum IP addresses per endpoint | 6 | Scaling lever for query volume |
| Queries per second per endpoint IP | Up to 10,000 | Determines when to add addresses |
| Resolver rules per Region | 1,000 | Rarely binding; relevant in large orgs |
| VPC associations per rule | 1,000 | Binds in very large landing zones |
| Endpoints per Region (in + out) | 600 | Confirms endpoints are regional |
6. Failure Modes and What They Look Like in Production
The most common failure is a forwarding rule whose target IP addresses are unreachable from the outbound endpoint's subnets. The symptom is a query timeout rather than an NXDOMAIN, and the diagnostic move is to check the route table for the subnet containing the outbound endpoint ENIs and confirm there is a route to the on-premises DNS server addresses over the Transit Gateway or VPN. This failure is easy to misdiagnose as an on-premises DNS problem because the query does leave AWS; it simply never arrives. A useful first check is to look at the Resolver query logs, which record the query, the rule that matched, and the response code, making the difference between "no rule matched" and "rule matched but target timed out" immediately visible.
The second failure is a rule that matches more than intended. A rule for internal will capture anything.internal, including names that should have resolved publicly or from a private hosted zone. The symptom is that a name which used to resolve now times out, and the cause is usually a rule added for one purpose that overlaps an existing namespace. Because rules are evaluated by most-specific match, the fix is either to narrow the rule or to add a more specific rule that restores the intended behavior. This is the DNS equivalent of a route table with an overly broad prefix, and it fails the same way: silently, for a subset of names.
The third failure is an endpoint with insufficient IP addresses for the query volume, which manifests as intermittent timeouts under load rather than a consistent failure. Because DNS is usually the first step of every new connection, the application-level symptom is a spike in connection errors that correlates with traffic, not with any single service. The diagnostic move is to look at Resolver endpoint metrics and query logs for throttling, and the fix is to add IP addresses to the endpoint. This failure is worth internalizing because it is the one that looks like an application bug and is actually a capacity problem.
The fourth failure is asymmetric resolution, where AWS can resolve on-premises names but on-premises cannot resolve AWS names, or vice versa. This is almost always a missing endpoint in one direction rather than a misconfiguration of an existing one, and it is the failure mode the exam is most likely to describe, because it tests whether you understand that inbound and outbound endpoints are separate components with separate purposes.
7. The Operational and SRE Angle
DNS is a dependency of nearly every other service, which makes it a poor candidate for best-effort monitoring. The metrics that matter are the Resolver endpoint metrics: query volume, the count of queries that resulted in a timeout or a SERVFAIL, and the health of the endpoint ENIs themselves. A useful alarm is one on the rate of SERVFAIL responses from an outbound endpoint, because that metric rises before applications start failing and gives you a window to act. A second alarm on the endpoint's query volume relative to its IP address count gives you early warning on the capacity failure described above.
Resolver query logging is the operational tool that makes the rest tractable. It can be sent to CloudWatch Logs, S3, or Kinesis Data Firehose, and it records the query name, the query type, the VPC, the rule that matched, and the response code. In an incident, this is the difference between guessing and knowing: you can see whether the query arrived, which rule handled it, and what the target returned. The cost consideration is volume, since every query is logged, so the usual pattern is to log to S3 with a lifecycle policy rather than to CloudWatch Logs for high-volume environments.
From an SLO perspective, hybrid DNS is a shared dependency that should have its own availability target, because it is on the critical path for every service that resolves a name across the boundary. The runbook shape follows from the failure modes: first check whether the query is arriving at the endpoint at all, then whether a rule matched, then whether the target was reachable, then whether the endpoint had capacity. That ordering moves from the most common cause to the least, and it maps cleanly onto the query logs, which is what makes it a runbook rather than a checklist.
One operational detail worth building in early is a synthetic check that resolves a known on-premises name from a known AWS subnet on a schedule. This is the DNS equivalent of the canary from Day 30: it detects a broken path before an application does, and it distinguishes a DNS failure from an application failure during an incident. Without it, the first signal of a DNS problem is usually a cascade of unrelated application alarms.
8. Edge Cases and Exam Gotchas
The single most common mistake is reversing the endpoint direction. The mnemonic that works is to name the endpoint after where the query is going, not where it is coming from: an inbound endpoint is inbound to AWS, and an outbound endpoint is outbound from AWS. If the scenario says on-premises clients need to resolve AWS names, the answer contains "inbound." If it says AWS resources need to resolve on-premises names, the answer contains "outbound." Candidates who reason from the perspective of the client rather than the query get this backwards under time pressure.
The second gotcha is forgetting that the endpoint is regional. A Resolver endpoint in us-east-1 does nothing for a VPC in us-west-2, and a multi-region design needs an endpoint in each Region that has VPCs needing hybrid resolution. This is the same regional-versus-global distinction that appears with Transit Gateway and Direct Connect, and it is a reliable source of wrong answers.
The third gotcha is assuming that a private hosted zone is automatically resolvable from on-premises once an inbound endpoint exists. The zone must be associated with the VPC that the inbound endpoint serves, and the on-premises resolver must be configured to forward to the endpoint IPs. Both halves are required, and the exam will describe a scenario where one half is missing.
The fourth gotcha is the shared-services account pattern. In a multi-account environment, the outbound endpoint and the rules belong in the central networking account, and the rules are shared with workload accounts via AWS RAM. A design that creates an endpoint in every account is functionally correct but is the wrong answer on cost and operational grounds, and the exam will usually include it as a distractor.
The fifth gotcha is the interaction with private hosted zone resolution order. A query for a name that exists both in a private hosted zone and publicly resolves to the private answer from within an associated VPC, which is usually the intent but occasionally is not. If the scenario requires the public answer from inside the VPC, the zone association is the thing to change, not the Resolver rule.
9. Resolver vs. the Alternatives It Gets Confused With
Route 53 Resolver is not the only way to make names resolve across a boundary, and the exam tests whether you can tell the alternatives apart. The most common confusion is with Route 53 private hosted zones themselves. A private hosted zone is the container for the records; Resolver is the mechanism that lets a resolver outside the VPC query those records. Creating a private hosted zone does not make it resolvable from on-premises, and creating an inbound endpoint does not create any records. They are complementary, and a scenario that mentions both is usually testing whether you know which one is missing.
The second confusion is with VPC peering and Transit Gateway. Those provide IP reachability, which is necessary but not sufficient for name resolution. A design that connects two VPCs with peering and expects names to resolve across them will fail unless the private hosted zone is associated with both VPCs or a Resolver rule bridges them. The exam uses this by describing a scenario where connectivity is already established and the problem is still that names do not resolve.
The third confusion is with the public DNS of the on-premises environment. If the on-premises names are publicly resolvable, no Resolver configuration is needed at all, and the correct answer is to do nothing. The exam will sometimes include a scenario where the names are public and the elaborate hybrid DNS answer is a distractor. The signal is a phrase like "the records are already in a public zone."
| Option | What it provides | Pick it when… |
|---|---|---|
| Route 53 private hosted zone | Record storage, resolvable inside associated VPCs | You need internal names for AWS resources only |
| Inbound Resolver endpoint | On-prem → AWS name resolution | On-premises clients must resolve AWS private names |
| Outbound Resolver endpoint + rule | AWS → on-prem name resolution | AWS resources must resolve on-premises names |
| VPC peering / Transit Gateway | IP reachability only | You need packets to flow, not names to resolve |
| Public DNS records | Global resolution | The names are not sensitive and can be public |
The practical rule is to ask what is missing: if packets cannot flow, the answer is a connectivity service; if packets flow but names do not resolve, the answer is a Resolver endpoint in the correct direction; if the names should not be private at all, the answer is a public zone. Most hybrid DNS exam questions are one of those three, and identifying which one is the whole task.
Hands-on Lab: Conditional Forwarding to an On-Premises Resolver (45 min)
This lab builds the AWS-to-on-premises direction end to end: an outbound endpoint, a forwarding rule for corp.internal, and the verification that a query from an EC2 instance in a workload VPC actually reaches an on-premises DNS server and returns an answer. It assumes a working hybrid path already exists — a Transit Gateway attachment or a VPN — because the lab is about DNS, not connectivity, and a broken path will make every step look like a DNS failure.
Step 1 — Establish the baseline. From an EC2 instance in the workload VPC, run dig erp.corp.internal and confirm it fails with a timeout or NXDOMAIN. Record the exact failure, because the difference between the two is the difference between "no rule matched" and "rule matched but the target was unreachable," and you will use that distinction later. Also run dig amazon.com to confirm the VPC resolver itself is healthy and recursing normally.
Step 2 — Create the outbound endpoint. In the central networking VPC, create a Route 53 Resolver outbound endpoint and select two subnets in different Availability Zones. Note the IP addresses assigned to the endpoint ENIs; you will need them if you later need to allow this traffic through a firewall or security group. Confirm the endpoint reaches a status of Operational before continuing, since rules associated with a non-operational endpoint will silently fail.
Step 3 — Create the forwarding rule. Create a Resolver rule with the domain name corp.internal, the rule type set to Forward, and two target IP addresses corresponding to the on-premises DNS servers. If your on-premises resolver requires encrypted DNS, set the protocol to DoH and the port to 443; otherwise leave the default UDP/TCP on port 53. Associate the rule with the workload VPC, not just the networking VPC, since association is what makes the rule visible to the resolver that the EC2 instance uses.
Step 4 — Verify the path. From the EC2 instance, run dig erp.corp.internal again. If it resolves, the rule and the path are both working. If it times out, check the route table for the subnets containing the outbound endpoint ENIs and confirm there is a route to the on-premises DNS server addresses over the Transit Gateway or VPN. This is the step where most failures occur, and it is a routing problem rather than a Resolver problem.
Step 5 — Enable query logging and inspect it. Turn on Resolver query logging to CloudWatch Logs for the workload VPC, repeat the query, and find the log entry. Confirm that it records the query name, the rule that matched, and the response code. This is the artifact you will use in an incident, and seeing it once makes the runbook concrete.
Step 6 — Test the failure mode deliberately. Remove one of the two target IP addresses from the rule and repeat the query. With a single target, the query should still succeed; then remove the second target and confirm the query fails with a timeout. This demonstrates that the rule's target list is the availability mechanism for the on-premises side, and it mirrors the two-AZ requirement on the AWS side.
Step 7 — Clean up. Delete the rule, delete the outbound endpoint, and disable query logging. Leaving an endpoint running costs money per hour regardless of traffic, which is itself a lesson about why the shared-endpoint pattern matters.
Scenario Question Drills (20 min)
Q1. On-premises servers need to resolve names in a private Route 53 hosted zone. What do you configure?
Q2. An EC2 instance in a workload VPC must resolve erp.corp.internal, which is served by an on-premises DNS server. The VPC is connected to on-premises via a Transit Gateway. What is required?
Q3. A forwarding rule for corp.internal is configured with two target IP addresses, but queries from the workload VPC time out. The on-premises DNS servers are confirmed healthy. What is the most likely cause?
Q4. A company has 40 workload accounts, all of which need to resolve the same on-premises namespace. What is the most cost-effective and operationally sound design?
Q5. A Resolver outbound endpoint has two IP addresses, both in the same Availability Zone. What is the risk?
Q6. An on-premises DNS team refuses to modify their resolver configuration. On-premises clients must resolve AWS private names. What is the correct assessment?
Q7. A forwarding rule for internal was added, and now a name that previously resolved from a private hosted zone times out. What happened?
Q8. Which Resolver feature lets you determine, during an incident, whether a query matched a rule and what the target returned?
Q9. A workload in us-west-2 needs to resolve on-premises names, but the only Resolver outbound endpoint is in us-east-1. What is the problem?
Q10. An application resolves a name that exists both in a private hosted zone and in public DNS. From inside an associated VPC, which answer is returned?
Q11. What is the maximum number of IP addresses a single Resolver endpoint can have, and why does it matter?
Q12. A company needs encrypted DNS on the wire between AWS and an on-premises resolver that only accepts HTTPS. What should they configure?
Q13. Two VPCs are connected by peering, and both need to resolve names from a private hosted zone owned by one of them. What is required?
Q14. An on-premises application can reach an AWS database by IP address but fails when configured with its hostname. Connectivity is confirmed working. What is the most likely gap?
Q15. A team wants early warning that hybrid DNS is degrading before applications fail. What should they alarm on?
Peek into Tomorrow
Today's design assumed that the thing being resolved is a name, and that the network path to the answer already exists. That assumption is comfortable when the two sides of the boundary are your own VPCs and your own data center, because you control the CIDR ranges on both ends and can route between them. It becomes uncomfortable the moment the consumer is a different organization with its own address plan, because now the question is not just how names resolve but whether the two networks can be connected at all without one side renumbering.
Tomorrow's topic, PrivateLink, is the answer to that narrower question, and it is narrower in a way that matters. An Interface Endpoint is an ENI in the consumer's VPC that represents a service in the provider's VPC, which means the consumer never needs a route to the provider's CIDR range and the two address spaces can overlap completely without conflict. That property is what makes PrivateLink the standard answer for exposing one microservice to many consumer accounts, and it is also why it is not a general-purpose replacement for Transit Gateway. The open question today leaves behind is where the boundary sits between service-level exposure and network-level connectivity, and tomorrow's comparison of PrivateLink against VPC peering is where that line gets drawn.
Sources
- Amazon Route 53 Developer Guide — Resolving DNS queries between VPCs and your network
- Route 53 Developer Guide — Forwarding outbound DNS queries to your network
- Route 53 Developer Guide — Resolving inbound DNS queries from your network
- Route 53 Developer Guide — Route 53 Resolver DNS Firewall
- Route 53 Developer Guide — Resolver query logging
- AWS General Reference — Route 53 service quotas
- AWS Whitepaper — Hybrid Cloud DNS Options for Amazon VPC
- AWS Whitepaper — Building a Scalable and Secure Multi-VPC AWS Network Infrastructure