Day 12 of 70 · Week 2
Day 12 / 70 Week 2 of 14 Phase 1: Multi-Account Governance & Networking

AWS Route 53 Resolver (Hybrid DNS)

🕑 ~58 min read · 2 services covered
Route 53 Resolver Inbound/Outbound Endpoints

Recap: From Private Links to Private Names

Day 11 established that Direct Connect gives you a dedicated private path into AWS, and that the VIF you choose determines what that path can reach: a Private VIF terminates on a single VPC, a Transit VIF terminates on a Transit Gateway and therefore reaches many VPCs, and a Public VIF reaches AWS public service endpoints. It also established that resilience is a property of the physical design, not the logical one — redundant connections at separate DX locations with BGP failover, not a single link with a backup configuration on paper.

Today extends that material in a direction the VIF discussion deliberately left open. A Transit VIF gives you IP reachability between on-premises networks and every VPC attached to the Transit Gateway, and BGP failover gives you a path that survives a link failure. Neither of those things gives you name resolution. An on-premises application that can ping a database's private IP address still cannot connect to it if the connection string uses a hostname, and an EC2 instance that can reach an on-premises server over the Transit VIF still cannot resolve corp.internal to an address. Route 53 Resolver is the piece that closes that gap, and it is the reason hybrid connectivity questions on the exam almost always have a DNS component hiding inside them.

Foundations You'll Need Today

DNS: Turning Names Into Addresses

Computers find each other by IP address, but humans remember names. DNS is the system that bridges the two: when an application is told to connect to db.prod.internal, something has to look that name up and hand back an IP address before a single packet can be sent. The component that performs the lookup is called a resolver, and the process of asking one resolver, which may in turn ask others, is called recursion. The answers themselves live in records, which are grouped into containers called hosted zones — think of a hosted zone as the file cabinet for one domain, and the records as the individual index cards inside it. Two failure responses are worth knowing by name because today's lesson distinguishes them: NXDOMAIN means "I asked and the name genuinely does not exist," while SERVFAIL or a timeout means "I could not get an answer at all." That difference is the difference between a missing record and a broken path, and it is the first thing you check when DNS misbehaves.

VPCs, Subnets, and the Address at .2

A VPC is your own private, isolated network inside AWS — a walled-off slice of the cloud where you decide the addressing. You carve that address space into subnets, each of which lives in one Availability Zone and holds a range of IP addresses written in CIDR notation, such as 10.0.0.0/16. CIDR notation is just a compact way of saying "this many addresses starting here": the number after the slash tells you how large the range is, so a smaller number means a bigger network. Every VPC automatically gets a built-in DNS resolver, and AWS reserves a fixed address for it at the base of the VPC's range plus two — which is why a VPC starting at 10.0.0.0 has its resolver at 10.0.0.2. Instances learn to use that address through the DHCP option set, which is the configuration AWS hands to each instance at boot telling it things like which DNS server to use. You never create this resolver; it is simply there, and today's endpoints exist to extend it beyond the VPC.

Elastic Network Interfaces (ENIs)

An elastic network interface is the virtual network card attached to a resource in a VPC. It holds one or more private IP addresses, it lives in exactly one subnet, and it is the thing that actually sends and receives packets — when you think "this server has IP 10.0.1.25," you are really thinking about its ENI. ENIs matter today because Resolver endpoints are not servers you log into; they are sets of ENIs that AWS places in the subnets you choose, each with its own IP address. That is why the lesson keeps talking about "the subnets containing the endpoint ENIs" — those subnets are where the DNS traffic physically originates, and therefore the subnets whose networking rules govern whether the query can leave.

Route Tables: How Packets Know Where to Go

A route table is a list of rules attached to a subnet that answers one question for every outgoing packet: where do I send this next? Each rule pairs a destination range with a target — for example, "traffic for 10.0.0.0/8 goes to the Transit Gateway." If no rule matches a packet's destination, the packet is dropped, and the sender experiences a timeout with no explanation. This is the single most important concept for today's troubleshooting, because a DNS query leaving an outbound endpoint is just an ordinary packet: it needs a matching route to the on-premises DNS server's address, or it dies silently inside the VPC. A perfectly configured Resolver rule with a missing route looks exactly like no rule at all.

Private vs. Public Hosted Zones

A public hosted zone holds records that anyone on the internet can look up — this is how amazon.com resolves for the whole world. A private hosted zone holds records that only resolve from inside the VPCs you explicitly associate with it, which is how you give internal resources names like db.prod.internal that the outside world cannot even see. The association is the whole mechanism: a private zone attached to a VPC answers queries from that VPC, and a private zone attached to nothing answers nobody. Today's lesson is largely about the gap this creates — a private zone is perfectly resolvable from inside AWS and completely invisible from on-premises until you build something to bridge the two.

With that grounding, here's why Route 53 Resolver exists and what problem it actually solves.

1. Why Hybrid DNS Is a Recurring Exam Problem

Hybrid DNS shows up on the SAP-C02 exam because it is the point where three otherwise independent design decisions collide: network topology, identity of names, and failure domain. A candidate who has memorized that Transit Gateway connects VPCs and that Direct Connect connects on-premises can still fail a question that asks why an application works from one subnet but not another, or why a migration cutover succeeded for IP-based connections and failed for hostname-based ones. The exam uses DNS as a discriminator precisely because it is invisible in a network diagram. Two architectures that look identical on a whiteboard behave completely differently depending on whether name resolution was designed or assumed.

The domain mapping is mostly Domain 1, the organizational complexity and hybrid connectivity domain, but the topic bleeds into Domain 2 as well. Any resilience question that involves a secondary region, a standby data center, or a failover runbook has a DNS dependency, because failover that changes an IP address without changing the name is not failover from the application's perspective. Route 53 Resolver also appears in migration scenarios, where the question is how to keep on-premises clients resolving the same names before, during, and after a cutover without editing application configuration. The recurring shape of the question is a constraint that rules out the obvious answer: you cannot use a public hosted zone because the names must not be publicly resolvable, you cannot edit every client's /etc/hosts, and you cannot put a NAT gateway in the path because the traffic must stay on private connectivity.

What makes the topic tractable is that the AWS-side mechanism is small. There are two endpoint types, one rule type, and a handful of quotas. The complexity is entirely in the direction of the query and in the routing that carries it. Once you can state, for any given query, which resolver asks which resolver and over which path, the exam questions become mechanical. The rest of this lesson is organized around building that directional intuition, because almost every wrong answer on a hybrid DNS question comes from reversing the direction of an endpoint.

2. How Resolver Actually Answers a Query

Every VPC has a Route 53 Resolver, and it is not something you create. It is reachable at the VPC CIDR base plus two — the classic 10.0.0.2 for a 10.0.0.0/16 VPC — and it is the resolver that every EC2 instance uses by default through the DHCP option set. This resolver answers queries for names in private hosted zones associated with the VPC, resolves public names by recursing against the internet, and, critically, is the component that Resolver endpoints extend. The endpoints do not replace the VPC resolver; they give it a way to talk to resolvers outside the VPC.

An inbound endpoint is a set of elastic network interfaces, one per subnet you specify, each with an IP address from that subnet. On-premises DNS servers are configured to forward queries for AWS-hosted names to those IP addresses. The query arrives at the ENI, is handed to the VPC resolver, and is answered from the private hosted zone associated with the VPC. The direction is on-premises to AWS. An outbound endpoint is also a set of ENIs, but it works the other way: the VPC resolver, when it encounters a name that matches a Resolver rule, forwards the query from those ENIs to the target IP addresses named in the rule. The direction is AWS to on-premises. Both endpoint types are regional constructs, both require at least two IP addresses for availability, and both are billed per endpoint-hour plus per query.

The rule type that matters for hybrid DNS is the forwarding rule, which pairs a domain name with a list of target IP addresses and optionally a port. Rules are associated with VPCs, and association is what makes the rule visible to that VPC's resolver. A rule for corp.internal with targets pointing at two on-premises DNS servers means that any query from an associated VPC for a name ending in corp.internal is forwarded out through the outbound endpoint rather than being recursed publicly. This is why the mechanism is described as conditional forwarding: the rule is a condition, and only matching queries take the outbound path. Everything else follows the default behavior.

The path the forwarded query takes is the part candidates most often get wrong. The outbound endpoint's ENIs live in your VPC subnets, so the query leaves the VPC as ordinary IP traffic from those ENI addresses. It reaches the on-premises resolver only if the route tables and the hybrid connectivity — the Transit VIF or the VPN attachment from Day 11 — actually carry traffic from those subnets to the on-premises DNS server addresses. A correctly configured Resolver rule with a broken route is indistinguishable from a rule that does not exist, and the failure looks like a timeout rather than a resolution error.

3. The Core Decision Boundary: Which Direction Is the Query Going?

Nearly every hybrid DNS scenario reduces to a single question: who is asking, and where does the answer live? If the asker is on-premises and the answer lives in a private hosted zone in AWS, you need an inbound endpoint. If the asker is in AWS and the answer lives on-premises, you need an outbound endpoint plus a forwarding rule. The two are not interchangeable, they are not alternatives to each other, and a design that needs both directions needs both endpoints. The exam exploits this by describing a scenario in terms of the application rather than the query direction, so the first move is always to restate the scenario as a query.

The second-order decision is whether the on-premises side is even capable of forwarding. An inbound endpoint only works if the on-premises DNS servers can be configured with a conditional forwarder pointing at the endpoint IPs. In an environment where the on-premises DNS is a managed appliance the team does not control, that may be impossible, and the design has to change — for example, by resolving the AWS names through a different mechanism entirely. The exam rarely states this constraint explicitly, but it appears as a phrase like "the on-premises DNS team will not modify their configuration," which is a signal that the inbound endpoint answer is wrong.

The third decision is scope. Resolver rules are associated with VPCs, and a rule associated with one VPC does not apply to another. In a multi-account landing zone, the standard pattern is to create the outbound endpoint and the rules in a central networking account, then share the rules with workload accounts using AWS RAM — the same sharing mechanism from Day 6. This is why hybrid DNS questions often have a multi-account flavor: the endpoint is expensive and should exist once, but the rules need to be visible everywhere.

Scenario shapeDirectionComponentWhere it lives
On-prem app resolves db.prod.internal in a private hosted zoneOn-prem → AWSInbound endpointVPC that owns the zone
EC2 instance resolves erp.corp.internal on-premAWS → on-premOutbound endpoint + forwarding ruleCentral networking VPC
Both directions neededBidirectionalBoth endpoint typesCentral networking VPC
On-prem DNS cannot be modifiedOn-prem → AWSNot solvable with inbound endpointRedesign required
Many accounts need the same on-prem namesAWS → on-premOutbound endpoint + RAM-shared rulesCentral networking account

4. Endpoint Configuration and What Each Choice Costs

The first configuration decision is how many IP addresses to give an endpoint and in which subnets. AWS requires at least two IP addresses per endpoint for availability, and the guidance is to place them in different Availability Zones. This is not a formality: an endpoint with both addresses in one AZ is a single-AZ dependency for every DNS query in the environment, and DNS failure is unusually total because it breaks new connections rather than degrading existing ones. The cost of the second AZ is one more ENI and one more IP address, which is trivial next to the blast radius of a zonal DNS outage.

The second decision is the protocol and port. Resolver endpoints support DNS over UDP and TCP on port 53, and they also support DNS over HTTPS on port 443 for outbound rules. The DoH option matters in environments where the on-premises resolver is exposed through a service that only accepts HTTPS, or where the security team requires encrypted DNS on the wire. Choosing DoH changes the target port in the rule and the protocol the on-premises side must accept; it does not change the direction or the rule semantics.

The third decision is rule scope and priority. A forwarding rule can be associated with many VPCs, and rules are evaluated by most-specific match, so a rule for corp.internal and a rule for hr.corp.internal can coexist with the more specific one winning. This is how you split a single on-premises namespace across two different DNS servers — for example, sending HR queries to a dedicated resolver while everything else goes to the general one. The tradeoff is operational: every additional rule is another object to keep in sync with the on-premises DNS team's zone delegation, and a stale rule produces timeouts rather than clean failures.

The fourth decision is whether to use a single outbound endpoint for everything or to segment by environment. A single endpoint in a shared networking VPC is cheaper and simpler, but it means production and development DNS traffic share a path and a failure domain. Segmenting gives you isolation at the cost of more endpoints and more rules. The exam generally rewards the shared-endpoint answer unless the scenario explicitly calls for isolation, because the question is usually testing whether you know that endpoints are regional and shareable rather than whether you can build the most elaborate design.

5. Sizing, Limits, and the Numbers Worth Knowing

Resolver endpoints are sized by IP address count, not by a throughput tier you select. Each IP address on an endpoint can handle a certain number of queries per second, and the practical guidance is to add addresses when you approach that ceiling rather than to over-provision up front. The relevant published figures are that each Resolver endpoint IP address supports up to 10,000 queries per second, and that an endpoint can have up to six IP addresses. That gives a single endpoint a theoretical ceiling of 60,000 queries per second, which is far beyond most enterprise DNS loads but is worth knowing because it tells you the scaling lever is address count.

The quota that catches people is the number of Resolver rules per Region, which is 1,000 by default, and the number of VPC associations per rule, which is also 1,000. Neither is likely to bind in a normal environment, but the association quota matters in a large landing zone where a single shared rule is associated with every VPC in the organization. The other quota worth remembering is that you can have up to 600 inbound and outbound endpoints per Region combined, which is effectively unlimited for design purposes but confirms that endpoints are a per-Region resource.

On the pricing side, the model is endpoint-hour plus query volume. Each endpoint IP address is billed per hour whether or not it serves traffic, and each query through an endpoint is billed per million queries. This is the reason the shared-endpoint pattern is the standard recommendation: an endpoint per VPC multiplies the hourly charge by the number of VPCs for no functional benefit, since a single endpoint in a shared VPC can serve rules associated with many VPCs. The exam does not usually ask for exact prices, but it does ask for the design that avoids unnecessary per-VPC endpoints, and the cost model is the justification.

ItemDefault / published valueWhy it matters
Minimum IP addresses per endpoint2Availability; place in separate AZs
Maximum IP addresses per endpoint6Scaling lever for query volume
Queries per second per endpoint IPUp to 10,000Determines when to add addresses
Resolver rules per Region1,000Rarely binding; relevant in large orgs
VPC associations per rule1,000Binds in very large landing zones
Endpoints per Region (in + out)600Confirms endpoints are regional

6. Failure Modes and What They Look Like in Production

The most common failure is a forwarding rule whose target IP addresses are unreachable from the outbound endpoint's subnets. The symptom is a query timeout rather than an NXDOMAIN, and the diagnostic move is to check the route table for the subnet containing the outbound endpoint ENIs and confirm there is a route to the on-premises DNS server addresses over the Transit Gateway or VPN. This failure is easy to misdiagnose as an on-premises DNS problem because the query does leave AWS; it simply never arrives. A useful first check is to look at the Resolver query logs, which record the query, the rule that matched, and the response code, making the difference between "no rule matched" and "rule matched but target timed out" immediately visible.

The second failure is a rule that matches more than intended. A rule for internal will capture anything.internal, including names that should have resolved publicly or from a private hosted zone. The symptom is that a name which used to resolve now times out, and the cause is usually a rule added for one purpose that overlaps an existing namespace. Because rules are evaluated by most-specific match, the fix is either to narrow the rule or to add a more specific rule that restores the intended behavior. This is the DNS equivalent of a route table with an overly broad prefix, and it fails the same way: silently, for a subset of names.

The third failure is an endpoint with insufficient IP addresses for the query volume, which manifests as intermittent timeouts under load rather than a consistent failure. Because DNS is usually the first step of every new connection, the application-level symptom is a spike in connection errors that correlates with traffic, not with any single service. The diagnostic move is to look at Resolver endpoint metrics and query logs for throttling, and the fix is to add IP addresses to the endpoint. This failure is worth internalizing because it is the one that looks like an application bug and is actually a capacity problem.

The fourth failure is asymmetric resolution, where AWS can resolve on-premises names but on-premises cannot resolve AWS names, or vice versa. This is almost always a missing endpoint in one direction rather than a misconfiguration of an existing one, and it is the failure mode the exam is most likely to describe, because it tests whether you understand that inbound and outbound endpoints are separate components with separate purposes.

7. The Operational and SRE Angle

DNS is a dependency of nearly every other service, which makes it a poor candidate for best-effort monitoring. The metrics that matter are the Resolver endpoint metrics: query volume, the count of queries that resulted in a timeout or a SERVFAIL, and the health of the endpoint ENIs themselves. A useful alarm is one on the rate of SERVFAIL responses from an outbound endpoint, because that metric rises before applications start failing and gives you a window to act. A second alarm on the endpoint's query volume relative to its IP address count gives you early warning on the capacity failure described above.

Resolver query logging is the operational tool that makes the rest tractable. It can be sent to CloudWatch Logs, S3, or Kinesis Data Firehose, and it records the query name, the query type, the VPC, the rule that matched, and the response code. In an incident, this is the difference between guessing and knowing: you can see whether the query arrived, which rule handled it, and what the target returned. The cost consideration is volume, since every query is logged, so the usual pattern is to log to S3 with a lifecycle policy rather than to CloudWatch Logs for high-volume environments.

From an SLO perspective, hybrid DNS is a shared dependency that should have its own availability target, because it is on the critical path for every service that resolves a name across the boundary. The runbook shape follows from the failure modes: first check whether the query is arriving at the endpoint at all, then whether a rule matched, then whether the target was reachable, then whether the endpoint had capacity. That ordering moves from the most common cause to the least, and it maps cleanly onto the query logs, which is what makes it a runbook rather than a checklist.

One operational detail worth building in early is a synthetic check that resolves a known on-premises name from a known AWS subnet on a schedule. This is the DNS equivalent of the canary from Day 30: it detects a broken path before an application does, and it distinguishes a DNS failure from an application failure during an incident. Without it, the first signal of a DNS problem is usually a cascade of unrelated application alarms.

8. Edge Cases and Exam Gotchas

The single most common mistake is reversing the endpoint direction. The mnemonic that works is to name the endpoint after where the query is going, not where it is coming from: an inbound endpoint is inbound to AWS, and an outbound endpoint is outbound from AWS. If the scenario says on-premises clients need to resolve AWS names, the answer contains "inbound." If it says AWS resources need to resolve on-premises names, the answer contains "outbound." Candidates who reason from the perspective of the client rather than the query get this backwards under time pressure.

The second gotcha is forgetting that the endpoint is regional. A Resolver endpoint in us-east-1 does nothing for a VPC in us-west-2, and a multi-region design needs an endpoint in each Region that has VPCs needing hybrid resolution. This is the same regional-versus-global distinction that appears with Transit Gateway and Direct Connect, and it is a reliable source of wrong answers.

The third gotcha is assuming that a private hosted zone is automatically resolvable from on-premises once an inbound endpoint exists. The zone must be associated with the VPC that the inbound endpoint serves, and the on-premises resolver must be configured to forward to the endpoint IPs. Both halves are required, and the exam will describe a scenario where one half is missing.

The fourth gotcha is the shared-services account pattern. In a multi-account environment, the outbound endpoint and the rules belong in the central networking account, and the rules are shared with workload accounts via AWS RAM. A design that creates an endpoint in every account is functionally correct but is the wrong answer on cost and operational grounds, and the exam will usually include it as a distractor.

The fifth gotcha is the interaction with private hosted zone resolution order. A query for a name that exists both in a private hosted zone and publicly resolves to the private answer from within an associated VPC, which is usually the intent but occasionally is not. If the scenario requires the public answer from inside the VPC, the zone association is the thing to change, not the Resolver rule.

9. Resolver vs. the Alternatives It Gets Confused With

Route 53 Resolver is not the only way to make names resolve across a boundary, and the exam tests whether you can tell the alternatives apart. The most common confusion is with Route 53 private hosted zones themselves. A private hosted zone is the container for the records; Resolver is the mechanism that lets a resolver outside the VPC query those records. Creating a private hosted zone does not make it resolvable from on-premises, and creating an inbound endpoint does not create any records. They are complementary, and a scenario that mentions both is usually testing whether you know which one is missing.

The second confusion is with VPC peering and Transit Gateway. Those provide IP reachability, which is necessary but not sufficient for name resolution. A design that connects two VPCs with peering and expects names to resolve across them will fail unless the private hosted zone is associated with both VPCs or a Resolver rule bridges them. The exam uses this by describing a scenario where connectivity is already established and the problem is still that names do not resolve.

The third confusion is with the public DNS of the on-premises environment. If the on-premises names are publicly resolvable, no Resolver configuration is needed at all, and the correct answer is to do nothing. The exam will sometimes include a scenario where the names are public and the elaborate hybrid DNS answer is a distractor. The signal is a phrase like "the records are already in a public zone."

OptionWhat it providesPick it when…
Route 53 private hosted zoneRecord storage, resolvable inside associated VPCsYou need internal names for AWS resources only
Inbound Resolver endpointOn-prem → AWS name resolutionOn-premises clients must resolve AWS private names
Outbound Resolver endpoint + ruleAWS → on-prem name resolutionAWS resources must resolve on-premises names
VPC peering / Transit GatewayIP reachability onlyYou need packets to flow, not names to resolve
Public DNS recordsGlobal resolutionThe names are not sensitive and can be public

The practical rule is to ask what is missing: if packets cannot flow, the answer is a connectivity service; if packets flow but names do not resolve, the answer is a Resolver endpoint in the correct direction; if the names should not be private at all, the answer is a public zone. Most hybrid DNS exam questions are one of those three, and identifying which one is the whole task.

Hands-on Lab: Conditional Forwarding to an On-Premises Resolver (45 min)

This lab builds the AWS-to-on-premises direction end to end: an outbound endpoint, a forwarding rule for corp.internal, and the verification that a query from an EC2 instance in a workload VPC actually reaches an on-premises DNS server and returns an answer. It assumes a working hybrid path already exists — a Transit Gateway attachment or a VPN — because the lab is about DNS, not connectivity, and a broken path will make every step look like a DNS failure.

Step 1 — Establish the baseline. From an EC2 instance in the workload VPC, run dig erp.corp.internal and confirm it fails with a timeout or NXDOMAIN. Record the exact failure, because the difference between the two is the difference between "no rule matched" and "rule matched but the target was unreachable," and you will use that distinction later. Also run dig amazon.com to confirm the VPC resolver itself is healthy and recursing normally.

Step 2 — Create the outbound endpoint. In the central networking VPC, create a Route 53 Resolver outbound endpoint and select two subnets in different Availability Zones. Note the IP addresses assigned to the endpoint ENIs; you will need them if you later need to allow this traffic through a firewall or security group. Confirm the endpoint reaches a status of Operational before continuing, since rules associated with a non-operational endpoint will silently fail.

Step 3 — Create the forwarding rule. Create a Resolver rule with the domain name corp.internal, the rule type set to Forward, and two target IP addresses corresponding to the on-premises DNS servers. If your on-premises resolver requires encrypted DNS, set the protocol to DoH and the port to 443; otherwise leave the default UDP/TCP on port 53. Associate the rule with the workload VPC, not just the networking VPC, since association is what makes the rule visible to the resolver that the EC2 instance uses.

Step 4 — Verify the path. From the EC2 instance, run dig erp.corp.internal again. If it resolves, the rule and the path are both working. If it times out, check the route table for the subnets containing the outbound endpoint ENIs and confirm there is a route to the on-premises DNS server addresses over the Transit Gateway or VPN. This is the step where most failures occur, and it is a routing problem rather than a Resolver problem.

Step 5 — Enable query logging and inspect it. Turn on Resolver query logging to CloudWatch Logs for the workload VPC, repeat the query, and find the log entry. Confirm that it records the query name, the rule that matched, and the response code. This is the artifact you will use in an incident, and seeing it once makes the runbook concrete.

Step 6 — Test the failure mode deliberately. Remove one of the two target IP addresses from the rule and repeat the query. With a single target, the query should still succeed; then remove the second target and confirm the query fails with a timeout. This demonstrates that the rule's target list is the availability mechanism for the on-premises side, and it mirrors the two-AZ requirement on the AWS side.

Step 7 — Clean up. Delete the rule, delete the outbound endpoint, and disable query logging. Leaving an endpoint running costs money per hour regardless of traffic, which is itself a lesson about why the shared-endpoint pattern matters.

Scenario Question Drills (20 min)

Q1. On-premises servers need to resolve names in a private Route 53 hosted zone. What do you configure?

A. A Route 53 Resolver outbound endpoint
B. A Route 53 Resolver inbound endpoint, with on-premises DNS forwarding queries to it
C. A public hosted zone instead
D. NAT gateway DNS forwarding
Correct answer: B. Inbound endpoints expose AWS private DNS to on-premises resolvers; outbound endpoints do the reverse. The direction of the query determines the endpoint type.

Q2. An EC2 instance in a workload VPC must resolve erp.corp.internal, which is served by an on-premises DNS server. The VPC is connected to on-premises via a Transit Gateway. What is required?

A. An inbound Resolver endpoint in the workload VPC
B. An outbound Resolver endpoint plus a forwarding rule for corp.internal, associated with the workload VPC
C. A private hosted zone for corp.internal in the workload VPC
D. Nothing — Transit Gateway resolves names automatically
Correct answer: B. The query originates in AWS and the answer lives on-premises, so the outbound direction is required. Transit Gateway provides IP reachability only, never name resolution.

Q3. A forwarding rule for corp.internal is configured with two target IP addresses, but queries from the workload VPC time out. The on-premises DNS servers are confirmed healthy. What is the most likely cause?

A. The rule needs to be associated with the VPC that owns the private hosted zone
B. The subnets containing the outbound endpoint ENIs have no route to the on-premises DNS server addresses
C. The endpoint needs a third IP address
D. Resolver rules only work with DNS over HTTPS
Correct answer: B. A correctly configured rule with an unreachable target produces a timeout, not an NXDOMAIN. The outbound endpoint's subnets must have a route to the on-premises resolver addresses.

Q4. A company has 40 workload accounts, all of which need to resolve the same on-premises namespace. What is the most cost-effective and operationally sound design?

A. Create an outbound endpoint and rule in each of the 40 accounts
B. Create one outbound endpoint and rule in a central networking account and share the rule with the workload accounts via AWS RAM
C. Create an inbound endpoint in each account
D. Use a public hosted zone for the on-premises namespace
Correct answer: B. Endpoints are billed per hour per IP address, so a single shared endpoint in a central account avoids multiplying the cost by 40. AWS RAM shares the rule across accounts.

Q5. A Resolver outbound endpoint has two IP addresses, both in the same Availability Zone. What is the risk?

A. None — the endpoint is still highly available because it has two addresses
B. The endpoint becomes a single-AZ dependency, so an AZ failure breaks all hybrid name resolution
C. The endpoint will not reach an Operational state
D. Queries will be limited to 10,000 per second total
Correct answer: B. Two addresses in one AZ protect against an ENI failure but not against an AZ failure. The standard guidance is to place endpoint IPs in separate AZs.

Q6. An on-premises DNS team refuses to modify their resolver configuration. On-premises clients must resolve AWS private names. What is the correct assessment?

A. An inbound endpoint will work without any on-premises change
B. An inbound endpoint requires the on-premises resolver to forward queries to it, so the design is not viable as stated
C. An outbound endpoint will solve it
D. A public hosted zone will solve it
Correct answer: B. An inbound endpoint only receives queries if the on-premises resolver is configured to forward to it. Without that change, the endpoint is unreachable from on-premises.

Q7. A forwarding rule for internal was added, and now a name that previously resolved from a private hosted zone times out. What happened?

A. The private hosted zone was deleted
B. The broad rule captures names that should have resolved from the private hosted zone, and the forwarded query times out
C. Resolver rules always override private hosted zones and cannot be fixed
D. The endpoint needs more IP addresses
Correct answer: B. Rules are evaluated by most-specific match, so an overly broad rule can capture names intended for a private hosted zone. Narrow the rule or add a more specific one.

Q8. Which Resolver feature lets you determine, during an incident, whether a query matched a rule and what the target returned?

A. VPC Flow Logs
B. Resolver query logging
C. CloudTrail data events
D. AWS Config
Correct answer: B. Resolver query logging records the query name, the rule that matched, and the response code, which is exactly the information needed to distinguish a missing rule from an unreachable target.

Q9. A workload in us-west-2 needs to resolve on-premises names, but the only Resolver outbound endpoint is in us-east-1. What is the problem?

A. None — Resolver endpoints are global
B. Resolver endpoints are regional, so a second endpoint is required in us-west-2
C. The rule must be recreated in us-west-2 but the endpoint can stay
D. The endpoint must be converted to an inbound endpoint
Correct answer: B. Resolver endpoints are regional resources. A VPC in another Region needs its own endpoint and rules in that Region.

Q10. An application resolves a name that exists both in a private hosted zone and in public DNS. From inside an associated VPC, which answer is returned?

A. The public answer, because public DNS takes precedence
B. The private hosted zone answer, because the zone is associated with the VPC
C. Both answers, and the client chooses
D. Neither, because the conflict causes an error
Correct answer: B. A private hosted zone associated with the VPC takes precedence for names it contains, which is usually the intent but occasionally is not.

Q11. What is the maximum number of IP addresses a single Resolver endpoint can have, and why does it matter?

A. Two, because that is the availability minimum
B. Six, and it is the scaling lever for query volume since each address supports up to 10,000 queries per second
C. Ten, and it is fixed by the VPC CIDR size
D. There is no limit
Correct answer: B. An endpoint supports up to six IP addresses, and each address handles up to 10,000 queries per second, so address count is how you scale an endpoint.

Q12. A company needs encrypted DNS on the wire between AWS and an on-premises resolver that only accepts HTTPS. What should they configure?

A. A forwarding rule with the protocol set to DoH and the port set to 443
B. An inbound endpoint with TLS termination
C. A VPN connection, which encrypts DNS automatically
D. Resolver does not support encrypted DNS
Correct answer: A. Outbound forwarding rules support DNS over HTTPS on port 443, which is the option for on-premises resolvers that only accept HTTPS.

Q13. Two VPCs are connected by peering, and both need to resolve names from a private hosted zone owned by one of them. What is required?

A. Nothing — peering propagates DNS automatically
B. Associate the private hosted zone with both VPCs, or bridge them with a Resolver rule
C. Create an inbound endpoint in each VPC
D. Create a public hosted zone instead
Correct answer: B. Peering provides IP reachability only. Name resolution requires the zone to be associated with both VPCs or a Resolver rule to bridge them.

Q14. An on-premises application can reach an AWS database by IP address but fails when configured with its hostname. Connectivity is confirmed working. What is the most likely gap?

A. The security group is blocking port 53
B. No inbound Resolver endpoint exists, so on-premises cannot resolve the AWS private name
C. The database needs a public endpoint
D. The Transit Gateway needs a new route table
Correct answer: B. IP reachability without name resolution is the classic hybrid DNS gap. The on-premises side needs an inbound endpoint to forward queries to.

Q15. A team wants early warning that hybrid DNS is degrading before applications fail. What should they alarm on?

A. EC2 CPU utilization in the workload VPC
B. The rate of SERVFAIL responses and query volume relative to endpoint IP address count
C. The number of private hosted zone records
D. CloudTrail API call volume
Correct answer: B. SERVFAIL rate rises before applications fail, and query volume against endpoint capacity gives early warning of the throttling failure mode.

Peek into Tomorrow

Today's design assumed that the thing being resolved is a name, and that the network path to the answer already exists. That assumption is comfortable when the two sides of the boundary are your own VPCs and your own data center, because you control the CIDR ranges on both ends and can route between them. It becomes uncomfortable the moment the consumer is a different organization with its own address plan, because now the question is not just how names resolve but whether the two networks can be connected at all without one side renumbering.

Tomorrow's topic, PrivateLink, is the answer to that narrower question, and it is narrower in a way that matters. An Interface Endpoint is an ENI in the consumer's VPC that represents a service in the provider's VPC, which means the consumer never needs a route to the provider's CIDR range and the two address spaces can overlap completely without conflict. That property is what makes PrivateLink the standard answer for exposing one microservice to many consumer accounts, and it is also why it is not a general-purpose replacement for Transit Gateway. The open question today leaves behind is where the boundary sits between service-level exposure and network-level connectivity, and tomorrow's comparison of PrivateLink against VPC peering is where that line gets drawn.

Sources