A custom DNS server in your DHCP options turns off private DNS for every endpoint in the VPC
Private hosted zones, interface endpoint private DNS, and Route 53 Resolver rules all live in one place: the VPC resolver at .2. Point the DHCP option set at your own DNS servers and every instance stops asking it. Nothing is deleted, nothing is disabled, and every one of those features quietly stops applying.
· Venerable Networks
A VPC has one resolver that knows about its private things. It sits at the base of the VPC CIDR plus two — 10.0.0.2 in a 10.0.0.0/16 — and it is also reachable at 169.254.169.253. Every private hosted zone associated with the VPC, every interface endpoint with private DNS enabled, and every Route 53 Resolver rule associated with the VPC is answered by that one address and by nothing else.
Instances find it through the DHCP option set, which defaults to AmazonProvidedDNS. Change that to point at your own servers — a corporate Active Directory, a pair of resolvers in a shared services VPC, an on-premises forwarder — and every instance in the VPC stops asking .2. Which means every private answer .2 was giving stops arriving.
What breaks, and how it looks
Three features share the dependency, and each fails in its own way.
Interface endpoint private DNS. With private DNS enabled on an endpoint for, say, sqs.us-east-1.amazonaws.com, the VPC resolver answers that name with the endpoint's private addresses. A custom DNS server has never heard of your endpoint and answers with the public addresses, as AWS's troubleshooting article states:
For private IPs, you must send the DNS queries to the Amazon provided DNS of the VPC where you created the interface endpoint. The Amazon provided DNS is the base of the VPC CIDR plus two.
The workload then tries to reach SQS at a public address from a subnet with no internet route and times out — the same symptom, line for line, as private DNS being off on the endpoint. Except here private DNS is on. The console shows it enabled. The endpoint is healthy. It is simply never consulted.
Private hosted zones. A zone associated with the VPC exists only in the VPC resolver. Ask a corporate DNS server for db.internal.example and it returns NXDOMAIN, or forwards to the public internet and gets NXDOMAIN from there. The zone is intact, the records are right, the association is in place, and no query ever reaches it.
Resolver rules. Forwarding rules — the ones that send corp.internal to on-premises — are evaluated by the VPC resolver. Bypass it and they are not evaluated. In the common case this is masked, because the custom DNS server is the on-premises resolver the rule was forwarding to, so corporate names still work. What silently stops working is everything else the rule set did: the system rules, the autodefined reverse-lookup rules, and any rule for a domain the corporate server does not hold.
Why the console cannot tell you
Nothing is misconfigured in the sense the console understands. The endpoint's private DNS attribute is true. The hosted zone's VPC association is present. The Resolver rules are associated. The DHCP option set is valid and the DNS servers in it answer queries. Every object reports the state its author intended.
The relationship that broke is one no object owns: which resolver the instances send their queries to. That is decided by the DHCP option set and consumed by the guest OS, and there is no AWS resource whose status reflects whether the private DNS features are actually being reached.
The one-line diagnosis
From any instance in the VPC:
cat /etc/resolv.conf
# nameserver 10.50.0.10
# nameserver 10.50.0.11If the nameservers are anything other than .2 of the VPC CIDR (or 169.254.169.253), private DNS features are not in the resolution path. Then prove it by asking both resolvers the same question:
dig +short sqs.us-east-1.amazonaws.com @10.0.0.2
# 10.0.11.47
# 10.0.12.112
dig +short sqs.us-east-1.amazonaws.com @10.50.0.10
# 3.231.xx.xx
# 52.4.xx.xxPrivate addresses from the VPC resolver, public ones from the custom server. The feature works. The instance is asking the wrong server.
The fix is not "use AmazonProvidedDNS"
Teams point DHCP at custom servers for reasons — Active Directory integration, on-premises name resolution, centralised logging — and reverting throws those away. The actual fix keeps the VPC resolver in the path and gets the custom resolution through it.
Resolver rules, outbound endpoint. Leave the DHCP option set at AmazonProvidedDNS. Create a Route 53 Resolver outbound endpoint in the VPC and a forwarding rule for each domain the corporate servers own — corp.example, the AD domain, reverse zones if needed — targeting those servers. Instances ask .2; .2 answers private hosted zones and endpoints itself and forwards the corporate domains onward. Every feature works because every query goes through the one resolver that knows about all of them. This is the architecture AWS's hybrid DNS guidance is built around, and it is the right answer in nearly every case.
Conditional forwarding on the custom server, the other direction. Keep the custom DNS servers in DHCP and configure them to forward amazonaws.com, your private zone names, and anything else AWS-resolved to the VPC's .2. It works, and it means the list of forwarded domains has to be maintained by hand on a server AWS does not manage, and every new endpoint or private zone is a change ticket against a different team. It is the answer when the DHCP change genuinely cannot be reverted.
Either way, the test is the same: dig an endpoint's service name from an instance with no @ and get a private address back.
The audit
Any VPC whose DHCP option set does not name AmazonProvidedDNS and which also has private DNS features is a candidate:
# DHCP option sets with custom name servers
aws ec2 describe-dhcp-options \
--query 'DhcpOptions[?DhcpConfigurations[?Key==`domain-name-servers` && !contains(Values[].Value, `AmazonProvidedDNS`)]].DhcpOptionsId' \
--output text
# VPCs using one of them, and whether any has a private-DNS endpoint
aws ec2 describe-vpcs --filters "Name=dhcp-options-id,Values=$DHCP_ID" \
--query 'Vpcs[].VpcId' --output text
aws ec2 describe-vpc-endpoints --filters "Name=vpc-id,Values=$VPC_ID" \
--query 'VpcEndpoints[?PrivateDnsEnabled==`true`].ServiceName' --output textA VPC that appears in both lists has endpoints whose private DNS is enabled and unreachable. The same applies to any private hosted zone associated with it.
The general shape
This is the third post on this site about the VPC resolver being a single point that several features depend on without saying so. A Resolver rule outranks a private hosted zone because both are evaluated in the same place and one wins. Private DNS off and this failure produce the same timeout from opposite causes — the resolver was asked and had nothing to say, versus the resolver had the answer and was never asked.
Underneath all three: the private DNS features of a VPC are properties of one resolver, not of the VPC. Anything that moves queries away from that resolver turns every one of them off at once, and no individual feature will report it.