The one network limit that does not grow when you scale up
Every instance can send 1,024 packets per second to the VPC resolver, per network interface, and the number is the same on a t3.nano and a u-24tb1.metal. Throttled queries never reach anything that logs, so the symptom is intermittent DNS failure with no evidence anywhere — and scaling up, the instinctive fix, changes nothing.
· Venerable Networks
Almost every network allowance on an EC2 instance scales with the instance. Bandwidth, packets per second, tracked connections — pick a bigger size and the ceiling rises. It is the reflex that makes "scale up" the first move when a network limit bites, and it usually works.
One allowance does not move. From the Amazon VPC quotas:
Each EC2 instance can send 1024 packets per second per network interface to Route 53 Resolver (specifically the .2 address, such as 10.0.0.2 and 169.254.169.253). This quota cannot be increased.
Per network interface. Not per vCPU, not per instance size, not adjustable. AWS's own instance-level network metrics post says it in one line, in a list of allowances that otherwise all grow with size: "All of these allowances get a bump as you increase instance size within the instance family, except for link local PPS."
Why the symptom has no evidence
A query that exceeds the allowance is dropped at the network interface, before it leaves the instance. That has a consequence for every tool you would normally reach for, and AWS states it directly in the troubleshooting article for exactly this failure:
VPC Flow Logs doesn't capture the traffic that applications send to Amazon DNS servers. […] Route 53 query logs capture only the traffic that reaches the VPC.2 resolver. Throttled DNS queries don't appear in query logs because the queries throttle at the network interface level.
So:
- Flow logs never record traffic to the Amazon DNS server at all, throttled or not. It is one of a short documented list of traffic flow logs exclude by design.
- Resolver query logs record queries that reached the resolver. A throttled query did not reach it, so it is not there.
- The resolver itself never saw the packet, so nothing on the AWS side registers a failure.
What the application sees is a DNS timeout, intermittently, under load, that clears on retry. What every log sees is a healthy resolver answering every query it received. The two are both true, and the gap between them is the packets that were dropped before anyone counted them.
The one place it is visible
The ENA driver counts it. On any instance with a recent driver:
ethtool -S eth0 | grep linklocal_allowance_exceeded
# linklocal_allowance_exceeded: 48213From the ENA metrics documentation, that counter is:
The number of packets dropped because the PPS of the traffic to local proxy services exceeded the maximum for the network interface. This impacts traffic to the Amazon DNS service, the Instance Metadata Service, and the Amazon Time Sync Service, but does not impact traffic to custom DNS resolvers.
A non-zero, rising linklocal_allowance_exceeded during the failures is the diagnosis. The CloudWatch agent can export it, which is the only way to alarm on it. Without that export, the number exists on the instance and nowhere else.
Two things in that definition deserve a second look. The allowance is shared by three services — DNS, the instance metadata service, and time sync — so an application that polls 169.254.169.254 aggressively is spending the same budget that its DNS lookups need. And the exemption for custom resolvers is the clue to the fix: the limit is on the link-local path to AWS's proxy services, not on DNS as such.
Why scaling up does nothing
The same AWS post that lists the ENA counters gives the advice for the other allowances — "go up an instance family (for example, a c5n.18xlarge instead of c5n.9xlarge)" — and then exempts this one. The 1,024 is a property of the interface, and a bigger instance has the same interface with the same number.
Which means the instinctive response makes the problem worse in the common case. The workload that generates the most DNS queries per second is the one with the most concurrency — more worker threads, more connections, more lookups — and the usual way to get more concurrency is a bigger instance. Scaling up raises the query rate and leaves the ceiling where it was.
What actually fixes it
AWS's recommendation in the partial DNS failure article is two options, and the first is the whole answer for most workloads:
Activate DNS caching on the instance. Increase the DNS retry timer on the application.
A local cache on the instance. The 1,024 counts packets to the resolver, not DNS lookups by the application. A cache in front of the resolver — systemd-resolved, nscd, dnsmasq, or the resolver built into the runtime — answers repeat lookups locally and sends the resolver only the misses. A fleet that looks up the same few hundred names turns tens of thousands of queries per second into tens, and the allowance stops being reachable.
This is also the explanation for why the problem appears at all. Linux does not cache DNS by default; glibc sends every getaddrinfo to the configured resolver. An application that resolves a hostname per request — common in service meshes, connection-per-call clients, and anything that disables keepalive — generates one resolver query per request, and 1,024 requests per second is not a large number.
More network interfaces. The allowance is per ENI, so a second interface is a second 1,024. This works and it is a blunt instrument: the application has to be made to spread queries across interfaces, which most cannot do. It is the right answer for a specific class of appliance and the wrong one for a web tier.
A Route 53 Resolver endpoint. The hybrid DNS whitepaper notes the different ceiling: "This limit is higher for Route 53 resolver endpoints, which have a limit of approximately 10,000 queries per second (QPS) per elastic network interface." Pointing instances at an inbound resolver endpoint's addresses instead of .2 moves the queries off the link-local path — the ENA definition said custom resolvers are exempt — and onto an interface with roughly ten times the budget. That is a bigger architectural change than a cache and it is the one that scales past a single instance's needs.
Retry tuning helps the application survive the drops. It does not reduce them. It belongs alongside a cache, not instead of one.
The audit
Any instance whose linklocal_allowance_exceeded is non-zero has hit this at some point. Across a fleet, with the CloudWatch agent exporting ENA metrics:
aws cloudwatch get-metric-statistics \
--namespace CWAgent --metric-name ethtool_linklocal_allowance_exceeded \
--dimensions Name=InstanceId,Value="$INSTANCE" \
--start-time "$(date -u -d '-24 hours' +%FT%TZ)" --end-time "$(date -u +%FT%TZ)" \
--period 3600 --statistics Maximum --query 'Datapoints[].Maximum' --output textAny non-zero value is a day on which some DNS queries were silently dropped. The instance did not log it, the resolver did not see it, and the application recorded a timeout it could not explain.
The broader lesson is about the one word in the quota that matters. Every other allowance says instance. This one says interface, and the difference is the whole failure: a limit that lives one layer below where you are looking, that scales with the one thing you cannot change by choosing a bigger box, and that drops its evidence before the evidence reaches anything that keeps it.