Your 25 Gbps instance does 5 Gbps to the internet, and that is documented
EC2 instance bandwidth is one number on the spec sheet and at least three in practice. Traffic through an internet gateway is capped at 5 Gbps on anything under 32 vCPUs, and a single flow is capped at 5 Gbps nearly everywhere. Neither is a fault, and both get filed as one.
· Venerable Networks
The spec sheet says 25 Gigabit. The transfer to S3 saturates it. The same transfer to a public endpoint on the internet stalls at 5 Gbps, and a ticket is opened against the NAT Gateway, the internet gateway, or the ISP, in that order.
None of them is at fault. The instance is doing what its documentation says it will do, on a page most people have never had a reason to read past the first paragraph.
The headline number is one of three
The bandwidth figure on the instance type page is the aggregate the instance can push to other instances in the Region. Two further limits apply depending on where the traffic is going and how many connections carry it.
| Traffic | Limit | Source |
|---|---|---|
| Many flows, to another instance in the Region | the instance's full bandwidth | spec sheet |
| Many flows, through an internet gateway | 5 Gbps under 32 vCPUs; 50% of bandwidth at 32+ | documented, below |
| One flow, anywhere not in a cluster placement group | 5 Gbps | documented, below |
The middle row is the one that produces tickets. In AWS's words:
When traffic goes through an internet gateway or a local gateway, the available bandwidth for multi-flow traffic depends on the instance type. Instance types with less than 32 vCPUs: limited to 5 Gbps. Instance types with more than 32 vCPUs: limited to 50% of the available bandwidth for the instance type.
So a c5.4xlarge — "up to 10 Gigabit", 16 vCPUs — gets 5 Gbps to the internet. A c5.18xlarge — 25 Gigabit, 72 vCPUs — gets 12.5 Gbps. Neither figure appears anywhere on the instance's own spec page. The worked example AWS gives is exactly this shape:
An
m5.16xlargeinstance has 64 vCPUs, so traffic to another instance in the Region can use the full bandwidth available (20 Gbps). However, traffic that goes through an internet gateway or a local gateway can use only 50% of the bandwidth available (10 Gbps).
The single-flow cap is the other half
The third row is independent of the second and bites a different workload. One TCP connection — one database replication stream, one scp, one backup job — is capped at 5 Gbps regardless of instance size unless both ends are in a cluster placement group:
When instances are not in the same cluster placement group, bandwidth for single-flow traffic is limited to 5 Gbps.
A flow here is a 5-tuple for TCP and UDP. For tunnel protocols it is narrower — AWS notes that for GRE or IPsec the flow is defined by the 3-tuple of source, destination and protocol — which is why a single IPsec tunnel between two instances cannot exceed 5 Gbps no matter how many connections it carries inside. Everything inside the tunnel is one flow to the network.
The documented ways past it are specific: a cluster placement group takes a single flow to 10 Gbps, ENA Express takes it to 25 Gbps between eligible instances in the same Availability Zone, and Multipath TCP spreads one logical connection across several real ones. None of these applies to the internet gateway path. For traffic leaving AWS, the two caps stack: a single flow is 5 Gbps, and all flows together are 5 Gbps on a small instance.
Why it reads as a fault
The symptom is a clean plateau. Throughput climbs and then sits at 5 Gbps with no errors, no retransmits worth noting, and no metric on the instance that says "limited". The NAT Gateway, if there is one, is well within its 100 Gbps. The internet gateway publishes nothing to blame. The ISP at the other end shows headroom.
Then someone runs the same transfer to an S3 endpoint, or to another instance, and it saturates the link. That comparison is what turns "the network is slow" into "the network is slow to the internet specifically", and from there the ticket goes to whoever owns the edge.
The instance is where the limit lives, and the instance has a counter for it. The ENA driver exposes allowance metrics on every instance, and bw_out_allowance_exceeded is the one that increments when outbound traffic is queued or dropped because it hit the instance's aggregate bandwidth allowance:
ethtool -S eth0 | grep -E 'bw_(in|out)_allowance_exceeded|pps_allowance_exceeded'
# bw_in_allowance_exceeded: 0
# bw_out_allowance_exceeded: 184733
# pps_allowance_exceeded: 0A rising bw_out_allowance_exceeded during the plateau is the limit identifying itself. Nothing downstream will.
The "up to" is also a limit
A separate trap sits on the same page and compounds this one. Instances documented as "up to" a bandwidth — generally 16 vCPUs and smaller — have a lower baseline and burst above it on a credit mechanism:
Instances can use burst bandwidth for a limited time, typically from 5 to 60 minutes, depending on the instance size.
A c5.large is "up to 10 Gigabit" with a baseline of 0.75 Gbps. A sustained transfer on one gets ten gigabits for a while and then drops to three-quarters of one, and the drop is not a fault either. The baseline is not on the console; it is in describe-instance-types:
aws ec2 describe-instance-types --instance-types c5.large c5.4xlarge c5.18xlarge \
--query 'InstanceTypes[].[InstanceType,NetworkInfo.NetworkPerformance,NetworkInfo.NetworkCards[0].BaselineBandwidthInGbps]' \
--output table
# | c5.large | Up to 10 Gigabit | 0.75 |
# | c5.4xlarge | Up to 10 Gigabit | 5.0 |
# | c5.18xlarge | 25 Gigabit | 25.0 |So the real question for an internet-bound workload on a c5.large is not "why is it slower than 10 Gbps" but "which of three limits is it at": the 0.75 Gbps baseline after burst credits run out, the 5 Gbps internet gateway cap during burst, or the 5 Gbps single-flow cap if it is one connection. The answer changes the fix.
What to do with it
Size for the path, not the spec sheet. If a workload's job is to move data to the internet, the number that matters is the internet gateway figure: 5 Gbps below 32 vCPUs, half the spec above. An instance type that looks like overkill for the compute may be the smallest one that reaches the throughput, and a 32-vCPU floor is a real design constraint for high-egress nodes.
Parallelise anything that is one flow. A single stream is 5 Gbps nearly everywhere. Tools that open multiple connections — aws s3 cp does by default, most backup agents can be configured to — sidestep the single-flow cap on the Region path entirely. They do not sidestep the internet gateway cap, which is aggregate.
Read the ENA counters before the ticket. bw_out_allowance_exceeded incrementing is a definitive answer and takes ten seconds to check. A ticket against the NAT Gateway or the ISP with that counter at zero is at least looking in a plausible place; with it climbing, the ticket is against the instance type.
Keep the baseline in the capacity model. "Up to" instances are honest about being burstable, but the baseline is the number that holds under sustained load, and it is not where anyone looks for it.
The pattern is the one that runs through the other limits that changed shape: a number that reads as a single figure is actually several, scoped by destination and by flow, and the scoping is documented in a place the symptom does not point at. The spec sheet is not wrong. It is just answering a narrower question than the one you are asking it.