Skip to content
Networking7 min read

A regional NAT Gateway makes the most common NAT mistake impossible to express

The zonal NAT Gateway has a placement rule, a per-AZ route table requirement, and a single-zone dependency, and each one is a documented outage. Regional mode does not fix those bugs. It removes the configuration they live in.

· Venerable Networks

The oldest NAT Gateway failure on this site is a gateway placed in the private subnet it serves. The apply succeeds, the gateway reports available, and every packet loops back into the route table that just sent it there. It is Lab 01 because it is the thing people actually do.

That failure has a precondition most explanations skip past: a zonal NAT Gateway has to be placed in a subnet, and the subnet has to be one whose route table reaches an internet gateway. The mistake is possible because the choice exists.

A regional NAT Gateway removes the choice. From the AWS documentation:

A regional NAT Gateway is a standalone resource with its own route table and you do not need a public subnet in your VPC to host a regional NAT Gateway, which reduces chances of misconfiguring private resources in subnets with public connectivity.

You create it against a VPC, not a subnet. There is no subnet to pick wrongly.

Three bugs, one configuration surface

The zonal model has three well-known failure modes, and each one lives in a specific piece of configuration that regional mode does not have.

Zonal failureWhere it livesWhat regional mode does with it
Gateway in the wrong subnet, routing loopthe subnet_id you choosethere is no subnet; the gateway is created against the VPC
Private subnets in AZ b and c routing to the gateway in AZ a, paying cross-AZ and depending on one zonea route table shared across AZs, or three route tables that driftedone NAT ID is valid from every zone, so a shared route table is correct
Scale into a new AZ, forget to add a gateway and a route, that zone has no egressthe step you repeat per AZthe gateway expands into the zone by itself

AWS's own description of the zonal workflow is worth quoting because it lists the chores that each become a place to get something wrong:

You first create zonal NAT Gateways per Availability Zone and host your NATs in public subnets. You then configure separate routes per Availability Zone from your private subnets to the NAT in that Availability Zone. You repeat this step every time your workloads expand to a new Availability Zone, for high availability. Additionally, you need to add routes for the internet gateway in the route table of your NAT subnet per Availability Zone.

Four actions, repeated per zone, for the lifetime of the VPC. Regional mode collapses them to one.

The route table nobody has to write

The mechanism that makes the public subnet unnecessary is that the regional gateway brings its own route table:

Once you create your regional NAT Gateway, AWS automatically creates a route table for it, which comes with a pre-configured route to the internet gateway.

This is the piece Lab 01 exists to teach. A zonal NAT Gateway needs two route tables to cooperate — the private subnet's, pointing at the gateway, and the gateway's subnet's, pointing at the internet gateway — and the loop happens when they turn out to be the same table. A regional gateway owns the second one. You cannot associate it with the wrong subnet because it is not associated with a subnet at all.

The same route table is where you add return routes for middleboxes, and it accepts a Transit Gateway as a target, so a regional gateway slots into a centralised egress design without a special case.

Expansion follows network interfaces, with a delay

The high-availability story is not a replica in every zone standing by. The gateway watches for workloads:

When you launch resources in a new Availability Zone, the regional NAT gateway detects the presence of an network interface (ENI) in that Availability Zone and automatically expands to that zone. Similarly, the NAT Gateway contracts from the Availability Zone that has no active workloads.

That design choice has a consequence worth planning around:

It may take your regional NAT Gateway up to 60 minutes to expand to a new Availability Zone after a resource is instantiated there. Until this expansion is complete, the relevant traffic from this resource is processed across zones by your regional NAT Gateway in one of the existing Availability Zones.

So a scale-out into a cold zone works immediately and pays cross-AZ data processing for up to an hour. That is a cost blip rather than an outage, and it is the opposite trade from zonal mode, where a cold zone with no gateway has no egress at all until somebody adds one. Which failure you would rather have during an incident is not a hard question.

Expansion is also what you are choosing between in the two modes. Automatic mode, which AWS recommends, manages addresses and zone expansion for you. Manual mode gives you control over addresses per zone and makes you responsible for the expansion — which reintroduces the per-zone chore the feature exists to remove, so it is for the case where you need specific addresses in specific zones and not a default.

The port headroom is the better argument

The availability case gets the attention. The capacity case is stronger for a lot of workloads.

A zonal gateway supports up to 8 IPv4 addresses. A regional one:

Your regional NAT Gateways support up to 32 IP addresses per Availability Zone (compared to 8 for zonal NAT gateways). Each IP address increases the limit on concurrent connections to a popular destination (identified by unique combination of destination IP, destination port and protocol) by 55,000.

If you have ever hit ErrorPortAllocation against a single busy destination — one third-party API, one managed database endpoint — the lever is more addresses on the gateway, not more gateways, and regional mode quadruples the ceiling per zone. The 55,000 figure is per address per destination, which is why this matters and why adding a second NAT Gateway for "connection headroom" never helped.

What it does not do

Two constraints decide whether this applies to you, and they are stated plainly:

Regional NAT gateways do not support private NAT. If you need private NAT, use zonal NAT gateways instead. Regional NAT gateways are not supported in constrained Availability Zones.

Private NAT — translating between overlapping ranges, or presenting a single routable address to on-premises — is zonal only. So an estate that uses private NAT keeps zonal gateways for that, and keeps every one of the three failure modes above for them. The guidance is otherwise unambiguous: consider regional gateways "for all use cases except those that require private connectivity."

And it is not cheaper. The hourly rate is still charged per Availability Zone the gateway is active in, so collapsing three zonal gateways into one regional gateway saves operational surface and route-table maintenance rather than money. The cost lever in NAT is still centralised egress, and the two are compatible.

Migrating resets connections

If you are converting an existing VPC, read the warning before the steps:

This will reset your existing connections. We recommend that you complete these steps in your maintenance window.

There are two documented paths and they trade the same thing. Create the regional gateway first, repoint the routes, delete the zonal ones — new addresses, connections reset at the route change. Or delete the zonal gateways first to release their Elastic IPs, create the regional gateway with those same addresses, then repoint — addresses preserved, but there is a gap with no NAT at all while you do it. If downstream allow-lists depend on your egress addresses, the second path is the one you need, and it is the one that needs a real maintenance window rather than a quiet moment.

Whichever path, every long-lived connection through the old gateways — database pools, message consumers, anything that holds a socket — will drop and reconnect. That is the same failure shape as the 350-second idle reset, arriving all at once instead of per connection, and the applications that handle one well handle the other.

The audit

A zonal gateway in a VPC that has no private-NAT requirement is a candidate. Finding them is one call:

aws ec2 describe-nat-gateways \
  --filter "Name=state,Values=available" \
  --query 'NatGateways[].[NatGatewayId,VpcId,SubnetId,ConnectivityType,AvailabilityMode]' \
  --output table

Rows with ConnectivityType: public and an AvailabilityMode other than regional are the ones carrying the three bugs for no reason. Rows with ConnectivityType: private are the ones that have to stay.

The general point is the one worth carrying. Most of this site's NAT Gateway content is about mistakes that are easy to make because the configuration invites them. The strongest fix for a class of misconfiguration is rarely a check that catches it. It is a resource that has nowhere to put it.