Lab 20: The Egress VPC That Routed Around Its Own NAT Gateway
Centralised egress, built to the reference architecture. The NAT Gateway is available, both attachments are available, every route is active, and the egress VPC itself reaches the internet fine. The spokes get nothing, and no log anywhere records a drop.
- Debugging time
- ~30 min
- Reading time
- 15 min
- Reported by
- Platform Engineering
- Tier
- Professional
All spoke VPCs lost internet access after cutover to the central egress VPC; egress VPC itself is fine
Reported by Platform Engineering
We cut the first spoke over to the new central egress VPC this afternoon. Its default route now goes to the transit gateway instead of its own NAT Gateway, which we deleted as planned. Since then, outbound from the spoke times out. Package installs, API calls, everything.
I have checked the obvious things. Both transit gateway attachments are available. The spoke's
route table sends 0.0.0.0/0 to the transit gateway. The transit gateway route table sends
0.0.0.0/0 to the egress attachment. The egress VPC's NAT Gateway is available with an Elastic IP.
I can log into a host in the egress VPC and curl the internet without any problem, so the egress VPC
can clearly get out.
I enabled flow logs on the egress side to see where it is being blocked and there is not a single
REJECT. Every record is ACCEPT. So nothing is blocking it, and it still does not work.
What you are working with
One Transit Gateway, two VPCs, one Availability Zone.
The spoke has no internet gateway and no NAT Gateway of its own any more. Its default route reaches the transit gateway, which delivers it to the egress VPC's attachment subnet. What that subnet's route table does with it is the one thing this diagram does not state, because it is the thing you are going to read. The control host sits in the public subnet with its own public address, which is why it can reach the internet and why that proves less than it seems to.
| Resource | Configuration |
|---|---|
| Workload | Spoke VPC, 10.200.1.50, private subnet, default route → transit gateway |
| Control host | Egress VPC public subnet, 10.201.1.50, has a public IPv4 address |
| Transit gateway | Spokes route table: 0.0.0.0/0 → egress attachment. Egress table: spoke prefix propagated |
| NAT Gateway | Egress VPC public subnet, available, Elastic IP attached |
| Public subnet route table | 0.0.0.0/0 → internet gateway; 10.200.0.0/16 → transit gateway |
| Attachment subnet route table | 0.0.0.0/0 → see Reproduction |
| Flow logs | Enabled on the egress VPC's attachment subnet, all traffic |
Every object is available. Every route is active. There is no security group or network ACL in the
path that denies anything.
- EC2
Workload
10.200.1.50 → 1.1.1.1:443
- GW
Transit gateway
spokes route table: 0.0.0.0/0 → egress attachment
- SUBNET
Egress attachment subnet
flow log: ACCEPT, src 10.200.1.50
- RTB
Attachment subnet route table
0.0.0.0/0 → the target it was given
Dropped — The packet is handed to a target that is valid for the route table and accepts it. The reply never comes. Nothing records why.
- DEST
1.1.1.1
never sees a SYN
The packet travels three hops correctly and is accepted into the egress VPC. The fourth hop is where it goes somewhere and does not come back, and the absence of a REJECT is the clue rather than the reassurance the reporter took it for.
Scope and constraints
- In scope: why the spoke's traffic reaches the egress VPC and never reaches the internet, when the egress VPC itself can.
- Out of scope: the transit gateway configuration, the spoke's route table, the NAT Gateway's health, security groups, and network ACLs. All correct, and the reproduction confirms each.
- The reporter is right that nothing is rejecting the traffic. Hold on to that: it is a real observation and it narrows the search considerably.
- The reporter's test from the egress VPC is also genuinely true and genuinely misleading. Work out why before you trust it.
- The fix is a one-line change to one route.
Deploy the broken state
cd lab-20-egress-vpc-igw-drop
terraform init
terraform applyThe transit gateway takes a few minutes and the attachments a few more — budget about eight minutes.
The spoke host starts a probe on boot that attempts one connection a minute to 1.1.1.1:443, so the
egress VPC's flow logs will have evidence waiting by the time you look.
SPOKE=$(terraform output -raw spoke_instance_id)
TGW_RTB=$(terraform output -raw tgw_subnet_route_table_id)
PUB_RTB=$(terraform output -raw public_subnet_route_table_id)
NATGW=$(terraform output -raw nat_gateway_id)
IGW=$(terraform output -raw internet_gateway_id)
FLOWS=$(terraform output -raw flow_log_group)Confirm everything the ticket claims
All of it is true. Confirm it so you stop re-checking it.
aws ec2 describe-transit-gateway-vpc-attachments \
--filters "Name=tag:Lab,Values=20-egress-vpc-igw-drop" \
--query 'TransitGatewayVpcAttachments[].[Tags[?Key==`Name`]|[0].Value,State]' --output table
# vn-lab-20-attach-spoke available
# vn-lab-20-attach-egress available
aws ec2 describe-nat-gateways --nat-gateway-ids "$NATGW" \
--query 'NatGateways[0].[State,NatGatewayAddresses[0].PublicIp]' --output text
# available 54.xx.xx.xxThen the reporter's reassuring test. Open a session on the control host —
terraform output -raw control_session_command — and reach the internet:
curl -s -m 5 https://checkip.amazonaws.com
# 3.xx.xx.xxIt works. Note the address it returns, because it is going to matter: it is not the NAT Gateway's Elastic IP. The control host has its own public address and is using it. This test proved that a host with a public address in a public subnet can reach the internet, which was never in question.
Read the flow logs the reporter read
aws logs filter-log-events --log-group-name "$FLOWS" \
--filter-pattern '"10.200.1.50" "1.1.1.1"' \
--query 'events[-6:].message' --output text
# 2 123456789012 eni-0ab… 10.200.1.50 1.1.1.1 41622 443 6 1 60 1727540400 1727540459 ACCEPT OK
# 2 123456789012 eni-0ab… 10.200.1.50 1.1.1.1 41634 443 6 1 60 1727540460 1727540519 ACCEPT OK
# 2 123456789012 eni-0ab… 10.200.1.50 1.1.1.1 41650 443 6 1 60 1727540520 1727540579 ACCEPT OKThe reporter was right. Every record is ACCEPT. Now read the records more carefully than they did.
Each line is one packet, 60 bytes — a lone SYN — from the spoke to 1.1.1.1, accepted into the
attachment subnet's network interface. There is no corresponding line in the other direction. Nothing
from 1.1.1.1 ever arrives for 10.200.1.50. The SYNs go in, are accepted, and the conversation ends
there.
ACCEPT in a flow log means the packet was permitted by the security group and network ACL. It says
nothing about what the route table did with it afterwards. So the evidence says: the traffic arrives in
the egress VPC, nothing filters it, and it is sent somewhere that never answers.
Follow the route
The attachment subnet's route table decides where the accepted packet goes next:
aws ec2 describe-route-tables --route-table-ids "$TGW_RTB" \
--query 'RouteTables[0].Routes[].[DestinationCidrBlock,GatewayId,NatGatewayId,TransitGatewayId,State]' \
--output table
# ---------------------------------------------------------------
# | 10.201.0.0/16 | local | None | None | active |
# | 0.0.0.0/0 | igw-0c4e… | None | None | active |
# ---------------------------------------------------------------The default route targets the internet gateway. Compare with the public subnet, where the NAT Gateway lives:
aws ec2 describe-route-tables --route-table-ids "$PUB_RTB" \
--query 'RouteTables[0].Routes[].[DestinationCidrBlock,GatewayId,NatGatewayId,TransitGatewayId,State]' \
--output table
# ---------------------------------------------------------------
# | 10.201.0.0/16 | local | None | None | active |
# | 0.0.0.0/0 | igw-0c4e… | None | None | active |
# | 10.200.0.0/16 | None | None | tgw-0f2a… | active |
# ---------------------------------------------------------------Same default route, same target. The two tables are nearly identical, and that is the tell: the attachment subnet's table looks like a copy of the public subnet's. The public subnet is supposed to route to the internet gateway — that is where the NAT Gateway's translated traffic goes out. The attachment subnet is supposed to route to the NAT Gateway, and nothing in this VPC does.
EGRESS_VPC=$(aws ec2 describe-nat-gateways --nat-gateway-ids "$NATGW" \
--query 'NatGateways[0].VpcId' --output text)
aws ec2 describe-route-tables --filters "Name=vpc-id,Values=$EGRESS_VPC" \
--query 'RouteTables[].Routes[?NatGatewayId!=`null`][]' --output json
# []The NAT Gateway is available, has an address, and is the target of nothing. Spoke traffic is being
sent straight past it to the internet gateway.
This debrief is part of Labs Pro
The root-cause analysis, packet-flow walkthrough, and Terraform remediation diff for this lab are available to Labs Pro members. One payment of $49, no subscription, and it covers every Pro lab now and later.
The brief, the reproduction steps, and the Terraform stay free — you can still solve this one yourself.
Already bought it? Sign in and it unlocks.