The Site-to-Site VPN tunnel that waits to be called
By default, AWS never initiates a Site-to-Site VPN tunnel. Your device has to. So a tunnel that only carries traffic AWS originates — monitoring, a scheduled job, a backup pull — comes up, idles, drops, and stays down until something on your side sends a packet. The fix is one setting that most people have never seen, and it requires IKEv2.
· Venerable Networks
A Site-to-Site VPN has two tunnels, both show UP, the runbook is written, and the connection is handed over. Three weeks later a scheduled job in AWS that pulls from an on-premises system fails overnight with a timeout. In the morning somebody logs in, pings the on-premises host from an instance, the ping succeeds — and the job runs fine when re-tried. It happens again the following week.
The tunnel was down when the job ran and up when the human looked. The human brought it up by looking.
AWS is the responder
From the tunnel initiation options:
By default, your customer gateway device must bring up the tunnels for your Site-to-Site VPN connection by generating traffic and initiating the Internet Key Exchange (IKE) negotiation process.
AWS's side does not initiate. It waits for your device to start the IKE exchange, and it answers. That is a reasonable default for the common case — on-premises users reaching AWS — because the first packet comes from the side that initiates. It is the wrong default for every flow that starts in AWS.
The same page describes what happens when nothing is flowing:
If you do not configure IKE initiation from the AWS side for your VPN tunnel and the VPN connection experiences a period of idle time (usually 10 seconds, depending on your configuration), the tunnel might go down.
Idle, then down, then waiting for your device to notice it needs to come back up. A customer gateway with no traffic to send has no reason to initiate, so the tunnel stays down until something on-premises wants AWS — or until a human on the AWS side generates traffic that reaches the on-premises device and provokes it to negotiate.
That is the ping that fixed the job. It was not a test. It was the repair.
The second default makes it stickier
Dead peer detection is how each side notices the other has gone. AWS sends a DPD message every 10 seconds and declares the peer dead after three go unanswered. What happens next is the second setting on the same page:
DPD timeout action: The action to take after dead peer detection (DPD) timeout occurs. By default, the IKE session is stopped, the tunnel goes down, and the routes are removed.
Routes removed. So a brief blip on the customer gateway — a reboot, a flap on the ISP link — tears down the IKE session on the AWS side, withdraws the routes, and then AWS waits, because it is the responder. If your device does not re-initiate on its own, the tunnel is down until something makes it try.
Both defaults point the same direction: AWS does not start tunnels and does not restart them.
The settings that change it
Two tunnel options, settable per tunnel on a new or modified connection:
Startup action: start. AWS initiates IKE for a new or modified connection, rather than waiting:
You can specify that AWS must initiate the IKE negotiation process instead.
And it does not give up:
If using IKE initiation from the AWS side of the VPN connection, it does not include a timeout setting. It will continuously try to establish a connection until one is made. Additionally, the AWS side of VPN connection will re-initiate IKE negotiation when it receives a delete SA message from your customer gateway.
DPD timeout action: restart. Instead of tearing down and waiting, AWS restarts the IKE session when DPD fires. Or none, to leave the session alone and let your own device's DPD decide.
In Terraform, per tunnel:
resource "aws_vpn_connection" "main" {
# ...
tunnel1_startup_action = "start"
tunnel1_dpd_timeout_action = "restart"
tunnel2_startup_action = "start"
tunnel2_dpd_timeout_action = "restart"
}With both set, a tunnel that drops is brought back by AWS, regardless of whether anything on-premises wants to talk. The scheduled job finds the tunnel up because AWS kept it up.
Three conditions on start
The rules and limitations list on that page is short and every item has bitten someone.
It is IKEv2 only.
IKE initiation (startup action) from the AWS side of the VPN connection is supported for IKEv2 only.
A connection negotiated with IKEv1 — still common on older customer gateway configurations — cannot use start. Changing IKE version is a change on both ends and usually a maintenance window.
AWS needs your device's public address.
To initiate IKE negotiation, AWS requires the public IP address of your customer gateway device. If you configured certificate-based authentication for your VPN connection and you did not specify an IP address when you created the customer gateway resource in AWS, you must create a new customer gateway and specify the IP address.
Certificate-authenticated connections can be created with a customer gateway that has no IP address — the certificate identifies the peer, so none is needed for AWS to respond. To initiate, AWS has to know where to send the first packet. A customer gateway's IP address cannot be edited after creation, so this means a new customer gateway and a modification to the connection.
A device behind NAT needs an identity.
If your customer gateway device is behind a firewall or other device using Network Address Translation (NAT), it must have an identity (IDr) configured.
When AWS initiates toward a NATed device, the address it sends to is the NAT's, and the responder has to identify itself by something other than that address. That is an IKEv2 IDr on the customer gateway, and it is a configuration most devices do not set by default.
The alternative, and why it is worse
The same page offers the other option:
To prevent this, you can use a network monitoring tool to generate keepalive pings.
An instance in the VPC pinging an on-premises address every few seconds keeps the tunnel from idling and, when it does drop, generates the traffic that provokes the customer gateway to renegotiate. It works. It also means the availability of your hybrid link depends on a cron job on an instance someone has to remember exists, and it does nothing about DPD teardown if the on-premises side is the one that went quiet.
start and restart are the same guarantee from the side that is actually managed. Use the keepalive only where IKEv1 or a missing customer gateway address rules them out, and treat it as the debt it is.
What to check
Whether your connection has either setting:
aws ec2 describe-vpn-connections --vpn-connection-ids "$VPN" \
--query 'VpnConnections[0].Options.TunnelOptions[].[OutsideIpAddress,StartupAction,DpdTimeoutAction,IkeVersions[].Value|join(`,`,@)]' \
--output table
# -----------------------------------------------------
# | 3.xx.xx.xx | add | clear | ikev1,ikev2 |
# | 52.xx.xx.xx | add | clear | ikev1,ikev2 |
# -----------------------------------------------------add is the default startup action — wait for the customer gateway. clear is the default DPD action — tear down and remove routes. A connection showing both defaults, carrying any flow that originates in AWS, is the scheduled job waiting to fail.
The behaviour the tunnel is supposed to have — up, and coming back up on its own — is described in the hybrid connectivity guide as the thing redundancy depends on. Two tunnels that both wait to be called are not two paths. They are two places to wait.