Ten kinds of traffic VPC Flow Logs never record
A flow log with no REJECT lines is routinely read as proof that nothing was blocked. AWS documents ten categories of traffic that flow logs exclude by design, and several of them are precisely the traffic you would be looking for when something is wrong.
· Venerable Networks
Flow logs are the first thing anyone enables when a VPC misbehaves, and the first thing anyone reads. The reading is usually "look for REJECT", and the conclusion when there is none is "nothing is being blocked, so the problem is somewhere else".
Both halves of that are weaker than they look. A REJECT is a security group or network ACL decision and nothing else — Lab 20 is built on an outage that produced nothing but ACCEPT. And before you get to interpreting what is in the log, there is the question of what was never going to be in it.
AWS maintains the list. From Flow log limitations:
Flow logs do not capture all IP traffic. The following types of traffic are not logged:
Ten items follow. Here they are, grouped by why each one matters.
The three that hide a diagnosis
Traffic to the Amazon DNS server. The first item on AWS's list, and the one with the sharpest consequence:
Traffic generated by instances when they contact the Amazon DNS server. If you use your own DNS server, then all traffic to that DNS server is logged.
Every query to the .2 resolver is invisible. So is every failed query — including the ones dropped by the 1,024 packets-per-second allowance, which never reach the resolver's own query log either. A DNS problem leaves no trace in flow logs by construction, and the absence of DNS traffic in a flow log is not evidence that an instance made no queries.
Traffic to and from 169.254.169.254 for instance metadata, and 169.254.169.123 for Amazon Time Sync. Two separate items on the list, one consequence: an instance hammering the metadata service — a credential-refresh loop, a misconfigured SDK, an agent polling for tags — generates traffic that shares the link-local packet allowance with DNS and shows up nowhere in flow logs. When DNS starts failing because something else is spending the allowance, flow logs cannot show you either half.
The three that hide a design
Traffic between an endpoint network interface and a Network Load Balancer network interface. This is PrivateLink's own hop. When a consumer reaches a service through an interface endpoint, the leg from the endpoint ENI to the provider's NLB is not logged — on either side. Flow logs on the consumer VPC stop at the endpoint ENI; flow logs on the provider VPC start at the NLB's targets. The hop between them belongs to AWS, and a connectivity problem on it is invisible to both parties' logs. The source address on that path is always the NLB, which is the other reason PrivateLink troubleshooting from flow logs is harder than it should be.
Traffic mirrored source traffic. Traffic Mirroring copies packets from a source ENI to a target. The copies appear in flow logs at the target. The mirroring itself — the fact that a source's traffic is being duplicated and sent elsewhere — does not appear at the source. A flow log cannot tell you that an interface is being mirrored. That is worth knowing from the security side as much as the operations side.
Traffic to the reserved IP address for the default VPC router. The .1 address in every subnet. An instance's traffic through the router is logged as it leaves the ENI; traffic to the router itself is not. In practice this rarely matters, but it means a probe against the gateway address is not something flow logs will confirm happened.
The four that are housekeeping
Windows license activation, DHCP, ARP, and — the newest item — traffic on a short-lived regional NAT gateway, which AWS describes as one "deleted a few minutes after creation". These are protocol chatter and transient resources; their absence rarely misleads. They are on the list because the list is complete, and a complete list is the point.
What this does to "the flow log is clean"
A clean flow log proves that the traffic it recorded was not rejected by a security group or network ACL. It proves nothing about:
| You are looking for | Flow logs show |
|---|---|
| DNS failures | nothing — DNS to the Amazon resolver is excluded |
| Metadata service problems | nothing — 169.254.169.254 is excluded |
| A PrivateLink connectivity fault | each side's ENI, never the hop between them |
| Whether an interface is being mirrored | nothing at the source |
| Routing failures | ACCEPT records with no reply — routing is not a filter decision |
The last row is not an exclusion; it is the thing ACCEPT actually means. But it belongs in the same table, because it is the other way a clean flow log gets read as a clean network.
The other limitations worth knowing
The same page carries a second list, of behaviours rather than exclusions, and three of them regularly cost people an afternoon:
You cannot change a flow log after creating it. Not the role, not the format, not the fields. To add pkt-srcaddr you delete the flow log and create another.
The default srcaddr and dstaddr are the ENI's primary address, not the packet's. Traffic to a secondary IP, or through an intermediate device, shows the interface's primary address in the default fields. The packet's real addresses are in pkt-srcaddr and pkt-dstaddr, which are not in the default format — so a flow log created with defaults, in front of a NAT instance or a NLB, is recording the wrong addresses and cannot be fixed without recreating it.
Records can be skipped. AWS states that some records "may be skipped during the aggregation interval" due to capacity constraints, and that the log-status field says so. A flow log is not a packet capture, and a gap in it is not evidence of a gap in the traffic.
Reading flow logs honestly
Three habits, in order of how often they would have shortened an incident.
Before concluding from absence, check the exclusion list. If the traffic you are looking for is DNS, metadata, time sync, PrivateLink's inner hop, or anything on that list, a flow log cannot show it to you and its silence means nothing. Reach for the tool that can: linklocal_allowance_exceeded for the link-local services, Resolver query logs for DNS that reached the resolver, a packet capture for the rest.
Read ACCEPT as "the filters let it through", not "it worked". Then read the byte count and the direction. Accepted packets with no return traffic are a routing or destination problem that the flow log has told you about by not rejecting anything.
Create flow logs with the fields you will need, because you cannot add them later. pkt-srcaddr, pkt-dstaddr, traffic-path, and log-status are the four most often missing when they are needed. The cost of including them is nothing; the cost of not having them is a new flow log and a wait for traffic.
The through-line is the same as the other failures that present as success: a log that is honest about what it recorded, read as if it recorded everything. Flow logs are a good tool. They are a good tool with a published list of what they do not do, and the list is short enough to know.