Lab 19: The Firewall Rule That Only Guarded Itself
A central Network Firewall with a rule that drops plaintext HTTP. Security tested it and signed it off. From every spoke VPC the firewall exists to protect, HTTP goes straight through — and the firewall logged every one of those flows as allowed.
- Debugging time
- ~45 min
- Reading time
- 16 min
- Reported by
- Cloud Security
- Tier
- Professional
Plaintext HTTP from spoke VPCs still reaches the internet after the egress block was signed off
Reported by Cloud Security
We rolled out the plaintext-HTTP egress block on the central firewall last week. The rule is a
standard Suricata drop: $HOME_NET to $EXTERNAL_NET on port 80. I tested it myself from our test
host before sign-off — curl http://… times out, HTTPS works, exactly as designed. The change ticket is
closed.
This morning a platform engineer showed me a spoke workload fetching a package index over plain HTTP.
It works first time, every time. I checked the firewall: state READY, policy attached, rule group
referenced, my rule is in it, logging is on. Nothing has changed since I tested it.
I re-ran my test from the test host. Still blocked. So the firewall is enforcing the rule for me and not for them, and I cannot see what is different about their traffic.
What you are working with
One Transit Gateway, two VPCs, one Availability Zone, one firewall endpoint.
The spoke has no internet gateway of its own. Its only way out is the transit gateway, which delivers everything to the inspection VPC, where the attachment subnet hands it to the firewall endpoint and the firewall hands it to the NAT Gateway. The control host in the inspection VPC takes the same path from one hop later. Both hosts' HTTP crosses the same firewall endpoint and meets the same rule.
| Resource | Configuration |
|---|---|
| Workload | Spoke VPC, 10.190.1.50, private subnet, default route → transit gateway |
| Control host | Inspection VPC, 10.192.3.50, private subnet, default route → firewall endpoint |
| Transit gateway | Spoke route table: 0.0.0.0/0 → inspection attachment. Inspection table: spoke prefix propagated |
| Firewall | READY, one endpoint, one policy |
| Firewall policy | Strict order, one stateful rule group, default action alert_established |
| Rule group | drop tcp $HOME_NET any -> $EXTERNAL_NET 80 (flow:to_server; …) |
| Logging | Alert and flow logs to CloudWatch Logs |
| NAT Gateway | Inspection VPC public subnet; returns to the spoke and to the control subnet go back via the firewall |
Both hosts reach the internet only through the firewall endpoint. The rule in the firewall is the one the ticket describes, and it is the only rule.
- EC2
Workload
10.190.1.50 → :80
- GW
Transit gateway
spokes route table: 0.0.0.0/0 → inspection attachment
- SUBNET
Attachment subnet
0.0.0.0/0 → firewall endpoint
- FILTER
Network Firewall
stateful engine: one drop rule evaluated, no match, default action applied
- GW
NAT Gateway → internet gateway
source rewritten to the Elastic IP
- DEST
checkip.amazonaws.com
answers with the NAT address
Six hops and six passes. The fourth is the one that was supposed to fail. The firewall evaluated the rule, decided it did not apply, passed the flow under its default action, and wrote a log line saying so.
Scope and constraints
- In scope: why the firewall drops HTTP from the control host and passes it from the spoke, when both flows cross the same endpoint and meet the same rule.
- Out of scope: the transit gateway routing, the VPC route tables, the NAT Gateway, and the firewall's health. All correct, and the reproduction confirms each of them.
- The reporter is right about everything they checked. The firewall is
READY, the policy is attached, the rule is present and correctly written, and it genuinely blocks the test host. Confirm these early so you stop re-checking them. - Nothing is broken. Both hosts reach the internet. One of them is permitted to do something it should not be, and nothing alarms, because the firewall considers the outcome correct.
- The fix adds a small amount of configuration to the policy and changes no rule.
Deploy the broken state
cd lab-19-network-firewall-home-net
terraform init
terraform applyThe firewall takes several minutes to reach READY, and the four route tables that target its endpoint
cannot be created until it does, so budget ten to twelve minutes before anything is testable. Session
Manager reaches both hosts through the firewall on 443, which the policy passes.
SPOKE=$(terraform output -raw spoke_instance_id)
CONTROL=$(terraform output -raw control_instance_id)
FW=$(terraform output -raw firewall_arn)
POLICY=$(terraform output -raw firewall_policy_arn)
RG=$(terraform output -raw rule_group_arn)
ALERTS=$(terraform output -raw alert_log_group)Confirm the firewall is healthy and the rule is in place
Do this first, because it is what the ticket claims and all of it is true.
aws network-firewall describe-firewall --firewall-arn "$FW" \
--query 'FirewallStatus.[Status,ConfigurationSyncStateSummary]' --output text
# READY IN_SYNC
aws network-firewall describe-firewall-policy --firewall-policy-arn "$POLICY" \
--query 'FirewallPolicy.[StatefulEngineOptions.RuleOrder, StatefulDefaultActions, StatefulRuleGroupReferences[].ResourceArn]' \
--output json
# [ "STRICT_ORDER", [ "aws:alert_established" ], [ "arn:aws:network-firewall:…:stateful-rulegroup/vn-lab-19-egress-controls" ] ]
aws network-firewall describe-rule-group --rule-group-arn "$RG" \
--query 'RuleGroup.RulesSource.RulesString' --output text
# drop tcp $HOME_NET any -> $EXTERNAL_NET 80 (msg:"Lab 19 - plaintext HTTP egress is not permitted"; flow:to_server; sid:1900001; rev:1;)Firewall ready and in sync. Strict order. The rule group is referenced. The rule says what the ticket says it says. Every box ticked.
Reproduce from the spoke
Open a session on the spoke workload — terraform output -raw start_session_command — and try the
thing that should be blocked:
curl -s -m 5 http://checkip.amazonaws.com
# 54.87.xxx.xxx
curl -s -m 5 -o /dev/null -w '%{http_code}\n' https://checkip.amazonaws.com
# 200Plaintext HTTP returns an address — the NAT Gateway's Elastic IP, which also confirms the path: the request left through the inspection VPC. HTTPS works too. From the spoke, the firewall is a transparent hop.
Reproduce from the control host
Open a second session on the control host — terraform output -raw control_session_command — and run
the same two commands:
curl -s -m 5 http://checkip.amazonaws.com
# (nothing, then exit code 28 — timed out)
curl -s -m 5 -o /dev/null -w '%{http_code}\n' https://checkip.amazonaws.com
# 200HTTP times out. HTTPS works. From the control host the rule is doing exactly what it was written to do.
This is the whole symptom in two shells. Same firewall endpoint, same policy, same rule, same destination, same port. One source is dropped and one is passed, and the only variable is which VPC the source address belongs to.
Read the firewall's own account
The firewall logged both decisions. Pull the alert log for port 80 and compare:
aws logs filter-log-events --log-group-name "$ALERTS" \
--filter-pattern '{ $.event.dest_port = 80 }' \
--query 'events[].message' --output text \
| jq -c '{src: .event.src_ip, action: .event.alert.action, sig: .event.alert.signature}'
# {"src":"10.192.3.50","action":"blocked","sig":"Lab 19 - plaintext HTTP egress is not permitted"}
# {"src":"10.190.1.50","action":"allowed","sig":"aws:alert_established action"}Two flows to port 80. The control host's matched the drop rule and was blocked. The spoke's matched
no rule at all — it fell through to the policy's default action, alert_established, which passed
it and wrote this line.
So the engine did not fail to enforce the rule on the spoke's traffic. It evaluated the rule and found
that the traffic did not match it. The rule has a source of $HOME_NET, and 10.190.1.50 is apparently
not in it.
Ask what the variables are
A Suricata rule is only as specific as its variables, so find out what they resolve to. Variables can be set in two places — on the rule group and on the policy:
aws network-firewall describe-rule-group --rule-group-arn "$RG" \
--query 'RuleGroup.RuleVariables' --output json
# null
aws network-firewall describe-firewall-policy --firewall-policy-arn "$POLICY" \
--query 'FirewallPolicy.PolicyVariables' --output json
# nullNeither defines HOME_NET. Nothing in this deployment says what $HOME_NET is, so it has whatever
value AWS gives it when nobody says. That value is the root cause.
This debrief is part of Labs Pro
The root-cause analysis, packet-flow walkthrough, and Terraform remediation diff for this lab are available to Labs Pro members. One payment of $49, no subscription, and it covers every Pro lab now and later.
The brief, the reproduction steps, and the Terraform stay free — you can still solve this one yourself.
Already bought it? Sign in and it unlocks.