Lab 21: The Hardening Change That Took The Site Down
A security review asked for HTTP to redirect to HTTPS, so it does. The application is healthy and serving. The Application Load Balancer marks every target unhealthy, fails open to nothing, and answers every client with 503 — because of a default nobody has ever had a reason to look at.
- Debugging time
- ~25 min
- Reading time
- 12 min
- Reported by
- Web Platform
- Tier
- Associate
Site returns 503 from the load balancer after the HTTPS redirect change; the web servers themselves are up
Reported by Web Platform
We deployed the change from the security review this afternoon: nginx now redirects all plaintext HTTP
to HTTPS. Standard return 301. It rolled out cleanly and I confirmed it on the box — curl to port 80
gets a 301 with the right Location, curl to 443 gets the page.
About two minutes later the site started returning 503 Service Temporarily Unavailable to everyone.
Not from nginx — the 503 is coming from the load balancer. The target group shows the instance as
unhealthy.
The instance is fine. I am logged into it. nginx is running, both ports answer, CPU is idle, nothing in the error log. I have not touched the load balancer, the target group, the listener or the security groups. The only change was the redirect, and the redirect is working exactly as specified.
What you are working with
One VPC, one Application Load Balancer, one target.
The load balancer spans two zones because an Application Load Balancer requires it; the second zone is otherwise empty. The listener forwards HTTP to a target group holding one instance. That instance is doing exactly what it was changed to do, and the target group has marked it unhealthy for it.
| Resource | Configuration |
|---|---|
| Load balancer | Application, two public subnets, security group admits 80 from the VPC |
| Listener | HTTP :80 → forward to the target group |
| Target group | Instance type, HTTP, port 80 |
| Health check | HTTP, path /, port traffic-port, success codes 200, interval 10s |
| Target | nginx: :80 returns 301 https://…, :443 serves lab-21-application-ok |
| Target health | unhealthy |
The health check configuration has not been changed since the target group was created. It is the default in every respect that matters.
- EC2
Load balancer node
GET / HTTP/1.1, User-Agent: ELB-HealthChecker/2.0
- FILTER
Target security group
port 80 from the load balancer: allowed
- DEST
nginx :80
answers immediately, as configured
- RTB
Health check evaluation
compare the response code to the success codes
Dropped — The target answered. The answer is valid HTTP. It is not the answer the health check was told to accept, and a health check does not follow where the answer points.
Every hop on the way to the target passes, and the target answers promptly. The failure is in the comparison afterwards. The target is marked unhealthy for giving a correct response to the wrong question.
Scope and constraints
- In scope: why a target that is up, serving, and answering health checks promptly is marked unhealthy, and why the load balancer answers 503.
- Out of scope: the security groups, the listener, the subnets, the instance, and nginx. All correct, and the reproduction confirms each.
- The reporter is right: the instance is healthy and the redirect works as specified. Both facts are true at the same time as the outage, which is the point.
- The fix is a small change to either the health check or the application. Which one is the design decision this lab is actually about.
Deploy the broken state
cd lab-21-alb-health-check-redirect
terraform init
terraform applyThe load balancer takes two to three minutes to provision, and the target needs two failed checks at a 10-second interval to be marked unhealthy, so the broken state is in place about a minute after the load balancer is active.
ALB=$(terraform output -raw alb_dns_name)
TG=$(terraform output -raw target_group_arn)
APP=$(terraform output -raw app_private_ip)Confirm the outage
From the control host — terraform output -raw start_session_command:
curl -s -o /dev/null -w '%{http_code}\n' "http://$ALB/"
# 503The load balancer answers 503. That code is the load balancer's own, not the application's: it means no healthy target was available to forward to.
Confirm the application is healthy
Same host, straight to the instance, bypassing the load balancer:
curl -s -o /dev/null -w '%{http_code} %{redirect_url}\n' "http://$APP/"
# 301 https://10.210.11.50/
curl -sk "https://$APP/"
# lab-21-application-okPort 80 redirects. Port 443 serves. The reporter is right on every count: the application is up, and the change does exactly what the review asked for.
Read the target's health in full
Not just the status — the reason:
aws elbv2 describe-target-health --target-group-arn "$TG" \
--query 'TargetHealthDescriptions[].TargetHealth' --output json
# [
# {
# "State": "unhealthy",
# "Reason": "Target.ResponseCodeMismatch",
# "Description": "Health checks failed with these codes: [301]"
# }
# ]That is the diagnosis in one field. The health check got a 301. It wanted something else.
Read what the health check wanted
aws elbv2 describe-target-groups --target-group-arns "$TG" \
--query 'TargetGroups[0].[HealthCheckProtocol,HealthCheckPath,HealthCheckPort,Matcher.HttpCode]' \
--output text
# HTTP / traffic-port 200GET / on the traffic port, over HTTP, and the only acceptable answer is 200. The application now
answers 301 to every such request, by design.
Watch it happen on the target
Open a session on the application host — terraform output -raw app_session_command — and tail
nginx's access log:
sudo tail -n 5 /var/log/nginx/access.log
# 10.210.1.137 - [30/Sep/2026:14:31:02 +0000] "GET / HTTP/1.1" 301 "ELB-HealthChecker/2.0"
# 10.210.2.201 - [30/Sep/2026:14:31:04 +0000] "GET / HTTP/1.1" 301 "ELB-HealthChecker/2.0"
# 10.210.1.137 - [30/Sep/2026:14:31:12 +0000] "GET / HTTP/1.1" 301 "ELB-HealthChecker/2.0"Every ten seconds, from each load balancer node, GET / arrives and gets a 301. nginx is answering
every health check, correctly, with a redirect. And nothing ever follows the redirect — there is no
corresponding request on port 443 from those addresses, because a health check does not follow a
Location header. It reads the code, compares, and records a failure.
This debrief is part of Labs Pro
The root-cause analysis, packet-flow walkthrough, and Terraform remediation diff for this lab are available to Labs Pro members. One payment of $49, no subscription, and it covers every Pro lab now and later.
The brief, the reproduction steps, and the Terraform stay free — you can still solve this one yourself.
Already bought it? Sign in and it unlocks.