Skip to content
AssociateProLoad Balancing

Lab 21: The Hardening Change That Took The Site Down

A security review asked for HTTP to redirect to HTTPS, so it does. The application is healthy and serving. The Application Load Balancer marks every target unhealthy, fails open to nothing, and answers every client with 503 — because of a default nobody has ever had a reason to look at.

Debugging time
~25 min
Reading time
12 min
Reported by
Web Platform
Tier
Associate
INC-1806SEV-1InvestigatingOpened 2026-09-30 14:22 UTC

Site returns 503 from the load balancer after the HTTPS redirect change; the web servers themselves are up

Reported by Web Platform

We deployed the change from the security review this afternoon: nginx now redirects all plaintext HTTP to HTTPS. Standard return 301. It rolled out cleanly and I confirmed it on the box — curl to port 80 gets a 301 with the right Location, curl to 443 gets the page.

About two minutes later the site started returning 503 Service Temporarily Unavailable to everyone. Not from nginx — the 503 is coming from the load balancer. The target group shows the instance as unhealthy.

The instance is fine. I am logged into it. nginx is running, both ports answer, CPU is idle, nothing in the error log. I have not touched the load balancer, the target group, the listener or the security groups. The only change was the redirect, and the redirect is working exactly as specified.

What you are working with

One VPC, one Application Load Balancer, one target.

Starting state — a load balancer, a target, and a redirect
VPC 10.210.0.0/16 Attached: Application Load Balancer. AZ a contains public-a 10.210.1.0/24, holding Control host; app 10.210.11.0/24, holding nginx. AZ b contains public-b 10.210.2.0/24, holding no resources.VPC 10.210.0.0/16AZ apublic-apublic10.210.1.0/24Control host10.210.1.50apppublic10.210.11.0/24Risk — target group reports unhealthynginx:80 → 301 https:// · :443 serves the siteAZ bpublic-bpublic10.210.2.0/24no resourcesApplication LoadBalancerHTTP :80 listener → targetgroup

The load balancer spans two zones because an Application Load Balancer requires it; the second zone is otherwise empty. The listener forwards HTTP to a target group holding one instance. That instance is doing exactly what it was changed to do, and the target group has marked it unhealthy for it.

ResourceConfiguration
Load balancerApplication, two public subnets, security group admits 80 from the VPC
ListenerHTTP :80 → forward to the target group
Target groupInstance type, HTTP, port 80
Health checkHTTP, path /, port traffic-port, success codes 200, interval 10s
Targetnginx: :80 returns 301 https://…, :443 serves lab-21-application-ok
Target healthunhealthy

The health check configuration has not been changed since the target group was created. It is the default in every respect that matters.

health check: GET / on port 80, every 10 seconds
  1. EC2

    Load balancer node

    GET / HTTP/1.1, User-Agent: ELB-HealthChecker/2.0

  2. FILTER

    Target security group

    port 80 from the load balancer: allowed

  3. DEST

    nginx :80

    answers immediately, as configured

  4. RTB

    Health check evaluation

    compare the response code to the success codes

    Dropped — The target answered. The answer is valid HTTP. It is not the answer the health check was told to accept, and a health check does not follow where the answer points.

Every hop on the way to the target passes, and the target answers promptly. The failure is in the comparison afterwards. The target is marked unhealthy for giving a correct response to the wrong question.

Scope and constraints

  • In scope: why a target that is up, serving, and answering health checks promptly is marked unhealthy, and why the load balancer answers 503.
  • Out of scope: the security groups, the listener, the subnets, the instance, and nginx. All correct, and the reproduction confirms each.
  • The reporter is right: the instance is healthy and the redirect works as specified. Both facts are true at the same time as the outage, which is the point.
  • The fix is a small change to either the health check or the application. Which one is the design decision this lab is actually about.

Deploy the broken state

cd lab-21-alb-health-check-redirect
terraform init
terraform apply

The load balancer takes two to three minutes to provision, and the target needs two failed checks at a 10-second interval to be marked unhealthy, so the broken state is in place about a minute after the load balancer is active.