Why 502 Bad Gateway Matters in 2024
If you operate modern web services, a sudden surge of 502 Bad Gateway errors can take down critical user flows, damage customer trust, and wreck SLAs. In 2024’s DevOps landscape—microservices, CDNs, managed load balancers, service meshes, and edge security layers—502s are common but solvable. This guide breaks down the causes, shows you exactly how to read your logs, and walks you through a systematic, reproducible troubleshooting workflow tailored for DevOps teams.
You’ll learn how to quickly isolate the layer causing the error, validate timeouts, decode reverse proxy logs, verify TLS and DNS, and make lasting fixes that prevent recurrence. Keep this as a runbook for your next incident.
What a 502 Bad Gateway Really Means
A 502 Bad Gateway means a gateway or proxy (e.g., NGINX, HAProxy, Envoy, Cloudflare, AWS ALB, or a service mesh sidecar) received an invalid response from an upstream server. “Invalid” can be:
- No response at all (timeout)
- TCP reset or connection refused
- Malformed HTTP response (headers/body corruption)
- TLS handshake failure
- Protocol mismatch (e.g., expecting HTTP/1.1, receiving HTTP/2-gRPC without proper translation)
Contrast with similar errors:
- 500 Internal Server Error: Your application failed while processing the request.
- 503 Service Unavailable: Often a capacity issue, maintenance mode, or upstream unavailable but acknowledged as such.
- 504 Gateway Timeout: The gateway didn’t receive a response in time; specifically a timeout condition.
A 502 is often a symptom of an upstream dependency problem or a misconfiguration in the gateway layer.
Common Causes of 502 Errors
- Upstream timeout: The backend took too long to respond; proxy terminated early.
- Connection refused/reset: App was not listening or crashed; firewall dropped packets.
- TLS mismatch: Cipher mismatch, invalid certs, SNI mismatch, or TLS version incompatibility.
- Header/body issues: Oversized headers or invalid chunked encoding causing parsing errors.
- DNS problems: Upstream resolves to wrong IP, stale DNS cache, or intermittent DNS failures.
- Resource exhaustion: File descriptor limits, connection pool saturation, CPU spikes, or memory pressure causing restarts.
- Protocol mismatch: gRPC behind HTTP/1.1 without proper grpc-web translation or Envoy config.
- WAF/CDN edge behavior: Provider returns 502 when origin misbehaves or violates policies.
- Kubernetes readiness/liveness misconfig: Traffic routed to a pod that isn’t ready.
- Network path disruptions: Security groups, NACLs, MTU mismatches, or NAT gateway issues.
- Misconfigured proxy parameters: Incorrect keepalive, buffer sizes, max header size, or upstream declarations.
Quick Triage Checklist (30–90 Seconds)
- Confirm blast radius: Is it global, a specific region, or only certain paths/users?
- Reproduce with curl from multiple vantage points (edge node, bastion, local).
- Check status pages (CDN, cloud provider) and recent deploys.
- Inspect gateway logs: NGINX/Envoy/HAProxy/Cloud LB metrics.
- Spot spikes in latency or error rate in dashboards (4xx/5xx split).
- Roll back the last change if there’s a strong correlation and high impact.
A Step-by-Step Troubleshooting Workflow
1) Reproduce and Scope
-
Use curl with verbose logging:
curl -v https://api.example.com/health curl -v --http2 https://api.example.com/endpoint -
Capture headers and status codes. If it’s intermittent, run multiple times or use a small script.
-
Check multiple clients/networks:
curl -v --resolve api.example.com:443:203.0.113.10 https://api.example.com/healthThis bypasses DNS and hits a specific IP directly, helping isolate DNS issues.
2) Identify the Layer Returning 502
- CDN headers (e.g., CF-*) indicate if a CDN is responding.
- Load balancer headers (e.g., X-ALB, Via) reveal edge participation.
- Gateway logs (NGINX, Envoy, HAProxy) will definitively show upstream issues.
If the CDN returns 502 but origin is fine, the issue may be CDN-origin connectivity or a WAF policy. If the reverse proxy returns 502, the upstream app or internal network may be at fault.
3) Review Reverse Proxy/Load Balancer Logs
-
NGINX access log (with upstream timing and status):
203.0.113.45 - - [03/Oct/2024:12:01:22 +0000] "GET /checkout HTTP/1.1" 502 150 "-" "curl/8.6.0" "upstream_response_time=1.995 request_time=2.000 upstream_status=502 upstream_addr=10.0.4.23:8080"Action: Compare upstream_response_time with proxy_read_timeout. If close to timeout, increase timeouts or optimize the backend.
-
HAProxy:
Oct 03 12:01:22 haproxy[1234]: 203.0.113.45:53214 [03/Oct/2024:12:01:20.005] fe~ be/srv1 2000/0/0/2000/2000 502 0 - - PR-- 1/1/0/0/0 0/0 "GET /checkout HTTP/1.1"The timers indicate connect/queue/TTFB/total. A long TTFB suggests backend slowness or timeout.
-
Envoy (if using service mesh or API gateway): Look for upstream_connect_failure, upstream_reset_before_response_started, or downstream_rq_5xx.
4) Inspect Application Logs and Health
- Check application logs at the time of errors for unhandled exceptions, slow queries, or GC pauses.
- Verify process health:
ps aux | grep your-app netstat -tulpn | grep 8080 lsof -i :8080 | wc -l - Validate health endpoints:
curl -s http://localhost:8080/health - Check resource utilization:
If file descriptors are exhausted, increase ulimit and tune connection pooling.top | head -20 free -m df -h ulimit -n
5) Verify Timeouts and Proxy Settings
Common NGINX configs impacting 502s:
upstream backend {
server 10.0.4.23:8080 max_fails=3 fail_timeout=10s;
keepalive 64;
}
server {
proxy_connect_timeout 3s;
proxy_send_timeout 60s;
proxy_read_timeout 60s; # Increase if upstream is legitimately slow
send_timeout 60s;
proxy_buffering on;
proxy_buffers 16 16k;
proxy_busy_buffers_size 64k;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_http_version 1.1;
proxy_set_header Connection "";
}
If your upstream uses FastCGI (PHP-FPM) or uWSGI, also review:
- fastcgi_read_timeout
- uwsgi_read_timeout
- fastcgi_buffers and buffer sizes
- client_header_buffer_size and large_client_header_buffers if header size issues surface
6) TLS Handshake and Certificate Issues
-
Validate TLS from the gateway to upstream (if TLS-terminated upstream):
openssl s_client -connect backend.example.internal:443 -servername backend.example.internal -tls1_2Check for certificate errors, SNI mismatch, or unsupported ciphers.
-
Ensure LB/Proxy and upstream agree on TLS versions. If you tightened TLS policies in 2024 (e.g., dropped TLS 1.0/1.1), some upstream services might still be misconfigured.
-
In AWS ALB/NLB with TLS targets, confirm the correct security policy and that the target presents a cert for the hostname you use.
7) DNS and Network Path
- Confirm DNS resolution is correct and fast:
dig backend.service.local +short - Check from inside the cluster or VPC:
nslookup backend.default.svc.cluster.local - Trace the route and MTU issues:
mtr -rw backend.internal ping -M do -s 1472 backend.internal # Detect MTU blackhole - Verify security groups, NACLs, and firewall rules aren’t intermittently dropping connections.
8) Kubernetes-Specific Checks
-
Readiness and liveness probes:
kubectl describe pod <pod> kubectl logs <pod> -c <container> --since=5m kubectl get endpoints <service>Ensure the Service has healthy endpoints and the readiness probe matches actual app readiness. If pods are cycling, the LB may send traffic to pods not yet ready, causing 502s.
-
Sidecar/mesh considerations (Istio/Linkerd/Consul/Envoy):
- Conflicting timeouts between sidecar and app
- HTTP/2 or gRPC configuration mismatches
- mTLS issues between sidecars
-
Node pressure:
kubectl describe node <node> | grep -i pressure -A2CPU or memory pressure may throttle your app pods, causing timeouts.
9) CDN and WAF Diagnostics
-
Cloudflare 502: Often “Bad Gateway” when origin refuses connections or times out. Check:
- Origin health and firewall rules
- Cloudflare’s “Origin Reachability” metrics
- Bypass CDN and connect directly to origin for isolation
-
AWS CloudFront: Look at Origin Response metrics, increase origin timeout, confirm custom headers and authentication.
-
WAF rules: Check for false positives blocking or interfering with valid traffic. Temporarily bypass/disable a rule to confirm suspicion.
10) Service Dependency Failures
Your gateway might be fine, but your app’s downstream calls (DB, cache, third-party APIs) slow down and cause your app to respond late or not at all, leading to a 502 at the gateway.
- Look at APM traces spanning gateway → app → DB/external API.
- Check DB connection pools (max size, timeouts) and slow query logs.
- Implement circuit breakers and bulkheads to prevent cascading failures.
Reading and Interpreting Logs Like a Pro
NGINX Error Log Patterns
-
Upstream timed out:
upstream timed out (110: Connection timed out) while reading response header from upstream, client: 203.0.113.45, server: api.example.com, request: "GET /checkout HTTP/1.1", upstream: "http://10.0.4.23:8080/checkout", host: "api.example.com"Action: Raise proxy_read_timeout or optimize upstream latency.
-
Upstream prematurely closed connection:
upstream prematurely closed connection while reading response header from upstreamAction: Check app errors or crashes, connection resets, or resource limits.
-
connect() failed (111: Connection refused):
connect() failed (111: Connection refused) while connecting to upstreamAction: Verify the app is listening, container port mapping, and security rules.
-
no live upstreams:
no live upstreams while connecting to upstreamAction: All upstream targets are unhealthy or misconfigured.
HAProxy Log Fields to Watch
- Termination state flags (e.g., PR, SD, CR) indicate where the connection ended.
- Latency fields show connect time, TTFB, and total time—triage whether the delay is during connect or response.
Envoy Access Logs
- Look for upstream_reset, reset_reason (e.g., connection_failure, overflow), and response_flags like UF (upstream connection failure), UC (upstream connection termination).
Actionable Fixes by Root Cause
-
Timeouts:
- Increase gateway timeouts slightly above your 99th percentile latency.
- Set explicit application timeouts and optimize slow paths.
- Use retry budgets with backoff for idempotent requests.
-
Connection exhaustion:
- Raise ulimit -n and tune connection pooling.
- Enable reuse/keepalive and right-size worker processes.
- Monitor active connections and backlog.
-
TLS issues:
- Align TLS versions and ciphers between LB and upstream.
- Ensure SNI is configured and certs are valid and not expiring.
- Use consistent CA trust stores across environments.
-
Header/body parsing:
- Increase large_client_header_buffers if you see “upstream sent too big header.”
- Ensure proper response encoding and compression only once (avoid double gzip).
-
DNS:
- Reduce TTL if targets change frequently.
- Ensure gateways and apps respect TTL and don’t cache indefinitely.
- Pin IPs temporarily with --resolve to isolate DNS while you fix it.
-
Kubernetes:
- Fix readiness probes to reflect true readiness (DB connectivity, migrations).
- Use maxUnavailable and PodDisruptionBudgets to avoid draining all healthy pods.
- Align sidecar and app timeouts; disable HTTP/2 if the upstream can’t handle it.
-
WAF/CDN:
- Adjust rulesets causing false positives.
- Increase origin timeouts if traffic is legitimately slow.
- Enable “origin shield” near your origin to reduce connection churn.
Practical Examples
Example 1: NGINX 502 Due to Upstream Timeout
Symptom:
- Sporadic 502s on POST /checkout with upstream timed out messages in logs.
Diagnosis:
- NGINX shows upstream_response_time ≈ 60s, proxy_read_timeout is 60s.
Fix:
- Increase proxy_read_timeout to 90s temporarily.
- Profile the application’s checkout flow; optimize DB joins; add caching.
- Final: restore timeout to 60s, keep p95 under 500ms and p99 under 2s.
Config change:
proxy_read_timeout 90s; # temporary
Example 2: ALB 502 Because of TLS Policy Mismatch
Symptom:
- After rotating certificates and updating ALB security policy to require TLS 1.2+, internal legacy upstream still uses TLS 1.0.
Diagnosis:
- openssl s_client fails with handshake failure when targeting upstream.
Fix:
- Update upstream to support TLS 1.2; temporarily relax policy via a maintenance flag or route legacy traffic differently.
Example 3: Kubernetes Readiness Misconfiguration
Symptom:
- New deployment yields a spike in 502s. Pods start receiving traffic immediately.
Diagnosis:
- readinessProbe only checks “/health” but ignores that the app warms caches and migrations for ~25s.
Fix:
- Adjust readinessProbe to check deeper dependencies or add initialDelaySeconds:
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 30
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
Example 4: Cloudflare 502 at the Edge
Symptom:
- Global 502 reports, but origin logs show no traffic.
Diagnosis:
- Cloudflare failing to connect to origin due to firewall rule on origin IP range after an infrastructure change.
Fix:
- Update firewall to allow Cloudflare IPs; add allowlists and automated updates.
Observability: What to Measure and Alert
-
SLOs:
- Availability SLO for critical endpoints (e.g., 99.9% over 30 days).
- Latency SLO (p95, p99) by endpoint.
-
Metrics:
- Upstream 5xx rate split by gateway vs. app.
- Gateway upstream_response_time and connect_time.
- Active connections, connection errors, TCP resets.
- DNS resolution latency and error rate.
- TLS handshake errors.
-
Tracing:
- End-to-end tracing (gateway → app → DB) to locate bottlenecks.
- Annotate deploys in traces to correlate spikes.
-
Logging:
- Structured logs with correlation IDs (X-Request-ID).
- Log upstream_addr, upstream_status, and timing.
-
Dashboards:
- Per-service 5xx heatmap across regions.
- Gateway error decomposition by reason (timeout, connect, TLS).
-
Alerting:
- Sudden increase in 5xx above baseline.
- Saturation signals (CPU > 90%, connection pool full, thread pool saturation).
- DNS and TLS anomaly detection.
Runbook: 502 Bad Gateway
When paged:
-
Validate impact:
- Check dashboards for region/path scope.
- Try curl from two locations, with and without CDN.
-
Identify layer:
- Inspect CDN/LB/proxy logs for upstream_status and reset reasons.
- If CDN returns 502, bypass CDN; if proxy returns 502, check upstream.
-
Gather logs:
- NGINX error/access logs, HAProxy/Envoy logs, application logs for the same timestamps.
-
Check timeouts and health:
- proxy_read_timeout, connect timeouts, upstream health endpoints.
- Pod readiness and endpoint lists in Kubernetes.
-
Test connectivity:
- openssl s_client for TLS, dig/nslookup for DNS, mtr/ping for network path.
- netstat/lsof for listening ports and FD counts.
-
Mitigate:
- Roll back recent config/deploy if correlated.
- Temporarily increase timeouts or autoscale.
- Disable problematic WAF rule or CDN feature temporarily.
-
Fix forward:
- Optimize slow code paths, tune pools/timeouts, correct TLS policies.
- Improve readiness probes and health checks.
- Add alerts for precursors (latency creep, connection saturation).
-
Post-incident:
- Document root cause and corrective actions.
- Add tests or canaries to detect reoccurrence.
- Update runbook and SLO error budgets.
Tools You’ll Use Frequently
- curl and httpie:
- curl -v, --http2, --resolve, --connect-timeout, --max-time
- dig, nslookup:
- dig +trace, verify TTLs and record correctness
- openssl:
- s_client for TLS handshake and SNI checks
- mtr, traceroute, ping:
- Network path and MTU verification
- tcpdump/Wireshark:
- Packet-level analysis for resets and handshake failures
- kubectl:
- logs, describe, get endpoints, port-forward
- APM/Tracing:
- Jaeger, Zipkin, OpenTelemetry collectors
- Log search:
- Query by correlation ID across gateway and app logs
Prevention Strategies for 2024 Architectures
-
Use retries judiciously:
- At gateway and client, only for idempotent methods.
- Add jittered backoff and enforce retry budgets.
-
Circuit breakers and timeouts:
- Fail fast to avoid cascading failures.
- Prefer per-call timeouts in the app to align with gateway timeouts.
-
Capacity management:
- Autoscale based on p95 latency and queue depth, not CPU alone.
- Pre-warm caches before scaling events or deploys.
-
Robust health checks:
- Make readiness checks reflect all critical dependencies.
- Drain connections gracefully during deploys and rollouts.
-
Consistent protocol usage:
- If using gRPC, ensure gateways are configured for HTTP/2 or grpc-web translation.
- Avoid mixed-mode proxies without explicit config.
-
TLS and cert hygiene:
- Automate cert renewals, validate SNI, monitor for expiry.
- Standardize TLS policies across LB, gateway, and services.
-
Observability baked in:
- Correlation IDs from edge to DB.
- Red/Golden signals on dashboards with on-call ownership.
-
Chaos and failure injection:
- Test timeouts, DNS failures, and dependency slowness regularly.
- Validate that alerts fire early and runbooks resolve issues quickly.
A Short Decision Tree
-
Is 502 returned by CDN?
- Yes → Bypass CDN; check origin connectivity, WAF rules, origin certs/IP allowlists.
- No → Check reverse proxy logs.
-
Proxy shows connect errors?
- Yes → App not listening, firewall, DNS wrong, or wrong port.
- No → Proxy shows read timeout?
- Yes → Increase proxy_read_timeout or speed up upstream; analyze app and DB performance.
- No → TLS/Protocol error?
- Yes → Align TLS versions/ciphers, enable SNI, or configure gRPC/HTTP/2 properly.
-
Kubernetes involved?
- Check readiness, endpoints, pod restarts, node pressure, and mesh sidecar logs.
Final Takeaways
- A 502 is rarely random. It’s an upstream or gateway mismatch in timing, protocol, or connectivity.
- Start at the edge and move inward: CDN → LB → gateway → app → dependencies.
- Logs are your map. Learn to read upstream timing, status, and reset reasons.
- Fixes often combine configuration (timeouts, TLS), reliability patterns (circuit breakers, retries), and performance work (DB/query optimization).
- Prevent recurrence with strong readiness checks, aligned timeouts, consistent protocols, and end-to-end observability.
With this step-by-step approach and the detailed checks above, you’ll cut your mean time to resolution, protect your SLOs, and keep 502s from surprising you again.