When something breaks in production, the first signal is usually an HTTP status code. Most of the time, four of them tell me almost everything I need to know about what just happened.
5xx is the server saying “this one is on me.” 4xx is the server saying “this one is on you.” 429 is the one 4xx every SRE watches.
Here’s how I read them.
| Code | Name | What it really means | Retry? |
|---|---|---|---|
| 500 | Internal Server Error | Something blew up on the server. No idea what. | Sometimes |
| 502 | Bad Gateway | Proxy got a bad response from upstream. | Yes, with backoff |
| 503 | Service Unavailable | Service is overloaded or in maintenance. | Yes, respect Retry-After |
| 429 | Too Many Requests | You sent too much. Slow your client down. | Yes, respect Retry-After |
500 is the one you don’t want. It’s the uncaught exception, the null pointer, the panic. Your code did something it didn’t expect. Logs are the only way home. If 500s spike, you have a bug, not a capacity problem.
502 is a finger pointing. Your edge or load balancer got something it couldn’t parse from the backend. Almost always the backend is at fault: it crashed, ran out of memory, took too long to respond, or sent malformed HTTP. Restarting nginx rarely helps because nginx is just the messenger. Check nginx error.log for the clue (upstream prematurely closed, connection refused, upstream timed out) and the backend’s own logs for root cause.
503 is honest. The service is telling you it can’t take more right now. Healthy systems return 503 on purpose when they hit capacity. It’s a load-shedding signal. Always include a Retry-After header so clients know when to come back.
429 is the throttle. Every API at scale uses it: AWS, GitHub, Stripe, Cloudflare. It’s a 4xx because the server is healthy, you’re just sending too much. If you see 429s in your client metrics, you’re the noisy neighbor. Add jitter, respect Retry-After, and back off before someone else does it for you.
Treat 500 as a bug. Treat 502 and 503 as capacity. Treat 429 as your problem.
At scale, the mix of these four codes is the dashboard I check first. A small rate of 503 and 429 is healthy. A small rate of 500 is not.