Not Becoming the Person Who Raises the Timeout on a 502
In one line
A 502 means "an invalid response was received" and a 504 means "no response within the time limit", and if you mix the two, you burn hours just lengthening timeouts.
Why this is a problem
What eats the most time in incident response is not failing to find the cause but being confident about the wrong cause. If you conclude that a 502 means "the backend is dead", you restart the backend, if that doesn't work you lengthen the timeout, and if that still doesn't work you add servers. Meanwhile the real causes, the response header size or a connection closing early, go untouched.
With numbers, this confidence turns into evidence. If you pull the top response times by URL and the 5xx distribution from the access log, and count connection refused and upstream timed out separately in the nginx error log, 502 and 504 separate by themselves. The time it takes to count these two numbers is one minute.
502 and 504 are different stories
| Code | Meaning | When nginx emits it | Common causes |
|---|---|---|---|
| 502 Bad Gateway | Received an invalid response from the upstream, or the connection was refused | Connection refused, response parsing failure, connection closed early | Backend down, backlog saturation, response header larger than the buffer |
| 504 Gateway Timeout | The upstream did not respond within the set time | proxy_read_timeout exceeded |
Slow query, delay in an external integration, lock wait |
| 503 Service Unavailable | No upstream is available | All servers were excluded as failed | Total outage, health judgment error |
The point where people go wrong most often here is this: concluding that a 502 means "the upstream is dead".
The 502 as the specification defines it is "the gateway received an invalid response from upstream", not "the upstream did not respond". A case where a response did arrive but the proxy could not process it is also a 502. The typical one is a response header larger than the proxy buffer.
Why this situation is nasty: the upstream health check keeps coming back healthy. Because the health check response has a small header. So it becomes "the server is fine but only certain users get a 502". Those users are usually users who log in through SSO and have large cookies.
proxy_buffer_size 16k;
proxy_buffers 4 32k;
proxy_busy_buffers_size 64k;
And remember this without fail: when a 502 occurs, lengthening the timeout has no effect at all. Timeouts are a 504 story. If you mix the two, you waste hours.
The diagnostic ladder — five curl steps
When the symptom is ambiguous, narrow it down in this order. Step 5 in particular halves the search space.
# 1. 이름이 풀리는가
getent hosts api.example.com
# 2. 포트가 열려 있는가 (ping 은 이 환경에서 안 되니 쓰지 않는다)
curl -s -o /dev/null -w '%{http_code}\n' --connect-timeout 3 http://api.example.com/
# 3. 상태코드와 헤더
curl -sSI https://api.example.com/health
# 4. 시간 분해
curl -s -o /dev/null -w 'dns=%{time_namelookup} conn=%{time_connect} \
tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total}\n' https://api.example.com/
# 5. ★ 프록시를 건너뛰고 업스트림에 직접, 원래 Host 헤더를 유지한 채
curl -sSI -H 'Host: api.example.com' http://127.0.0.1:8080/health
The interpretation of the result of step 5 is the key.
- A direct call to the upstream is healthy → the problem is between the proxy and the upstream (configuration, headers, buffers, timeouts)
- A direct call to the upstream also fails → the proxy is innocent. Look at the application or the DB
If you leave out -H 'Host: ...', this comparison is meaningless. Where virtual host routing applies,
a different Host sends you to an entirely different app.
The time breakdown of step 4 is also useful.
time_appconnectitself is long → the TLS handshake. Suspect certificate chain validation or an OCSP lookup- A gap opens between
time_appconnectandtime_starttransfer→ the application/DB. It is not the network
The minimum for reading GC logs
[2026-08-19T02:14:33.221+0900][12.334s] GC(41) Pause Full (System.gc()) 486M->402M(512M) 812.443ms
What to read from this one line.
Pause Full— it is a Full GC. If frequent, it is a problem486M->402M(512M)— before reclaim → after reclaim (total heap). If the after-reclaim figure is close to the total, memory is really short. It is a different problem from reclaim working well but running often812.443ms— the application stops for this long. This is often the real cause of response time spikes
The criterion for judgment is not the absolute value but the trend and pattern.
- Usage keeps rising even after a Full GC → suspect a memory leak → heap dump
- Full GCs are rare but Young GCs are very frequent → excessive creation of short-lived objects → a code problem
- Concentrated only at a specific time (batch time) → that batch is the culprit
When an OOM occurs
The response to a java.lang.OutOfMemoryError is completely different depending on its kind.
| Message | Meaning | First action |
|---|---|---|
Java heap space |
Heap shortage or leak | Analyze a heap dump. If you blindly raise -Xmx, the leak remains |
Metaspace |
Class metadata shortage | Suspect a classloader leak from repeated deployments. Restart and observe |
GC overhead limit exceeded |
Spending most of the time on GC but reclaiming almost nothing | In effect a leak. Heap dump |
unable to create native thread |
Thread creation failed | Not the heap but an OS limit/thread leak. Thread dump |
The last row is especially confusing. It is a case of an OOM that is not a heap problem.
Raising -Xmx can actually make things worse (when the heap grows, the memory for threads shrinks).
And a dump can only be obtained at the moment it blows.
If you haven't put in -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=<경로> (the path) in advance,
it becomes "let's look when it reproduces", and it usually doesn't reproduce.
Leave evidence before restarting
The principle of incident response is recovery first. But in the 1–2 minutes before you press the restart button, you can secure these three.
PID=$(pgrep -f tomcat | head -1)
jcmd $PID Thread.print > /tmp/threads_$(date +%H%M%S).txt
jcmd $PID GC.heap_info > /tmp/heap_$(date +%H%M%S).txt
cp /opt/tomcat/logs/catalina.out /tmp/catalina_$(date +%H%M%S).out
If you save those 1–2 minutes, you will never find the cause. If "it worked once we restarted" repeats three times, the fourth time even a restart won't fix it, and by then there is no evidence at all.
What you should actually pull from the logs
Four things to pull from the access log immediately.
- Status code distribution — since when has 5xx increased
- Top URLs by response time — where is it slow
- Error distribution per upstream — is only a particular server a problem
- Requests per minute — is the traffic surge the cause or the result
Number 4 is important. The increase in requests during an outage may be the cause, but it is also often a result of users refreshing repeatedly because it doesn't work. If you look at the start time precisely, you can tell them apart.
What you see in the field
The nastiest form is "the server is fine but only certain users get a 502". The health check response has a small header so it always passes, and only some users with large session cookies exceed the proxy buffer and get a 502. The monitoring dashboards are all green then, so user inquiries are the only signal.
And you need the habit of leaving evidence before a restart. Heap dumps and thread dumps disappear the moment you restart, and you can't know the cause until the same outage happens again.