TT Lab
Get started
Learn Learning paths Courses

Tomcat & nginx Operations

Not Becoming the Person Who Raises the Timeout on a 502

Continue in TT Lab

In one line

A 502 means "an invalid response was received" and a 504 means "no response within the time limit", and if you mix the two, you burn hours just lengthening timeouts.

Why this is a problem

What eats the most time in incident response is not failing to find the cause but being confident about the wrong cause. If you conclude that a 502 means "the backend is dead", you restart the backend, if that doesn't work you lengthen the timeout, and if that still doesn't work you add servers. Meanwhile the real causes, the response header size or a connection closing early, go untouched.

With numbers, this confidence turns into evidence. If you pull the top response times by URL and the 5xx distribution from the access log, and count connection refused and upstream timed out separately in the nginx error log, 502 and 504 separate by themselves. The time it takes to count these two numbers is one minute.

502 and 504 are different stories

Code Meaning When nginx emits it Common causes
502 Bad Gateway Received an invalid response from the upstream, or the connection was refused Connection refused, response parsing failure, connection closed early Backend down, backlog saturation, response header larger than the buffer
504 Gateway Timeout The upstream did not respond within the set time proxy_read_timeout exceeded Slow query, delay in an external integration, lock wait
503 Service Unavailable No upstream is available All servers were excluded as failed Total outage, health judgment error

The point where people go wrong most often here is this: concluding that a 502 means "the upstream is dead".

The 502 as the specification defines it is "the gateway received an invalid response from upstream", not "the upstream did not respond". A case where a response did arrive but the proxy could not process it is also a 502. The typical one is a response header larger than the proxy buffer.

Why this situation is nasty: the upstream health check keeps coming back healthy. Because the health check response has a small header. So it becomes "the server is fine but only certain users get a 502". Those users are usually users who log in through SSO and have large cookies.

proxy_buffer_size       16k;
proxy_buffers         4 32k;
proxy_busy_buffers_size 64k;

And remember this without fail: when a 502 occurs, lengthening the timeout has no effect at all. Timeouts are a 504 story. If you mix the two, you waste hours.

The diagnostic ladder — five curl steps

When the symptom is ambiguous, narrow it down in this order. Step 5 in particular halves the search space.

# 1. 이름이 풀리는가
getent hosts api.example.com

# 2. 포트가 열려 있는가 (ping 은 이 환경에서 안 되니 쓰지 않는다)
curl -s -o /dev/null -w '%{http_code}\n' --connect-timeout 3 http://api.example.com/

# 3. 상태코드와 헤더
curl -sSI https://api.example.com/health

# 4. 시간 분해
curl -s -o /dev/null -w 'dns=%{time_namelookup} conn=%{time_connect} \
tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total}\n' https://api.example.com/

# 5. ★ 프록시를 건너뛰고 업스트림에 직접, 원래 Host 헤더를 유지한 채
curl -sSI -H 'Host: api.example.com' http://127.0.0.1:8080/health

The interpretation of the result of step 5 is the key.

If you leave out -H 'Host: ...', this comparison is meaningless. Where virtual host routing applies, a different Host sends you to an entirely different app.

The time breakdown of step 4 is also useful.

The minimum for reading GC logs

[2026-08-19T02:14:33.221+0900][12.334s] GC(41) Pause Full (System.gc()) 486M->402M(512M) 812.443ms

What to read from this one line.

The criterion for judgment is not the absolute value but the trend and pattern.

When an OOM occurs

The response to a java.lang.OutOfMemoryError is completely different depending on its kind.

Message Meaning First action
Java heap space Heap shortage or leak Analyze a heap dump. If you blindly raise -Xmx, the leak remains
Metaspace Class metadata shortage Suspect a classloader leak from repeated deployments. Restart and observe
GC overhead limit exceeded Spending most of the time on GC but reclaiming almost nothing In effect a leak. Heap dump
unable to create native thread Thread creation failed Not the heap but an OS limit/thread leak. Thread dump

The last row is especially confusing. It is a case of an OOM that is not a heap problem. Raising -Xmx can actually make things worse (when the heap grows, the memory for threads shrinks).

And a dump can only be obtained at the moment it blows. If you haven't put in -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=<경로> (the path) in advance, it becomes "let's look when it reproduces", and it usually doesn't reproduce.

Leave evidence before restarting

The principle of incident response is recovery first. But in the 1–2 minutes before you press the restart button, you can secure these three.

PID=$(pgrep -f tomcat | head -1)
jcmd $PID Thread.print            > /tmp/threads_$(date +%H%M%S).txt
jcmd $PID GC.heap_info            > /tmp/heap_$(date +%H%M%S).txt
cp /opt/tomcat/logs/catalina.out    /tmp/catalina_$(date +%H%M%S).out

If you save those 1–2 minutes, you will never find the cause. If "it worked once we restarted" repeats three times, the fourth time even a restart won't fix it, and by then there is no evidence at all.

What you should actually pull from the logs

Four things to pull from the access log immediately.

  1. Status code distribution — since when has 5xx increased
  2. Top URLs by response time — where is it slow
  3. Error distribution per upstream — is only a particular server a problem
  4. Requests per minute — is the traffic surge the cause or the result

Number 4 is important. The increase in requests during an outage may be the cause, but it is also often a result of users refreshing repeatedly because it doesn't work. If you look at the start time precisely, you can tell them apart.

What you see in the field

The nastiest form is "the server is fine but only certain users get a 502". The health check response has a small header so it always passes, and only some users with large session cookies exceed the proxy buffer and get a 502. The monitoring dashboards are all green then, so user inquiries are the only signal.

And you need the habit of leaving evidence before a restart. Heap dumps and thread dumps disappear the moment you restart, and you can't know the cause until the same outage happens again.