Create 502 and 504 Yourself and Tell Them Apart by Evidence
Goal
You produce 502 and 504 yourself, and then learn to pin down the cause not from the status code but from the phrase in error.log.
In particular, you learn to distinguish two cases that are the same 504 but hit different timeouts.
Why it matters
What eats the most time in incident response is being confident about the wrong cause.
If you conclude that a 502 means "the backend is dead" and raise a timeout, nothing happens,
because in a failure where the connection was refused there was no time to wait at all.
Conversely, a 504 doesn't mean you can fix it by raising proxy_read_timeout either.
If it was hit at the connection stage, even raising the read timeout to 30 seconds still cuts the request off
after 2 seconds. Step 5 of this lab is exactly that scene.
What makes that difference is just one phrase after while in one line of the log.
Steps
- Save
/root/ngi/upstream.pyand start it in the background. The app comes up on port 9101 and port 9109 becomes a port that does not accept connections. Nothing is started on port 9999. Check:/on port 9101 returns 200,/slow?d=2returns 200 after 2 seconds, and the/bighdr?n=8000response has anX-Session-Blobheader. - Write
/root/ngi/nginx.confand start nginx. Listen on port 8088 and set up three paths — proxy/app/to port 9101,/dead/to port 9999 and/hang/to port 9109. Puterror_logat/root/ngi/logs/error.logat warn level, put the access log at/root/ngi/logs/access.log, and include all of$upstream_response_time,$request_timeand$upstream_statusin its format. Setproxy_connect_timeout 2s;andproxy_read_timeout 3s;. Check:http://127.0.0.1:8088/app/must return 200. (You can open it with the browser preview.) - Call
http://127.0.0.1:8088/dead/and create/root/ngi/case-refused.txt. It has three lines.status=<받은 상태 코드> phrase=<error.log 에서 원인을 지목하는 구절 그대로> timeout=<실제로 발동한 타임아웃 이름, 없으면 none> - Call
http://127.0.0.1:8088/app/slow?d=10, measure how many seconds it takes to be cut off, and create/root/ngi/case-read.txt. It has four lines.status=<상태 코드> phrase=<error.log 의 while 로 시작하는 구절 그대로> timeout=<발동한 타임아웃 지시자 이름> elapsed=<끊기기까지 걸린 초, 정수> - Inside the
/hang/block, putproxy_read_timeout 30s;, reload, and then callhttp://127.0.0.1:8088/hang/. Even though you gave the read 30 seconds, the request is cut off much sooner. Measure the time and create/root/ngi/case-connect.txtin the same format as in step 4. The status code is the same as in step 4, buttimeoutmust be different. - Calling
http://127.0.0.1:8088/app/bighdr?n=8000gives a 502. After confirming the phrase that names the cause, enlarge the proxy buffers so that it returns 200. If you enlarge only one directive, nginx won't start, so read the emerg message and fix them together. Create/root/ngi/case-header.txt. It has four lines.
Even after the fix, the failures in steps 3–5 must still reproduce as before.before=<고치기 전 코드> after=<고친 뒤 코드> phrase=<error.log 에서 원인을 지목하는 구절 그대로> fix=<헤더 크기를 직접 정하는 지시자 이름> - Create
/root/ngi/verdict.csv. The first line iscase,status,fix, followed by four lines.caseisrefused,read-timeout,connect-timeoutandbig-header, andfixis one ofnone,proxy_connect_timeout,proxy_read_timeoutandproxy_buffer_size. Decide which pairs with which based on what you saw in steps 3–6. - Write
/root/ngi/rca.md. It must have five h2 headings,## 현상,## 증거,## 원인,## 조치and## 재발방지, and the body must quote as they are theelapsednumbers from steps 4 and 5 and the names of the two timeout directives.
Notes
- Start:
nginx -c /root/ngi/nginx.conf -p /root/ngi - Syntax check:
nginx -t -c /root/ngi/nginx.conf -p /root/ngi - Re-apply:
nginx -s reload -c /root/ngi/nginx.conf -p /root/ngi - Measuring time:
curl -s -o /dev/null -w '%{http_code} %{time_total}\n' <URL> - Starting the upstream:
cd /root/ngi && nohup python3 upstream.py > up.err 2>&1 & - Common mistake 1: leaving
error_logat theerrorlevel. All the warning lines in this course disappear. - Common mistake 2: seeing a 502 and raising the timeout first. In a failure where the connection was refused there was no time to wait at all.
- Common mistake 3: forgetting the slash after
proxy_pass, so that the upstream path becomes/app/slow. As in the example above, the address must end with a root slash for the path to become/slow.
Start the failure generator
Save /root/ngi/upstream.py and start it in the background.
The app comes up on port 9101 and port 9109 becomes a port that does not accept connections.
Nothing is started on port 9999.
Check: / on port 9101 returns 200, /slow?d=2 returns 200 after 2 seconds,
and the /bighdr?n=8000 response has an X-Session-Blob header.
You need three upstreams — an app that answers normally, a port that never accepts connections, and a port that nobody listens on. Start the first two together from one Python file. After starting, check each: that the app returns 200, and that the port that doesn't accept connections really hangs.
Set up a proxy that leaves evidence
Write /root/ngi/nginx.conf and start nginx.
Listen on port 8088 and set up three paths — proxy /app/ to port 9101,
/dead/ to port 9999 and /hang/ to port 9109.
Put error_log at /root/ngi/logs/error.log at warn level,
put the access log at /root/ngi/logs/access.log, and include all of
$upstream_response_time, $request_time and $upstream_status in its format.
Set proxy_connect_timeout 2s; and proxy_read_timeout 3s;.
Check: http://127.0.0.1:8088/app/ must return 200.
(You can open it with the browser preview.)
The port is 8088. The two logs are the whole point of this lab, so get the paths and levels exactly right — if you set error_log to the error level, the warnings in the later labs vanish entirely. The access log format must include three things: the time the upstream took, the total time, and the upstream status code.
A dead upstream — 502
Call http://127.0.0.1:8088/dead/ and create /root/ngi/case-refused.txt.
It has three lines.
status=<받은 상태 코드>
phrase=<error.log 에서 원인을 지목하는 구절 그대로>
timeout=<실제로 발동한 타임아웃 이름, 없으면 none>
See which code you get when you proxy to a dead port, and what the first word in error.log is at that moment. If you ask yourself whether there is any room for a timeout to come into play here, you will know what to write in the timeout field.
A slow upstream — 504, read timeout
Call http://127.0.0.1:8088/app/slow?d=10, measure how many seconds it takes to be cut off,
and create /root/ngi/case-read.txt. It has four lines.
status=<상태 코드>
phrase=<error.log 의 while 로 시작하는 구절 그대로>
timeout=<발동한 타임아웃 지시자 이름>
elapsed=<끊기기까지 걸린 초, 정수>
Call the path that takes 10 seconds to respond, but measure how many seconds it takes to be cut off. That number of seconds is written as is somewhere in the configuration file. If you read what comes after while in the phrase in error.log, you can tell for certain which timeout it is.
The timeout you raised was not the timeout that fired
Inside the /hang/ block, put proxy_read_timeout 30s;, reload, and then
call http://127.0.0.1:8088/hang/. Even though you gave the read 30 seconds,
the request is cut off much sooner. Measure the time and create /root/ngi/case-connect.txt
in the same format as in step 4.
The status code is the same as in step 4, but timeout must be different.
First put proxy_read_timeout 30s inside the /hang/ block and reload. You have declared that you will give the read 30 seconds. Even so, measure how many seconds it takes for the request to be cut off, and see how the phrase after while in error.log differs from step 4. The status code and the error number are the same.
The server is fine but only certain users get a 502
Calling http://127.0.0.1:8088/app/bighdr?n=8000 gives a 502.
After confirming the phrase that names the cause, enlarge the proxy buffers so that it returns 200.
If you enlarge only one directive, nginx won't start, so read the emerg message and fix them together.
Create /root/ngi/case-header.txt. It has four lines.
before=<고치기 전 코드>
after=<고친 뒤 코드>
phrase=<error.log 에서 원인을 지목하는 구절 그대로>
fix=<헤더 크기를 직접 정하는 지시자 이름>
Even after the fix, the failures in steps 3–5 must still reproduce as before.
There is a path that bloats the response header to 8000 bytes. See which code you get when only the header has grown. When you fix it, enlarging only one buffer directive keeps nginx from starting at all — read what that emerg message says to enlarge along with it.
A four-line verdict table
Create /root/ngi/verdict.csv. The first line is case,status,fix, followed by
four lines. case is refused, read-timeout, connect-timeout and
big-header, and fix is one of none, proxy_connect_timeout,
proxy_read_timeout and proxy_buffer_size.
Decide which pairs with which based on what you saw in steps 3–6.
Collect the four failures into one table. In the fix column, write the name of the directive in the nginx configuration that actually resolves that failure; but there is one that nginx configuration cannot resolve — write none on that line.
An incident report with numbers in it
Write /root/ngi/rca.md. It must have five h2 headings, ## 현상, ## 증거, ## 원인,
## 조치 and ## 재발방지, and the body must quote as they are
the elapsed numbers from steps 4 and 5 and the names of the two timeout directives.
The value of the report is in the numbers. Quote exactly the seconds you measured in the earlier steps. For prevention, instead of writing 'be careful', write which directive you will set to which value and which log phrase you will attach an alert to.