TT Lab
Get started
Learn Learning paths Courses

nginx Incident Response

Create 502 and 504 Yourself and Tell Them Apart by Evidence

Continue in TT Lab

Goal

You produce 502 and 504 yourself, and then learn to pin down the cause not from the status code but from the phrase in error.log. In particular, you learn to distinguish two cases that are the same 504 but hit different timeouts.

Why it matters

What eats the most time in incident response is being confident about the wrong cause. If you conclude that a 502 means "the backend is dead" and raise a timeout, nothing happens, because in a failure where the connection was refused there was no time to wait at all. Conversely, a 504 doesn't mean you can fix it by raising proxy_read_timeout either. If it was hit at the connection stage, even raising the read timeout to 30 seconds still cuts the request off after 2 seconds. Step 5 of this lab is exactly that scene. What makes that difference is just one phrase after while in one line of the log.

Steps

  1. Save /root/ngi/upstream.py and start it in the background. The app comes up on port 9101 and port 9109 becomes a port that does not accept connections. Nothing is started on port 9999. Check: / on port 9101 returns 200, /slow?d=2 returns 200 after 2 seconds, and the /bighdr?n=8000 response has an X-Session-Blob header.
  2. Write /root/ngi/nginx.conf and start nginx. Listen on port 8088 and set up three paths — proxy /app/ to port 9101, /dead/ to port 9999 and /hang/ to port 9109. Put error_log at /root/ngi/logs/error.log at warn level, put the access log at /root/ngi/logs/access.log, and include all of $upstream_response_time, $request_time and $upstream_status in its format. Set proxy_connect_timeout 2s; and proxy_read_timeout 3s;. Check: http://127.0.0.1:8088/app/ must return 200. (You can open it with the browser preview.)
  3. Call http://127.0.0.1:8088/dead/ and create /root/ngi/case-refused.txt. It has three lines.
    status=<받은 상태 코드>
    phrase=<error.log 에서 원인을 지목하는 구절 그대로>
    timeout=<실제로 발동한 타임아웃 이름, 없으면 none>
    
  4. Call http://127.0.0.1:8088/app/slow?d=10, measure how many seconds it takes to be cut off, and create /root/ngi/case-read.txt. It has four lines.
    status=<상태 코드>
    phrase=<error.log 의 while 로 시작하는 구절 그대로>
    timeout=<발동한 타임아웃 지시자 이름>
    elapsed=<끊기기까지 걸린 초, 정수>
    
  5. Inside the /hang/ block, put proxy_read_timeout 30s;, reload, and then call http://127.0.0.1:8088/hang/. Even though you gave the read 30 seconds, the request is cut off much sooner. Measure the time and create /root/ngi/case-connect.txt in the same format as in step 4. The status code is the same as in step 4, but timeout must be different.
  6. Calling http://127.0.0.1:8088/app/bighdr?n=8000 gives a 502. After confirming the phrase that names the cause, enlarge the proxy buffers so that it returns 200. If you enlarge only one directive, nginx won't start, so read the emerg message and fix them together. Create /root/ngi/case-header.txt. It has four lines.
    before=<고치기 전 코드>
    after=<고친 뒤 코드>
    phrase=<error.log 에서 원인을 지목하는 구절 그대로>
    fix=<헤더 크기를 직접 정하는 지시자 이름>
    
    Even after the fix, the failures in steps 3–5 must still reproduce as before.
  7. Create /root/ngi/verdict.csv. The first line is case,status,fix, followed by four lines. case is refused, read-timeout, connect-timeout and big-header, and fix is one of none, proxy_connect_timeout, proxy_read_timeout and proxy_buffer_size. Decide which pairs with which based on what you saw in steps 3–6.
  8. Write /root/ngi/rca.md. It must have five h2 headings, ## 현상, ## 증거, ## 원인, ## 조치 and ## 재발방지, and the body must quote as they are the elapsed numbers from steps 4 and 5 and the names of the two timeout directives.

Notes

Start the failure generator

Save /root/ngi/upstream.py and start it in the background. The app comes up on port 9101 and port 9109 becomes a port that does not accept connections. Nothing is started on port 9999. Check: / on port 9101 returns 200, /slow?d=2 returns 200 after 2 seconds, and the /bighdr?n=8000 response has an X-Session-Blob header.

You need three upstreams — an app that answers normally, a port that never accepts connections, and a port that nobody listens on. Start the first two together from one Python file. After starting, check each: that the app returns 200, and that the port that doesn't accept connections really hangs.

Set up a proxy that leaves evidence

Write /root/ngi/nginx.conf and start nginx. Listen on port 8088 and set up three paths — proxy /app/ to port 9101, /dead/ to port 9999 and /hang/ to port 9109. Put error_log at /root/ngi/logs/error.log at warn level, put the access log at /root/ngi/logs/access.log, and include all of $upstream_response_time, $request_time and $upstream_status in its format. Set proxy_connect_timeout 2s; and proxy_read_timeout 3s;. Check: http://127.0.0.1:8088/app/ must return 200. (You can open it with the browser preview.)

The port is 8088. The two logs are the whole point of this lab, so get the paths and levels exactly right — if you set error_log to the error level, the warnings in the later labs vanish entirely. The access log format must include three things: the time the upstream took, the total time, and the upstream status code.

A dead upstream — 502

Call http://127.0.0.1:8088/dead/ and create /root/ngi/case-refused.txt. It has three lines.

status=<받은 상태 코드>
phrase=<error.log 에서 원인을 지목하는 구절 그대로>
timeout=<실제로 발동한 타임아웃 이름, 없으면 none>

See which code you get when you proxy to a dead port, and what the first word in error.log is at that moment. If you ask yourself whether there is any room for a timeout to come into play here, you will know what to write in the timeout field.

A slow upstream — 504, read timeout

Call http://127.0.0.1:8088/app/slow?d=10, measure how many seconds it takes to be cut off, and create /root/ngi/case-read.txt. It has four lines.

status=<상태 코드>
phrase=<error.log 의 while 로 시작하는 구절 그대로>
timeout=<발동한 타임아웃 지시자 이름>
elapsed=<끊기기까지 걸린 초, 정수>

Call the path that takes 10 seconds to respond, but measure how many seconds it takes to be cut off. That number of seconds is written as is somewhere in the configuration file. If you read what comes after while in the phrase in error.log, you can tell for certain which timeout it is.

The timeout you raised was not the timeout that fired

Inside the /hang/ block, put proxy_read_timeout 30s;, reload, and then call http://127.0.0.1:8088/hang/. Even though you gave the read 30 seconds, the request is cut off much sooner. Measure the time and create /root/ngi/case-connect.txt in the same format as in step 4. The status code is the same as in step 4, but timeout must be different.

First put proxy_read_timeout 30s inside the /hang/ block and reload. You have declared that you will give the read 30 seconds. Even so, measure how many seconds it takes for the request to be cut off, and see how the phrase after while in error.log differs from step 4. The status code and the error number are the same.

The server is fine but only certain users get a 502

Calling http://127.0.0.1:8088/app/bighdr?n=8000 gives a 502. After confirming the phrase that names the cause, enlarge the proxy buffers so that it returns 200. If you enlarge only one directive, nginx won't start, so read the emerg message and fix them together. Create /root/ngi/case-header.txt. It has four lines.

before=<고치기 전 코드>
after=<고친 뒤 코드>
phrase=<error.log 에서 원인을 지목하는 구절 그대로>
fix=<헤더 크기를 직접 정하는 지시자 이름>

Even after the fix, the failures in steps 3–5 must still reproduce as before.

There is a path that bloats the response header to 8000 bytes. See which code you get when only the header has grown. When you fix it, enlarging only one buffer directive keeps nginx from starting at all — read what that emerg message says to enlarge along with it.

A four-line verdict table

Create /root/ngi/verdict.csv. The first line is case,status,fix, followed by four lines. case is refused, read-timeout, connect-timeout and big-header, and fix is one of none, proxy_connect_timeout, proxy_read_timeout and proxy_buffer_size. Decide which pairs with which based on what you saw in steps 3–6.

Collect the four failures into one table. In the fix column, write the name of the directive in the nginx configuration that actually resolves that failure; but there is one that nginx configuration cannot resolve — write none on that line.

An incident report with numbers in it

Write /root/ngi/rca.md. It must have five h2 headings, ## 현상, ## 증거, ## 원인, ## 조치 and ## 재발방지, and the body must quote as they are the elapsed numbers from steps 4 and 5 and the names of the two timeout directives.

The value of the report is in the numbers. Quote exactly the seconds you measured in the earlier steps. For prevention, instead of writing 'be careful', write which directive you will set to which value and which log phrase you will attach an alert to.