TT Lab
Get started
Learn Learning paths Courses

Network Troubleshooting

Full Triage: A Payments API Outage

Continue in TT Lab

Goal

You narrow down an outage with three overlapping faults layer by layer to pin down the causes, fix them, and leave a postmortem report in the prescribed format. You use all the tools you learned in this course at once.

Why it matters

Real outages often have more than one cause. If you fix one and say "it isn't better" and change direction, you start suspecting even what you already fixed. So it is ultimately faster to check each layer to the end and write down everything you find. And after fixing, you finish the action only when you confirm the success with the very command that reproduced the failure.

Steps

  1. Reproduce the outage situation with bash /opt/fixtures/nt-triage/start-incident.sh and read /opt/fixtures/nt-triage/incident.md. Create the /root/triage directory.
  2. Save the result of a 3-count ping to your own interface IP to /root/triage/l3.txt. Loss must be 0%.
  3. Knock on ports 9101 and 9102 of your own interface IP separately, and write the results on two lines in /root/triage/l4.txt. 9101=<open|refused|timeout> / 9102=<open|refused|timeout>
  4. Check the name resolution result of pay-api.labhub.local and write it on two lines in /root/triage/name.txt. RESOLVED=<getent 가 돌려준 IP> / SOURCE=<그 값이 적혀 있는 파일의 절대 경로> (RESOLVED= followed by the IP getent returned, and SOURCE= followed by the absolute path of the file where that value is written)
  5. Check the listening address of the 9101 service and write it on one line in /root/triage/bind.txt. The format is BIND=<리슨 주소:포트> (BIND= followed by the listening address:port). (For example: BIND=127.0.0.1:9101)
  6. Request /ok on the reachable side (9102) and write the status code on one line, code=<코드> (code= followed by the code), in /root/triage/http.txt.
  7. Fix two things.
    • Correct the pay-api.labhub.local entry in /etc/hosts to this server's real interface IP.
    • Restart the 9101 service so that it listens on 0.0.0.0. (Terminate the existing process) Then confirm that curl -s http://pay-api.labhub.local:9101/ok succeeds and save its output to /root/triage/fixed.txt.
  8. Make /root/triage/report.txt with the following 6 lines. SYMPTOM=remote-timeout / LAYER1=name / LAYER2=bind / CAUSE_NAME=<원래 hosts 에 적혀 있던 잘못된 IP> / CAUSE_BIND=127.0.0.1 / FIXED=yes (CAUSE_NAME= is followed by the wrong IP that was originally written in hosts)

Notes

Reproducing the outage situation

Reproduce the outage situation with bash /opt/fixtures/nt-triage/start-incident.sh and read /opt/fixtures/nt-triage/incident.md. Create the /root/triage directory.

If you run the fixture's start script, the situation is created. Read the report as well.

Checking L3 reachability

Save the result of a 3-count ping to your own interface IP to /root/triage/l3.txt. Loss must be 0%.

You must check by IP, not by name, so that it does not mix with the name problem. Use your own interface address.

Judging the result per port

Knock on ports 9101 and 9102 of your own interface IP separately, and write the results on two lines in /root/triage/l4.txt. 9101=<open|refused|timeout> / 9102=<open|refused|timeout>

The results of the two ports differ. Write refused, timeout, and open exactly distinguished.

Pinning down the name resolution problem

Check the name resolution result of pay-api.labhub.local and write it on two lines in /root/triage/name.txt. RESOLVED=<getent 가 돌려준 IP> / SOURCE=<그 값이 적혀 있는 파일의 절대 경로> (RESOLVED= followed by the IP getent returned, and SOURCE= followed by the absolute path of the file where that value is written)

Compare the IP getent returns with this server's real IP. You also have to find where that value is written.

Pinning down the binding problem

Check the listening address of the 9101 service and write it on one line in /root/triage/bind.txt. The format is BIND=<리슨 주소:포트> (BIND= followed by the listening address:port). (For example: BIND=127.0.0.1:9101)

The Local Address column of the ss output is the answer. Tell which of the two services is the problem.

Checking the application layer

Request /ok on the reachable side (9102) and write the status code on one line, code=<코드> (code= followed by the code), in /root/triage/http.txt.

Check the reachable side first to prove that the service itself is fine.

Fix and verify

Fix two things.

You must fix both. After fixing, be sure to check again on the outside address.

Postmortem report

Make /root/triage/report.txt with the following 6 lines. SYMPTOM=remote-timeout / LAYER1=name / LAYER2=bind / CAUSE_NAME=<원래 hosts 에 적혀 있던 잘못된 IP> / CAUSE_BIND=127.0.0.1 / FIXED=yes (CAUSE_NAME= is followed by the wrong IP that was originally written in hosts)

The format is fixed. Each value must be something you actually confirmed in the earlier steps.