TT Lab
Get started
Learn Learning paths Courses

In Front of an Unfamiliar System

Pinning the Symptom to a Number

Continue in TT Lab

Goal

You will be able to turn a report of "it doesn't work sometimes" into a measured value and a falsifiable hypothesis within 20 minutes.

Why it matters

Customers do not give you bug reports. They give you pain reports. Turning that into a measurable problem is entirely your job, and the first tool for it is the reproduction command. The moment you pin the failure rate down as a number, two things come into being. First, you see the same facts as the customer. Second, you get a yardstick for saying "it is fixed" later. What was fixed without a yardstick cannot be known to be fixed.

Next is the hypothesis, and there is one rule here. If what you would see when it is wrong is not decided, it is not a hypothesis. "It seems like a network problem" survives whatever result comes out, so it cannot reduce the candidates. Only a sentence that can die narrows the investigation scope. And if you do not write down what you ruled out, 30 minutes later you find yourself rechecking a candidate you have already erased.

Steps

  1. Create the /root/hypothesis directory.
  2. Run /opt/app/flaky.py so that it responds on 127.0.0.1:8001.
  3. Call /quote 20 times in a row and write the number of non-200 responses to /root/hypothesis/baseline.txt.
  4. Write that failure rate as a percentage integer to /root/hypothesis/rate.txt. Only the number, without a symbol.
  5. Save the body string of a failed response to /root/hypothesis/reason.txt.
  6. In /root/hypothesis/h1.md, write a hypothesis that contains three lines, 가설:, 맞다면:, and 틀리면: (the Korean words for "hypothesis", "if correct", and "if wrong", each followed by a colon).
  7. In /root/hypothesis/ruled_out.txt, write the layers you ruled out. Each line starts with one of network, auth, app, and data, followed by the evidence. At least two lines.
  8. Call /quote 40 times in a row and write the number of failures to /root/hypothesis/confirm.txt.

Notes

Create the hypothesis working directory

Create the /root/hypothesis directory.

Gather the ruled-out list and the hypothesis document in one place. Later they become half of the report.

Bring up the quote API

Run /opt/app/flaky.py so that it responds on 127.0.0.1:8001.

If you run /opt/app/flaky.py with python3, it waits on 127.0.0.1:8001. Add & at the end so that it does not hold the shell.

Call 20 times and count the failures

Call /quote 20 times in a row and write the number of non-200 responses to /root/hypothesis/baseline.txt.

The combination of curl's -o /dev/null -w '%{http_code}' prints only the status code. Loop 20 times with seq and count the ones that are not 200.

Write the failure rate as a percentage

Write that failure rate as a percentage integer to /root/hypothesis/rate.txt. Only the number, without a symbol.

This is the step that turns an adjective into a number. Convert how many out of 20 into a percentage integer and write it without a symbol.

Capture the body of a failed response

Save the body string of a failed response to /root/hypothesis/reason.txt.

From the status code alone you do not know why it failed. Drop -o /dev/null and take the response body as is.

Write a falsifiable hypothesis

In /root/hypothesis/h1.md, write a hypothesis that contains three lines, 가설:, 맞다면:, and 틀리면: (the Korean words for "hypothesis", "if correct", and "if wrong", each followed by a colon).

You need three lines: the hypothesis, if correct, and if wrong. If the third line does not get written, that is not yet a hypothesis but a feeling.

Record the layers you ruled out

In /root/hypothesis/ruled_out.txt, write the layers you ruled out. Each line starts with one of network, auth, app, and data, followed by the evidence. At least two lines.

Of network / auth / app / data, write at least two layers that you have checked and erased so far, each with why you erased it.

Measure again with the same yardstick

Call /quote 40 times in a row and write the number of failures to /root/hypothesis/confirm.txt.

Whether it is random or regular is settled by enlarging the sample. Measure again with 40 and check that the ratio holds.