TT Lab
Get started
Learn Learning paths Courses

I Break It — A Chaos Lab Where the Hypothesis Comes First

Hypothesis, measurement and conclusion, each written at its own moment

Continue in TT Lab

One-line summary

The product of an experiment is not a failure but a note. That note is useful only when the hypothesis, the measurement, and the conclusion agree with one another.

Why this is needed

Injecting a failure is not hard in itself. One command is enough. The hard part is creating something that is still there next week. After an experiment ends, people's memories fade astonishingly fast, and even their direction changes. They remember a 76% success rate as "it almost all failed, didn't it," and conversely remember an experiment that clearly failed as "I think it held up well back then." That is why an experiment records its three pieces at different points in time in writing.

How it works

First, write the hypothesis before you break anything. If you write it afterward, it is not a hypothesis but a summary of the result. After seeing the result, people cannot tell the feeling of "I knew it" from an actual prediction. So the order itself is part of the method. In the hypothesis, you write not only what will happen but also why you expect it and when you will stop.

Second, measure inside the experiment window. If you measure separately after causing the failure, you measure numbers from after recovery has already started. The order must be: start the load first, cause the failure inside that window, and read the result after the window closes.

Third, the conclusion is the result of comparing the hypothesis with the measurement. There is one rule here that matters most. A refuted hypothesis is not a failure. On the contrary, a result different from what you expected is the most expensive piece of information that experiment can earn. If you wrote "raising replicas to three means not a single request will break" and in fact a few did break, what you learned that day is the fact that "the number of replicas alone is not enough." An experiment note with only confirmed written in its conclusions usually means the hypothesis was written afterward.

실험 한 건의 최소 구성
  가설      availability: ok / latency: slower        + 왜 + 중단 조건
  측정      성공률 0.993 · p95 291ms (기준선 33ms)
  결론      관측 ok / slower → 가설과 같음 → confirmed + 무엇을 바꿀 것인가

Keeping these three pieces in separate files is also part of the method. If you keep writing over a single file, it quickly becomes blurry when the hypothesis was written and which number belongs to which experiment. You do not touch the hypothesis file once the experiment has started, the measurement file is written in one piece by the tool together with the cluster at that moment, and the conclusion file is created separately at the very end. When someone reads this note later, being able to answer "when was this sentence written?" at the file level is most of what makes it trustworthy.

You also decide in advance the rule for naming numbers. Whether a success rate of 0.993 is "fine" or "worse" is read differently by different people. The lab in this course calls a success rate of 0.98 or higher ok, 0.5 or higher degraded, and anything below that down, and calls latency slower if it is at least twice the baseline p95. If the rules come first, the conclusion does not turn into a quarrel.

What you see in the field

An experiment note must also record its limitations. In this lab the tool creates the observation file, but that file is in the end editable text inside the VM. So grading does not look only at the numbers but also at the traces actually left in the cluster: the ReplicaSet newly created with each deployment, and the Pods that are alive now. Numbers can be made up, but a ReplicaSet left by a deployment that never happened cannot. It is the same in the field. Numbers in a report should always be written with a path by which they can be checked again.

The last column of the note is "so what will we change?" If this column is empty, the experiment ends as a spectacle. What you change could be code or configuration, but often it is an operational procedure such as "add one more alert criterion on the latency side." In particular, incidents where the success rate is fine but only latency gets worse are not caught by existing alerts at all, so if the experiment confirmed that, creating one more alert on the spot is the precise follow-up. Ideally five experiments yield five different decisions, and if the same sentence is written in all five columns, it usually means only one experiment was really done.

And when the experiment ends, you must always restore things to how they were, or to a better design that reflects what you learned. An experiment that is not reverted is handed to the next person simply as a failure. If you learn how to look into a running Pod, checking right after recovery also gets faster.

Once notes accumulate, the next step is automation. You run the experiments you used to do by hand at fixed times and make them stop by themselves when steady state is left, and the starting point is always one experiment done once by hand. Only when a human-readable note exists first can a machine later rewrite that note.

What you will do in the next lab

You run five experiments on a real k3s. You measure the steady-state baseline, write the hypotheses for all five experiments at once, and then kill a Pod, squeeze the CPU, and starve the memory. At the end, you fix the design to reflect what you learned and take the same attack again.