TT Lab
Get started
Learn Learning paths Courses

I Break It — A Chaos Lab Where the Hypothesis Comes First

If "normal" isn't written as a number, it's just an incident

Continue in TT Lab

One-line summary

Before you break anything, "normal" must be written down as numbers. Without those numbers, an experiment cannot be told apart from an incident.

Why this is needed

The phrase heard most often in production retrospectives is "it was a bit slow back then." Nobody can say how slow "a bit slow" was, or what it was normally. What happens if you inject a failure in that state? The screen turns red, people rush in, someone hits rollback, and afterward nobody can say what was learned. That is not an experiment; it is just an incident.

What sets chaos engineering apart is not the breaking. It is that before breaking anything, you write down first what you will measure, what you believe will happen, and when you will stop. That is why the Principles of Chaos Engineering put "define the system's normal behavior as a measurable output" as their first item. They insist that you define it not by internal state (CPU utilization, heap size) but by output visible from the outside — the proportion of requests that succeed and the time it takes for responses to come back — because internal metrics sometimes wobble when there is no failure, and sometimes look fine during a real failure.

How it works

Steady state is usually written along two axes: availability and latency. These are different axes, so you write them together. Even if 99.9% of requests succeed, users feel that something is broken if the p95 response is 3 seconds, and conversely, if responses take 20 milliseconds but one in twenty fails, that is also broken. If you look at only one axis, you miss half of the experiment's result.

You write latency not as an average but as a quantile. An average dilutes the slow tail. If five out of a hundred requests take 2 seconds and the rest take 20 milliseconds, the average is 120 milliseconds and looks fine, but the p95 is 2 seconds, which shows the problem as it is. Users do not experience the average. They experience their own single request.

정상 상태 기술의 예 (이 코스의 실습에서 실제로 쓰는 형식)
  availability_ratio_min : 0.98      20초 동안 정적 경로 200회 중 196회 이상 성공
  p95_ms_max             : 70        CPU 를 쓰는 경로의 95분위 응답이 70밀리초 이하

You also write down an abort condition. It is a sentence such as "if the success rate drops below half, roll back immediately." During an experiment, judgment gets clouded. If you just watch a little longer, the cause seems about to appear, and rolling back means setting everything up again, which is a hassle. For that moment you need a sentence written while you were calm.

Finally, you set the blast radius: only one replica, only inside one namespace, only 1% of traffic. An experiment only needs to be as large as it takes to confirm the hypothesis; anything bigger adds risk with nothing gained.

You also decide the length of the measurement window and the number of samples in advance. If the window is too short, you miss a fast-recovering failure entirely, and if there are too few samples, a single failure shifts the success rate by 5%, so anything you look at seems significant. The lab in this course collects 200 samples by sending 10 availability requests per second over a 20-second window. One failure corresponds to 0.5%, so the 0.98 boundary means "up to four failures are acceptable," which makes the interpretation clear. Knowing how many requests a boundary number corresponds to, rather than just the number itself, builds the sense for reading an experiment.

What you see in the field

In Kubernetes, "normal" is split into several layers, so you need to be extra careful. A Pod being Running and being Ready are different facts, and being Ready and being carried in a Service's endpoints to receive traffic is yet another fact. A container's readiness is decided by the readinessProbe, and the result shows up as the Pod's conditions. Even with three green lights on, a user's request can fail.

That is why in practice you do not take cluster state as your "steady state." You take the result of sending real requests from the outside as the steady state, and use cluster state as supporting evidence that explains it. If you reverse this order, you get the familiar scene where every dashboard is green and only the customer support desk gets busy.

There is one more mistake commonly made when measuring a baseline: taking a state that has not yet settled as the baseline. Right after a deployment, caches are empty and images have just been pulled, so responses are slower than usual. If you take those numbers as the baseline, every comparison made afterward becomes loose. You must follow two rules: measure after the rollout has finished and the probes have stabilized, and touch nothing while measuring the baseline. That is also why this lab's recording tool refuses to record if a new Pod appears within the baseline measurement window.

One more thing: the steady state must be a statement the team has agreed on. A standard set by one person will always wobble when an incident happens. You must be able to explain where the number "p95 70 milliseconds" came from, and why it is not 60 or 100, and that explanation is usually a combination of measurement and judgment such as "I measured it normally and got 33 milliseconds, and users cannot feel up to double that." A standard that has only a number and no rationale changes next quarter for no reason.

What to check in the next quiz

Check why steady state is defined by output visible from the outside, why latency is written as a quantile rather than an average, and what role the abort condition and the blast radius each play in an experiment.