TT Lab
Get started
Learn Learning paths Courses

It failed, but the exit code was 0

Pin the bug down with a test

Continue in TT Lab

Summary

Reproduce a bug in an operations tool with a failing test first, and only then fix it. For tools that depend on files, time, environment variables, and output, use pytest's tmp_path, monkeypatch, and capsys to confine those dependencies inside the test.

Why this matters

A log rotation tool was supposed to "keep only 3", but 4 remained. The owner fixed the code, and the following week the same bug came back in a different form: this time a file was deleted even with --dry-run. What the two accidents have in common is that the person who fixed them left no evidence that they were fixed. The evidence is a test. A test that reproduces the bug is red before the fix and green after it, and from then on it acts as a gatekeeper that keeps the same bug from returning.

How it works

pytest finds the test_*.py files, runs the test_* functions in them, and records a failure when an assert is false. To test a tool, the tool has to be importable. This is where the insistence on the main(argv) -> int shape in the earlier module pays off. If you call it as import rotate; rotate.main([str(d), "--keep", "3"]), you can test the whole tool without starting a new process.

Files. The tmp_path fixture gives each test function its own temporary directory as a pathlib.Path. You create files in it, run the tool, and look at the result, and pytest cleans up when the test ends. A test that touches the real /var/log is not a test; it is an accident.

Time. A condition such as "files older than 7 days" depends on time.time(). You cannot wait for real time to pass, so you swap it out with monkeypatch. When you replace an attribute as in monkeypatch.setattr(rotate.time, "time", lambda: 1_700_000_000), it is restored when that test ends. The documentation lists setattr, delattr, setitem, setenv, delenv, and chdir, and they all follow the same principle: change global state only inside the test.

Output. If the tool reports with print what it deleted, that sentence is also a contract. The readouterr() method of the capsys fixture returns the stdout and stderr printed during the test. You check it as in assert "delete old.log" in captured.out.

Several sets of input. If you want to run the same logic with keep set to 0, 1, 3, and 10, you do not copy the function four times. The parametrize decorator creates one test for each argument in the list. When one fails, the test ID (test_keep[3]) records which value failed.

import pytest, rotate

@pytest.mark.parametrize("keep", [0, 1, 3, 10])
def test_keep_leaves_exactly_keep_files(tmp_path, keep):
    for i in range(5):
        (tmp_path / f"app-{i}.log").write_text("x")
    assert rotate.main([str(tmp_path), "--keep", str(keep)]) == 0
    assert len(list(tmp_path.glob("*.log"))) == min(5, keep)

def test_dry_run_deletes_nothing(tmp_path, capsys):
    (tmp_path / "a.log").write_text("x")
    (tmp_path / "b.log").write_text("x")
    rotate.main([str(tmp_path), "--keep", "0", "--dry-run"])
    assert sorted(p.name for p in tmp_path.glob("*.log")) == ["a.log", "b.log"]
    assert "a.log" in capsys.readouterr().out

Order. When you receive a bug report, (1) write a test that reproduces it, which must fail when run against the current code. If it does not fail, the reproduction is wrong, and if you fix the code in that state you cannot tell what you fixed. (2) Fix the code. (3) The test passes. (4) All the other tests pass too, which is evidence that the fix did not break anything else. Skipping step (1) of these four is the reason "the same bug comes back next week".

Leaving results behind. CI reads files, not human eyes. pytest --junitxml=report.xml leaves the number of tests, the number of failures, and the elapsed time as XML, and that becomes the pass condition of the pipeline. There is also an exit code: 0 if everything passes and 1 if even one fails.

What it looks like in the field

The most common case is "the test looks at a real directory". It calls os.chdir or uses a fixed path such as /tmp/test, so two tests interfere with each other and the result depends on the order. tmp_path gives a different path for every test, so this problem does not occur. The second is a test that creates time with time.sleep(2): it is slow and fails when CI is busy. Time is something you swap out. The third is "I fixed it, so the test can come later". Later never comes. If you write the failing test first, the time spent fixing actually shrinks, because you already have the reproduction in hand.

What you will do in the next lab

You copy the log rotation tool /opt/fixtures/pyops/buggy/rotate.py, which has two planted defects, into your working directory and proceed in this order: a pure function test, a reproduction (failing) with tmp_path, the fix, a --dry-run reproduction, the fix, parametrize, capsys, freezing time with monkeypatch, and a --junitxml report. The grader also runs your tests against the original defective code to confirm that they really catch the bugs.