TT Lab
Get started
Learn Learning paths Courses

Testing Tools in Practice

Coverage Is a Map, Not a Target

Continue in TT Lab

Summary

Coverage is a map that tells you where you have not looked. The moment you make 100% the goal, the map becomes squares to fill in, and from then on tests that verify nothing multiply.

Why this matters

You have probably seen a test like this.

def test_create_user():
    u = create_user("kim")
    assert u is not None

This test runs every line of create_user but checks nothing. Coverage reaches 100%, and the test passes even if the name is not saved.

Conversely, the truly risky places are usually outside coverage: error-handling branches, boundary values, concurrent execution. Those are the places where writing tests was too much trouble, and that is exactly why bugs live there.

What to test first

Decide priority by the cost of being wrong.

Priority What Why
1 Money, permissions, data deletion If wrong, it cannot be undone
2 Pure functions with many branches Cheap to verify many cases
3 Boundary values and error paths Places people rarely write tests for
4 One integration path To check that the pieces fit together
5 The UI The most expensive and the most often broken

This order is what the test pyramid says: wide at the bottom (unit), thin at the top (E2E).

How to read coverage

Look at the missing lines, not the number.

pytest --cov=mymod --cov-report=term-missing
Name        Stmts   Miss  Cover   Missing
mymod.py       42      6    86%   17-19, 28, 51-52

Missing is the answer. If 17–19 is error handling, that is a place to fill in, and if 51–52 is logging, you do not have to fill it in. A person makes that judgment.

Turning on branch coverage makes it more accurate.

pytest --cov=mymod --cov-branch

If you pass through if x: only on the true side, line coverage is 100% but branch coverage is 50%.

Common misconceptions

"If the tests pass, it is correct": a test can only show that a bug exists; it cannot prove that none exists. Passing means no more than "it is correct for the cases I thought of".

"A slow test is still a test": if it is slow, nobody runs it. A test nobody runs is worse than no test (it makes people believe one exists). The whole unit test suite should finish within a few seconds.

Criteria for choosing what to test

You cannot test everything. You judge along two axes: the damage if it is wrong and the probability of being wrong.

Changes often Rarely changes
High damage Must test. This is priority 1 Test. The goal is regression prevention
Low damage Not necessary to test Do not test

Payment, authentication, and data deletion are in the upper left. UI wording and log formats are in the lower right. If you fill in the lower right just to raise the coverage number, only the maintenance cost grows and the accidents stay the same.

The test pyramid and its counterexample

The basic form puts the slow, expensive tests at the top and the fast, cheap ones at the bottom.

      /\      E2E — 느리다(분), 잘 깨진다, 그러나 진짜를 본다
     /      /----\    통합 — DB·큐를 실제로 띄운다(testcontainers)
   /        /--------\  단위 — 밀리초, 로직만

But this shape is not always right. In services with thin logic and thick integration (CRUD APIs, data pipelines), unit tests hardly pay off. If you imitate the DB with mocks, you miss the very SQL errors that matter. In those places it is right to build up a thick layer of integration tests.

More important than the shape is "if this test breaks, is it a real problem?" A test that breaks with only a small change to the implementation blocks refactoring. Such a test is worse than none.

Three properties of a good test

It is deterministic. The same input gives the same result. If it relies on time, random numbers, or order, it fails intermittently, and an intermittent failure soon gets ignored. Inject the time (a clock argument) and fix the seed of random numbers.

It is independent. It must pass even if you change the order or run just one. If it depends on data left by an earlier test, parallel execution becomes impossible.

The intent is in the name. Not test_1 but test_returns_404_when_order_belongs_to_ another_user. You should be able to tell what broke just by reading the list of failures.

# 시각을 주입하면 결정적이 된다
def is_expired(token, now=None):
    now = now or datetime.now(timezone.utc)
    return token.exp < now

def test_expired_token_is_rejected():
    t = Token(exp=datetime(2026, 1, 1, tzinfo=timezone.utc))
    assert is_expired(t, now=datetime(2026, 1, 2, tzinfo=timezone.utc))

What really matters in practice

The value of a test is whether the cause is immediately visible when it fails.

# 나쁨 — 왜 틀렸는지 모른다
assert result == expected

# 좋음 — 무엇이 다른지 보인다
assert result.status == 200, f"응답: {result.status} {result.body[:200]}"

pytest automatically shows both sides of assert a == b. So splitting a condition into small pieces is better than writing a long message.