TT Lab
Get started
Learn Learning paths Courses

Debugging in Practice

When the Documentation and the Code Disagree, Which Do You Believe

Continue in TT Lab

One-line summary

The most dangerous defect in customer code is not the code that throws an exception but the code that quietly returns a plausible value.

Why this is needed

A bug that blows up is easy. The traceback tells you where the problem is, monitoring sounds the alarm, and nobody trusts the result.

The scary one is a bug that returns 200. An API where /sum?a=2&b=3 returns -1 has a normal status code, so it is green on the dashboard, there are no errors in the logs, and it is discovered months later when an accountant says the numbers do not add up. By then the wrong value has already been copied to several downstream systems.

The code an FDE takes over from a customer is usually in this state. The original author has left, there are no tests, and there is documentation, but it differs from the code.

How it works

When the documentation and the code differ, which should you trust? The answer is to trust neither, and to trust the observed behavior. But you do not throw the documentation away. The documentation is the only clue to the "original intent," and the places where the code and the documentation disagree are your list of bug candidates.

The practical procedure goes like this.

1. Turn the spec into an input/output table. Write a few lines from what the documentation says, in the form "for this input, this output." If the documentation does not say, ask the customer. Without this table, you have no criterion for deciding what is a bug.

2. Try boundary values. 0, negatives, a single element, empty values, a missing required parameter. Most defects show up not at average inputs but at the boundaries. If avg is right with three elements and wrong with one, the divisor was written wrongly.

3. Distinguish the kinds of failure. Returning 500 for a bad request and returning 400 are completely different problems. 400 means "you sent it wrong," and 500 means "I broke while handling it." When this distinction collapses, downstream retry logic misbehaves. Clients treat 5xx as retryable, so they end up retrying a request that will fail forever, endlessly.

4. Suspect the points where data gets involved. It is very common for the code to be right while one malformed row mixed into the input data kills the whole run. The spec says "skip invalid rows," but the code has no such branch.

What you see in the field

A handed-over batch script often arrives with the customer's guess that "it seems to have slowed down because the data grew." When you open it, it is not slow but dying outright, and the cause is not that the data grew but that rows of a form never seen before got mixed into the newly arrived data.

This is where the FDE's judgment splits. If you delete that row and move on, the same thing repeats next week. If you fix the code to skip it as the spec says and make it log the skipped rows, next time the cause comes out in one minute.

You do not need to flatly deny the customer's guess. Put it on the verification list and check it together. If you reject the customer's hypothesis at the door, at the next outage that person will not tell you what they observed.

What to do after the fix

Finding and fixing the defect is not the end of the job. Having fixed one bug in handed-over code is a signal that more bugs of the same kind are likely. There is no reason the original author made that mistake only once.

So right after the fix, you do three things.

1. Find every instance of the same pattern. If you found one place where the divisor was written wrongly, open every other aggregate function in the same file. If you found one except: pass that swallows exceptions, search the whole repository for that form. This work usually takes a few minutes and almost always turns up a second and a third.

2. Leave a test that catches the bug. Not a test that shows the fixed code is right, but a test that would have failed before the fix. This distinction matters. The former merely records the current state, and only the latter prevents regression. If you write the test first and see with your own eyes that it fails before fixing, this distinction is kept automatically.

3. Write down why nobody knew until now. The answer to this question tells you what to fix next. If the status code was 200 and so the dashboard was green, that is not a code defect but an observability defect. It means there was no way at all to notice when a calculation result was wrong, and if so, you need assertions that check value ranges or a reconciliation procedure downstream.

There is also an order when reporting to the customer. What was wrong, since when, how far the impact extends, what was fixed, and what was added so it does not happen again. Of these, the one people are most curious about and FDEs most often leave out is the third. If the wrong values have already flowed to other systems, undoing that is far bigger than the code fix, and that judgment is the customer's to make. If you do not know the extent of the impact, it is always better to write that you do not know and propose how to find out than to slide past it quietly and have it discovered later.

What you will do in the next lab

Against a calculation API taken over without any handover, you compare the spec written in the documentation with the actual behavior, find and fix four defects, and make the right status codes go out for bad requests.