The Limit Went Up and Nobody Raised It
In one line
An audit trail is not about storing values but about leaving records in a form that can be disputed later. You keep an append-only record instead of in-place UPDATEs, make the same event always the same bytes with canonical serialization, and then add a hash chain and a key to give it the property that "tampering in the middle shows."
Why this was needed
The moment audit response becomes hard at a bank is usually not when a value is wrong. It is when the value is right but nobody can explain how it got there. Consider a system where the credit limit table is updated with the single line UPDATE customer_limit SET limit_amount = ? WHERE customer_id = ?. The moment this statement runs, the old value disappears. All that remains are the updated_at and updated_by columns, and even those are only the trace of the last time, so you cannot tell how many times it passed through before that.
In the field, this problem shows up like this. At the month-end check, the total limit is 480 million won higher than at the beginning of the month. When you count the change applications, the number does not match. For some, there is no application at all, and for some, the approved amount written on the application differs from the current value. If you ask the person in charge, you usually get an honest answer — "we reflected them in bulk with a batch." What that batch did is recorded nowhere. The most expensive thing in an audit is not a wrong number but a number you cannot explain.
What the US NIST SP 800-92 log management guide says repeatedly about handling logs is the same thing (full PDF). Logs do not end with generation; generation, retention, protection, and verification must be designed as one bundle, and a log whose integrity is not kept loses much of its value as evidence. The concrete retention periods in Korean financial practice differ by institution, so this lab does not cover them and focuses only on how to build records so that you can prove them later.
How it works
First, you keep an append-only table. You keep separately a table that overwrites state (the current limit) and a table that accumulates events (the audit record). The current value is a derivative for fast lookup, and the truth is the list of events.
Second, you make events the same bytes. A hash is computed over bytes, so {"a":1,"b":2} and {"b":2,"a":1} give different hashes even though they are the same event. RFC 8785 JSON Canonicalization Scheme pinned this problem down as a specification. It sorts object keys, removes whitespace around delimiters, and serializes as UTF-8. In Python, if you turn on the three options sort_keys=True, separators=(",", ":"), and ensure_ascii=False of the json module, you get the same result for most practical payloads. However, RFC 8785 defines key sorting by UTF-16 code units and Python sorts by code points, so they diverge if you use supplementary-plane characters such as emoji as keys. In practice, it is safer to restrict keys to ASCII.
Third, you link them into a chain. Each entry carries the previous entry's hash as prev_hash, and hashes its own content together with prev_hash to make entry_hash (hashlib). If you fix a line in the middle, that line's entry_hash goes off, and if you recompute that line's hash and slot it in, the prev_hash of the later entry goes off. A chain does not make things "impossible to fix." It makes them show when fixed.
Fourth, you put a key on top of the chain. With only a hash chain, someone who can access the whole record can recompute from start to finish and swap it out wholesale. If you make a MAC for each entry with the hmac module and keep that key somewhere other than the record, someone who does not know the key cannot recompute. If the key is in the same DB, this property disappears entirely.
app_log(원본) ──▶ canon(사건) ──▶ sha256 ──▶ entry_hash ──┐
▲ │ prev_hash 로 다음 항목에
└──────── prev_hash ◀────┘
entry_hash + prev_mac ──▶ HMAC(열쇠) ──▶ entry_mac
What it looks like in the field
The most common failure is a log without a correlation ID. One limit change passes through the teller app, the limit service, the core batch, and the audit relay in turn, but each system leaves only its own events. When you later ask "where did this change start," the only way is to join lines with similar times by eye. If you attach an ID to a request and pass it through the whole path, the investigation time shrinks from hours to minutes. Conversely, if one system drops that ID midway, everything after that point goes back to guessing.
The second is trusting an export copy received from a vendor as it is. That a chain is attached does not mean it was verified. If you recompute the received file from the start, the places where only the content was fixed and the places where entries are missing entirely come to light. The two have different symptoms. If you fix the content, that line's hash goes off, and if you remove an entry, the link of the next line goes off.
What really matters in practice
- A record must be recomputable. Do not store only the judgment result; leave the rules and inputs used in the computation.
- Verification must be idempotent. Verifying the same file twice must give the same conclusion, and verification must not fix the record.
- Keep the key somewhere other than the record, with file permission 600. If you keep it inside the same DB, the chain becomes decoration.
- When you hand over evidence, hand it over as a bundle. Put the file list, each file's hash, the chain's head hash, and the verification result in one document.
What you will do in the next lab
You build a limit snapshot of Handeul Bank (fictional) yourself, and bring out in numbers the changes made without an application and those reflected differently from the approved value. Then you build a canonical serialization tool and fit it to the vectors the grader shuffles differently each time, and bind 400 log lines into a hash chain. You find the tampered places in a vendor export copy, stop recomputation forgery with HMAC, regroup events per request with the correlation ID, and produce an evidence bundle with hashes attached.