TT Lab
Get started
Learn Learning paths Courses

The Language of Banking

Finding the Bad Row After Loading Is Already Too Late

Continue in TT Lab

In one line

A settlement or transfer instruction file received from a counterparty institution is checked against a format contract before it is loaded, and if even one line is off, the whole file is sent back. Rejecting at the door is always cheaper than reconciling after loading.

Why this was needed

A bank's day begins with a file and ends with a file. It receives and loads the day's instruction file that the counterparty institution uploaded by FTP in the early morning, and in the evening it returns our side's results in the same way. The way incidents occur on this route is always similar. The file arrived, the load batch ended "successfully," and yet the count in the trailer the counterparty sent differs from the count we loaded.

What comes next is the real problem. The lines that already went in have been picked up by other batches. The balances changed and notifications went out. To undo it, you must write correcting entries and recount which lines went in and which were left out. A job that would have taken 5 minutes if you had not accepted one file becomes a day's reconciliation and apology calls.

So a gatekeeper must stand on the receiving path. The gatekeeper's job is simple. It looks only at whether this file kept the format contract we agreed on. If not, it loads not a single line and sends it back with the reason attached. There is no option of loading only half.

How it works

A format contract usually has four layers.

First, the file's outer shape. Name rules, encoding, line endings. In domestic financial-sector external files, fixed-width files in the EUC-KR family still remain in large numbers. In Python, you use the cp949 codec, and the standard encodings table lists this codec's aliases as 949, ms949, and uhc and classifies the language as Korean. The difference that one Korean character is 2 bytes in CP949 and 3 bytes in UTF-8 leads straight into the next layer.

Second, the record structure. For fixed width, you look at exactly how many bytes a line is. The mistake that happens most often here is padding by character count. If the recipient name field is 20, that is not twenty characters but twenty bytes. If you treat "Kim Seojun" (a three-character Korean name) as three characters and pad with seventeen spaces, that line deviates from the specified width, and the fields after it shift wholesale. As the Unicode HOWTO explains, the length of a Python string is the number of code points, and what is written into a file is the encoded bytes. The cutting positions must be on bytes, not str.

For CSV, RFC 4180 is the standard. This document's Category is Informational, and it states itself that it "does not specify an Internet standard of any kind," but it is widely used as a de facto consensus. The core is three lines. A field containing a line break, a double quote, or a comma is enclosed in double quotes. A double quote inside an enclosed field is written as two double quotes. A record ends with CRLF. So the moment you cut with line.split(","), a line whose recipient name contains a comma gets one more field, and a fragment of the name lands in the account position. If you leave it to the csv module, commas and line breaks inside quotes are not counted as delimiters.

Third, the header and the trailer. You match the count and total written in the T record at the end of the file against the values you actually counted. This is the cheapest device for catching truncation in transit. If the count matches but the total differs, it is not truncation but a changed value, and since they are different incidents, you keep separate error codes for them.

Fourth, field-level specifications and duplicates. You look at the shape of the transaction number, the account format, and the amount range. And the same file coming twice is not an exception but routine. When the counterparty does not receive our response, it uploads the same file again. If the file's sha256 is the same, it is a simple resend, and if the sequence number is the same but the content differs, it is an incident. If you keep the file hash with hashlib, you can tell these two apart.

수신함 ──▶ 이름·인코딩 ──▶ 레코드 폭·인용 ──▶ 트레일러 ──▶ 필드 ──▶ 중복
                 │              │               │           │        │
                 └──────────── 하나라도 걸리면 파일 전체를 거절 ──────┘

What it looks like in the field

The first form you meet is encoding that breaks silently. The specification is CP949, but the counterparty institution changes its system and starts sending in UTF-8. With luck, decoding ends with an exception and is caught right away. With bad luck, some bytes are interpreted as different characters, and only the names go in broken as they are. So you look not only at the encoding but also at the byte width. When one Korean character changes from 2 bytes to 3, the line length changes, so the width check catches an encoding incident one more time.

The second is the temptation of partial loading. Someone is sure to say, "998 items are fine, so do we have to send everything back because of 2 items?" You must send it back. If you load only the 998, the counterparty's trailer of 1000 and our ledger of 998 remain mismatched, and in the correction file the counterparty sends the next day, those 2 items come in again, and this time they are duplicates. If you keep it all-or-nothing, the incident ends in a day.

The third is rejection without a reason. If you send it back with only "format error," the counterparty's contact does not know what to fix. If you return together a machine-readable rejection file stating which field of which line was caught and why, the counterparty can feed that file straight into its own batch.

What really matters in practice

What you will do in the next lab

The 80-byte record layout and the error code names and exit codes used in the next lab are assumptions of this lab. Specifications differ by institution, and what is the source of truth in the field is the agreement document exchanged with the counterparty institution. What does not change are the order of the layers checked and the all-or-nothing principle.

You build yourself the day's inbox sent by four counterparty institutions and stack up step by step the gatekeeper that will stand in front of it. Starting with trailer matching, you attach the byte width of fixed width, encodings that differ from the specification, RFC 4180 quoting, field validation and rejection files, and resend judgment. The grader makes its own inbound files with a different institution code and count each time, runs your gatekeeper, and compares the judgments. At the end, you process your own inbox and leave a receipt report.