TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

The Same Receipt Was Used Twice Under Two Names

Continue in TT Lab

In one line

Checking claim supporting documents is not counting whether all the files arrived but judging whether the file really is that document by its content, not its name. The format is judged by magic bytes, sameness by content hash, and mismatches between records and the real files by a two-way comparison.

Why this was needed

In an insurance claim, documents are the basis for the payment decision. But when there are several intake windows and several channels, the state of the files coming into the same system varies. A person in charge uploads a scanner-made PNG with .pdf attached, files are sent with the insured person's name in the file name, and the same receipt is submitted to two claims with only the name changed. The reviewer looks only at what opens on screen, so none of the three is easily caught by eye.

If you leave the check to human eyes, three things collapse at the same time. Claims are processed as complete when documents are missing (the payment basis is empty), the same document is used twice as a payment basis (double payment), and personal information carried in file names spreads into lists, backups, and logs. If any of the three is found later, undoing it is very expensive. Conversely, an automatic check at intake time is cheap — all it needs is one rules table, the first few bytes of the file, and a hash.

How it works

Keep the checklist as data, not code. The documents required differ by claim type, and that list moves whenever products and regulations change. If you put it in a table with two columns, type and document kind, the checker only has to read the table, and even when regulations change, you do not fix the code. In a store like SQLite where column types are loose, it is also good to know what the storage class of a value is (SQLite data types).

Separate formats by the leading bytes. The extension is a label a person attached and guarantees nothing. The media type registration entry of RFC 8118 pins down that a PDF starts with %PDF-, and section 3.1 of RFC 2083 defines that the first eight bytes of a PNG are, in decimal, 137 80 78 71 13 10 26 10. A JPEG starts with FF D8 FF, and a ZIP has 03 04 after PK. What the Linux file command does is, in the end, to look at this table.

There is one thing you must keep here. These files have NUL bytes mixed in, so you must not apply grep to them. GNU grep treats input with NUL mixed in as a binary file and either prints only "Binary file matches" or quietly finds nothing. Formats that start with NUL fail for certain on real files. Look with od -c, or read it in Python with open(path, "rb").read(16) and compare.

Sameness is judged by content hash. You can never say from the file name whether it is the same document. If you compute the sha256 of the file contents with hashlib and group equal values, documents that were resubmitted under a changed name show up as one group. When putting a hash into a list people read, writing it in hexadecimal is usual, and the document that sorts out the places of base16, base32, and base64 when you must put bytes into text as they are is RFC 4648. For code that handles paths, using pathlib instead of string operations reduces mistakes in splitting extensions and names.

Records and real files are matched in both directions. A record that has no file and a file that has no record are different incidents. The former is an empty payment basis, and the latter is personal information of unknown ownership left behind. A checker that looks in only one direction always misses half.

doc_rule(규칙표) ─┐
                  ├─▶ 빠진 서류      ─┐
submission(기록) ─┤                    ├─▶ 청구별 보완 요청 목록
                  ├─▶ 고아 / 없는 파일 ┤
inbox(실물) ──────┴─▶ 형식 · 해시 · 이름 ┘

What it looks like in the field

Getting the start date of the deadline calculation wrong is very common. If you count from the filing date, the later a case was filed, the more its deadline extends. The start date must be the accident date, and that rule is written in one line of a document and kept together with the checker. In this lab, the deadline is 30 days from the accident date, but that is not a real statutory deadline and is this lab's assumption. It differs by product and company, so in the field you must always check that company's regulations.

Personal information in file names also often stays as it is. The content gets de-identified while the name is left untouched in many cases, and list screens, backup indexes, error logs, and even report file names carry that name as it is. So that the check report itself does not spread the same value once more, write only counts and rules in the report and do not put in the values.

What really matters in practice

What you will do in the next lab

You build 60 claims and 189 inbox files of Daon Insurance (fictional) yourself, and find missing documents with the rules table. You build a format judge that does not trust extensions and produce a judgment table for the whole inbox, group resubmitted documents by content hash, and count the personal information carried in file names. You compute the deadline from the accident date, match the records and the real files in both directions, and finish with a per-claim follow-up request list and a report.