TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

Checking That a Claim Document Is Really That Document

Continue in TT Lab

Goal

You check the required documents by claim type against a rules table, judge file formats by magic bytes, find resubmitted documents by content hash, and sweep through personal information in file names, submission deadlines, and orphan files to produce a per-claim follow-up request list.

Why it matters

Documents are the basis for the payment decision, but the file name does not tell you whether the file really is that document. A pdf extension gets attached to a scanner-made image, and the same receipt is used in two claims with only the name changed. An automatic check at intake time is cheap, and finding it after payment is very expensive. All you need is one rules table, the first few bytes of the file, and a hash. The files in this inbox have NUL bytes mixed in. If you apply grep, it finds nothing or only tells you it is a binary file. Look with od or read as bytes in Python.

Steps

  1. Create /root/docs/claims.db and /root/docs/inbox by saving the material script as /root/docs/gen_docs.py and running it.
  2. In /root/docs/missing.txt write the missing documents found by matching the rules table against the intake records.
  3. In /root/docs/filetypes.tsv produce a judgment table for the whole inbox, using /root/docs/sniff.py to judge the formats.
  4. In /root/docs/reuse.txt write the resubmitted documents, grouped by content hash.
  5. In /root/docs/namescan.txt write the count of personal information carried in file names.
  6. In /root/docs/overdue.txt write the results of computing the submission deadline from the accident date.
  7. In /root/docs/orphans.txt write the results of matching the records and the real files in both directions.
  8. In /root/docs/followup.json produce the per-claim follow-up request list, and in /root/docs/docs_report.md write the report.

Notes

Generate the claims table and the inbox

Create /root/docs/claims.db and /root/docs/inbox by saving the material script as /root/docs/gen_docs.py and running it.

First create /root/docs and run it with python3 inside. There are three tables, claim, doc_rule, and submission, and 189 files are created in the inbox. Do not modify the files that get created.

Run the required-documents checklist by type

In /root/docs/missing.txt write incomplete_claims, missing_docs, missing_inj, missing_ill, and missing_med, found by matching the rules table against the intake records.

doc_rule holds the document kinds needed for each claim type. For each claim, subtract the kinds in submission from that set and what is missing remains. The count per type is the value from counting the claims that have missing documents split by type.

Separate formats by the leading bytes, not the extension

Create /root/docs/sniff.py and in /root/docs/filetypes.tsv produce the judgment table for the whole inbox.

A PDF starts with %PDF-, a PNG with the decimal 137 80 78 71 13 10 26 10, a JPEG with FF D8 FF, and a ZIP with PK 03 04. If it matches none, it is unknown. Always open the file as bytes. The table splits the file name, extension, and judgment into three cells with tabs.

Find where the same document was reused

Group identical files by content hash and in /root/docs/reuse.txt write reused_groups, reused_files, and affected_claims.

File names can be changed freely, so group by sha256. Count only groups of two or more. The number of affected claims is the count of distinct claim_id values traced back through the filename in submission.

Count the personal information carried in file names

In /root/docs/namescan.txt write name_in_filename, rrn_in_filename, and flagged_files.

However well you mask the content, the file name stays as it is on list screens, in backup indexes, and in error logs. If it contains 2 to 4 Hangul characters, it is a name, and if it has the shape of six digits, a hyphen, and seven digits, it is a number. flagged_files is the number of files flagged by either one.

Count the deadline from the accident date

In /root/docs/overdue.txt write late_docs, late_claims, and overdue_missing.

The start date is the accident date, not the filing date. Count the documents submitted more than 30 days after the accident date, and count the claims that have at least one such document. overdue_missing is the number of claims that have missing documents and whose deadline passed before 2026-09-17.

Match records and real files in both directions

In /root/docs/orphans.txt write orphan_files, missing_files, and claims_with_missing_file.

Files only in the inbox and files only in the records are different incidents. The former is personal information of unknown ownership left behind, and the latter is an empty payment basis. Take the set difference of the two sets in both directions.

Produce the per-claim follow-up request list

In /root/docs/followup.json put each claim's missing, no_file, reused, and type_mismatch, and in /root/docs/docs_report.md write a five-section report.

Write the four reasons separately within a claim. Claims with no reason at all do not go in the list. Put the file lists in sorted. The report section titles are what was checked, missing documents, format and reuse, personal information in file names, and recommendations. Do not copy personal information values into the report.