Checking That a Claim Document Is Really That Document
Goal
You check the required documents by claim type against a rules table, judge file formats by magic bytes, find resubmitted documents by content hash, and sweep through personal information in file names, submission deadlines, and orphan files to produce a per-claim follow-up request list.
Why it matters
Documents are the basis for the payment decision, but the file name does not tell you whether the file really is that document. A pdf extension gets attached to a scanner-made image, and the same receipt is used in two claims with only the name changed. An automatic check at intake time is cheap, and finding it after payment is very expensive. All you need is one rules table, the first few bytes of the file, and a hash. The files in this inbox have NUL bytes mixed in. If you apply grep, it finds nothing or only tells you it is a binary file. Look with od or read as bytes in Python.
Steps
- Create /root/docs/claims.db and /root/docs/inbox by saving the material script as /root/docs/gen_docs.py and running it.
- In /root/docs/missing.txt write the missing documents found by matching the rules table against the intake records.
- In /root/docs/filetypes.tsv produce a judgment table for the whole inbox, using /root/docs/sniff.py to judge the formats.
- In /root/docs/reuse.txt write the resubmitted documents, grouped by content hash.
- In /root/docs/namescan.txt write the count of personal information carried in file names.
- In /root/docs/overdue.txt write the results of computing the submission deadline from the accident date.
- In /root/docs/orphans.txt write the results of matching the records and the real files in both directions.
- In /root/docs/followup.json produce the per-claim follow-up request list, and in /root/docs/docs_report.md write the report.
Notes
- Use only five format names: pdf, png, jpg, zip, unknown. sniff.py takes one file path as an argument and prints one line with the format name.
- filetypes.tsv has one line per file, with three tab-separated cells: the file name, the extension (lowercase, taken from the name), and the judged format.
- The submission deadline is 30 days from the accident date. This is this lab's assumption and not a statutory deadline. Whether the deadline has passed is counted with 2026-09-17 as today.
- The rule for judging personal information in file names: if it contains 2 to 4 Hangul characters, it is a name, and if it has the shape of six digits, a hyphen, and seven digits, it is treated as a number.
- Attach to sqlite3 with
sqlite3 -readonly /root/docs/claims.db "SELECT ...". Look into the inbox withod -c /root/docs/inbox/K0001_claim_form.pdf | head -2. - Common mistakes: judging the format by extension, taking the filing date as the start date of the deadline, matching records and real files in only one direction, and copying personal information values into the report as they are.
Generate the claims table and the inbox
Create /root/docs/claims.db and /root/docs/inbox by saving the material script as /root/docs/gen_docs.py and running it.
First create /root/docs and run it with python3 inside. There are three tables, claim, doc_rule, and submission, and 189 files are created in the inbox. Do not modify the files that get created.
Run the required-documents checklist by type
In /root/docs/missing.txt write incomplete_claims, missing_docs, missing_inj, missing_ill, and missing_med, found by matching the rules table against the intake records.
doc_rule holds the document kinds needed for each claim type. For each claim, subtract the kinds in submission from that set and what is missing remains. The count per type is the value from counting the claims that have missing documents split by type.
Separate formats by the leading bytes, not the extension
Create /root/docs/sniff.py and in /root/docs/filetypes.tsv produce the judgment table for the whole inbox.
A PDF starts with %PDF-, a PNG with the decimal 137 80 78 71 13 10 26 10, a JPEG with FF D8 FF, and a ZIP with PK 03 04. If it matches none, it is unknown. Always open the file as bytes. The table splits the file name, extension, and judgment into three cells with tabs.
Find where the same document was reused
Group identical files by content hash and in /root/docs/reuse.txt write reused_groups, reused_files, and affected_claims.
File names can be changed freely, so group by sha256. Count only groups of two or more. The number of affected claims is the count of distinct claim_id values traced back through the filename in submission.
Count the personal information carried in file names
In /root/docs/namescan.txt write name_in_filename, rrn_in_filename, and flagged_files.
However well you mask the content, the file name stays as it is on list screens, in backup indexes, and in error logs. If it contains 2 to 4 Hangul characters, it is a name, and if it has the shape of six digits, a hyphen, and seven digits, it is a number. flagged_files is the number of files flagged by either one.
Count the deadline from the accident date
In /root/docs/overdue.txt write late_docs, late_claims, and overdue_missing.
The start date is the accident date, not the filing date. Count the documents submitted more than 30 days after the accident date, and count the claims that have at least one such document. overdue_missing is the number of claims that have missing documents and whose deadline passed before 2026-09-17.
Match records and real files in both directions
In /root/docs/orphans.txt write orphan_files, missing_files, and claims_with_missing_file.
Files only in the inbox and files only in the records are different incidents. The former is personal information of unknown ownership left behind, and the latter is an empty payment basis. Take the set difference of the two sets in both directions.
Produce the per-claim follow-up request list
In /root/docs/followup.json put each claim's missing, no_file, reused, and type_mismatch, and in /root/docs/docs_report.md write a five-section report.
Write the four reasons separately within a claim. Claims with no reason at all do not go in the list. Put the file lists in sorted. The report section titles are what was checked, missing documents, format and reuse, personal information in file names, and recommendations. Do not copy personal information values into the report.