TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

Same Receipt, Slightly Different Characters, and the System Calls It a New Claim

Continue in TT Lab

In one line

Finding duplicate claims is not finding identical strings but finding what becomes identical after normalization and keeping only the genuine duplicates among them. The hard and expensive part is not finding but reducing false positives.

Why this was needed

There are several windows for receiving insurance claims. Photos are uploaded through the app, entered again on the web, submitted on paper at a counter, and sometimes arrive by fax. When the connection drops, the app sends the same request again. So the same single medical visit remaining as several lines in the intake system is not an exception but an everyday occurrence.

Two different things are mixed here. One is an exact duplicate. Two lines with even the same claim number come in. This is a resubmission, and grouping by claim number ends it. The other is a de facto duplicate. The claim number is different but it is the same receipt. After submitting through the app, it seemed processing was slow, so it was submitted once more at the counter. The characters here are not exactly the same. The app receives the name without a space and the counter clerk writes it with a space between the surname and the given name ("KimSeojun" versus "Kim Seojun"). The app copies the receipt number as is, including the hyphen, and the counter system stores it in full-width characters.

If you miss the second, the same money goes out twice. But an accident in the opposite direction is quieter. If you block as a duplicate the claim of someone who really did receive treatment twice at the same hospital on the same day, the customer cannot get paid without knowing why and calls the call center. Half the work of building a detector is reducing these false positives.

How it works

The order is always the same. Normalization → blocking → similarity → threshold → tolerance → false-positive rules.

Normalization. First fold away the wobble in notation. Among the Unicode normalization forms, NFKC is a form that does compatibility decomposition followed by canonical composition (UAX #15), and it folds compatibility characters that differ only in shape, such as full-width alphanumerics, into the ordinary form. Measured in the lab image, unicodedata.normalize("NFKC", "RC-2026") returns "RC-2026". Add removal of spaces and removal of hyphens to this, and the notations that differ by counter gather into one value. The unicodedata module and the guide to handling Unicode explain these forms.

Blocking. Comparing 250,000 claims against each other gives 30 billion pairs. So you group only those that share the same value and compare only within that. In this lab the normalized hospital name and the service date are used as the block key. Blocking is not free. If the block key is wrong, that pair never becomes a candidate to begin with. It often happens that wobble in how the hospital name is written splits the blocks, so use the normalized value for the block key too.

Similarity. Use SequenceMatcher from difflib. ratio() is 2.0*M / T, where T is the total number of elements in the two sequences and M is the number of matching elements. It is 1.0 if identical and 0.0 if nothing overlaps. There is one trap. If the second sequence is 200 or more elements, the automatic junk heuristic turns on and treats elements taking up more than 1% as junk. Names and receipt numbers are short so it does not matter, but when comparing long text you should think about autojunk=False. And this heuristic is asymmetric, so the result can differ if you swap A and B.

Threshold and tolerance. Take a weighted sum of the per-field similarities to produce a score and cut it with two thresholds. Above the higher one is a duplicate, the middle is the band a person looks at, and below is distinct. Then you multiply in tolerances for amount and date. Even if the score is 1.0, if the amount is 4,000 won apart it may not be the same receipt. Do not block automatically; send it to a person.

False-positive rules. Receipts received one after another on the same day at the same hospital have consecutive numbers. By string similarity, only one character of twelve differs, so you get 0.93, and since the name and hospital are also the same, the combined score exceeds the threshold. This is not a duplicate but an adjacent document. If only the last chunk of digits of the number differs by 1 to 5, take it out of the duplicate candidates.

청구 25만 ──▶ 정규화 ──▶ 블록 ──▶ 블록 안의 쌍 ──▶ 점수 ──▶ 허용 오차 ──▶ 오탐 규칙
                                                         │            │          │
                                                    duplicate      review     distinct

What it looks like in the field

The most common mistake is grouping by claim number only. Then resubmissions are caught and not one de facto duplicate is. And that fact does not show on screen. A report saying "0 duplicates" comes out every day, and it gets caught in audit after payment.

The second is producing only a score and leaving no basis. Even when a reviewer receives "0.94," they do not know what to check. If you leave a list of which fields matched to raise the score, the reviewer looks only at those fields with their own eyes and finishes in 5 seconds. The judgment result is not the automation's output but a person's input.

The third is setting the threshold without data. If you set it because 0.9 looks good, you start operating without knowing how many cases a day go to people. If investigation staff handle 30 a day and 200 come over, that list is ignored within a few days. The threshold is not a rule but a contract with staffing.

What really matters in practice

What you will do in the next lab

The weights used in the next lab (0.45, 0.35, 0.20), the thresholds (0.92, 0.80), and the tolerances (100 won, 0 days) are this lab's assumptions. In practice you set them from past judgment history and the investigation staff's daily throughput, and attach a version number to the ruleset and record every change.

You build a day's claim table of the intake system yourself, and stack up step by step from exact duplicates through normalization, blocking, similarity, tolerance, and removal of consecutive-number false positives. The grader builds its own claim table each time with different patients, hospitals, and receipt numbers, runs your detector, and compares the judgments. At the end you render judgments on your own table and produce a review list and a report that the review staff can look at as they are.