TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

We removed the names, and the person was still identifiable

Continue in TT Lab

In one line

De-identification is not erasing names but making it so that even the combination of the remaining values does not point to one person, and whether that has been achieved is confirmed not by feeling but by counting k.

Why this was needed

A request comes in asking for data, saying they want to analyze why claim payments are delayed. The requester may be an in-house analytics team or an outside consultancy. It seems enough to erase the name and member number and hand it over, but very often a person is revealed by what remains. If date of birth, sex, and the postal code of where they live are attached, most people can recognize their own data, and someone who has an alumni roster or an in-house address book can recognize other people's data too. This is the quasi-identifier problem. A direct identifier points to a person alone, and quasi-identifiers point to a person by gathering several together.

Health information has one more layer on top of this. The diagnosis code is a value that is essential for analysis and at the same time the most sensitive. In Korea, the Personal Information Protection Act separately defines information about health as sensitive information and has it handled more strictly (Article 23), so it is hard to get export approval with the explanation "we only erased the names."

How it works

The US HIPAA de-identification standard is often cited as a model for checklist items. 45 CFR 164.514 gives two paths. One is expert determination and the other is the Safe Harbor. The Safe Harbor lists 18 identifiers to remove, and three rules within it are worth noting.

This regulation applies to US healthcare organizations and their partners, and it does not apply as it is to Korean insurers. But it is still useful as a checklist for confirming "what did we miss."

The yardstick that measures the risk remaining even if you follow the whole list is k-anonymity. You count how many rows share the same quasi-identifier combination, and call the minimum of that k. If k is 1, a person who knows that combination can pick out exactly one row. There are only two ways to raise k. Blur values (generalization) or remove rows (suppression). Generalization loses less information, and suppression is certain but reduces the data available for analysis. So the order matters. Generalize first, and suppress only the small combinations that still remain. If you do the opposite, the rows you throw away snowball.

원본:      1971-04-02 · F · 80712  →  이 조합을 가진 줄이 1개 (k=1)
일반화:    1971 · F · 807          →  이 조합을 가진 줄이 9개 (k=9)
작은 지역: 1943 · M · 000          →  인구가 적어 앞자리까지 지움
억제:      조합이 5개 미만인 줄은 반출하지 않음

What it looks like in the field

The most common incident is judging it safe by looking at only one file. If you place last month's extract and this month's extract side by side, the combinations of people who appear in both shrink and narrow each other down. That is why you leave in the export record what was released, when, to whom, and by which rule. Without a record, you cannot even recompute the risk when the next request comes.

The second is being reassured by k alone. Even with k of 5, if the diagnosis codes of those five rows are all the same, you may not know who it is, but what illness that person claimed for is revealed. This problem, where the value itself leaks, is not blocked by k. For sensitive columns, it is better to separately look at how many kinds of values there are within a combination.

The third is sending along the key to reverse it. The re-identification provision of the same article says that even when you attach a code that allows reversal, that code must not be derived from the original value and the method must not be disclosed. Using the member number as it is, or hashing the member number and attaching it, violates this condition.

What really matters in practice

What you will do in the next lab

You build 800 synthetic claims, strip the direct identifiers and count k, and then apply date generalization, grouping of the elderly, and handling of small regions and rare diagnosis codes in order. Then you suppress the combinations with k under 5 to make the exported copy, try re-identification yourself to measure how many people that person hid among, and leave an export record and a checklist. At each step the grader recomputes the result from the previous step's file and checks line by line.