We removed the names, and the person was still identifiable
In one line
De-identification is not erasing names but making it so that even the combination of the remaining values does not point to one person, and whether that has been achieved is confirmed not by feeling but by counting k.
Why this was needed
A request comes in asking for data, saying they want to analyze why claim payments are delayed. The requester may be an in-house analytics team or an outside consultancy. It seems enough to erase the name and member number and hand it over, but very often a person is revealed by what remains. If date of birth, sex, and the postal code of where they live are attached, most people can recognize their own data, and someone who has an alumni roster or an in-house address book can recognize other people's data too. This is the quasi-identifier problem. A direct identifier points to a person alone, and quasi-identifiers point to a person by gathering several together.
Health information has one more layer on top of this. The diagnosis code is a value that is essential for analysis and at the same time the most sensitive. In Korea, the Personal Information Protection Act separately defines information about health as sensitive information and has it handled more strictly (Article 23), so it is hard to get export approval with the explanation "we only erased the names."
How it works
The US HIPAA de-identification standard is often cited as a model for checklist items. 45 CFR 164.514 gives two paths. One is expert determination and the other is the Safe Harbor. The Safe Harbor lists 18 identifiers to remove, and three rules within it are worth noting.
- Remove geographic units smaller than a state. The exception is the first three digits of the postal code, but only when the population of the areas sharing those leading digits exceeds 20,000. If it is 20,000 or fewer, it says to change even the first three digits to
000. - Remove all elements of dates except the year. Date of birth, admission date, discharge date, and date of death all fall here.
- Group all ages over 89 into one cell. Very old ages are rare in themselves, so a person can be narrowed down from the year alone.
- The last, 18th item is "any other unique number, characteristic, or code." It is a provision acknowledging the holes that remain even if you follow the whole list, and the same provision also attaches the condition that you must not release it "knowing that it could be combined with other information to identify an individual."
This regulation applies to US healthcare organizations and their partners, and it does not apply as it is to Korean insurers. But it is still useful as a checklist for confirming "what did we miss."
The yardstick that measures the risk remaining even if you follow the whole list is k-anonymity. You count how many rows share the same quasi-identifier combination, and call the minimum of that k. If k is 1, a person who knows that combination can pick out exactly one row. There are only two ways to raise k. Blur values (generalization) or remove rows (suppression). Generalization loses less information, and suppression is certain but reduces the data available for analysis. So the order matters. Generalize first, and suppress only the small combinations that still remain. If you do the opposite, the rows you throw away snowball.
원본: 1971-04-02 · F · 80712 → 이 조합을 가진 줄이 1개 (k=1)
일반화: 1971 · F · 807 → 이 조합을 가진 줄이 9개 (k=9)
작은 지역: 1943 · M · 000 → 인구가 적어 앞자리까지 지움
억제: 조합이 5개 미만인 줄은 반출하지 않음
What it looks like in the field
The most common incident is judging it safe by looking at only one file. If you place last month's extract and this month's extract side by side, the combinations of people who appear in both shrink and narrow each other down. That is why you leave in the export record what was released, when, to whom, and by which rule. Without a record, you cannot even recompute the risk when the next request comes.
The second is being reassured by k alone. Even with k of 5, if the diagnosis codes of those five rows are all the same, you may not know who it is, but what illness that person claimed for is revealed. This problem, where the value itself leaks, is not blocked by k. For sensitive columns, it is better to separately look at how many kinds of values there are within a combination.
The third is sending along the key to reverse it. The re-identification provision of the same article says that even when you attach a code that allows reversal, that code must not be derived from the original value and the method must not be disclosed. Using the member number as it is, or hashing the member number and attaching it, violates this condition.
What really matters in practice
- Removing direct identifiers is only the start. Until you count k over the remaining combinations, nothing has been proven.
- Generalization first, suppression later. You must also report the share of rows thrown away to be able to say "how much did we lose."
- Try re-identification yourself. A de-identification you have not attacked with your own data is unverified.
- Leave an export record. The hashes of the original and the exported copy, the rules applied, the purpose, and even the retention period.
- When citing a standard, also write its scope of application. If you write another country's regulation as if it were ours, you get caught first in an audit.
What you will do in the next lab
You build 800 synthetic claims, strip the direct identifiers and count k, and then apply date generalization, grouping of the elderly, and handling of small regions and rare diagnosis codes in order. Then you suppress the combinations with k under 5 to make the exported copy, try re-identification yourself to measure how many people that person hid among, and leave an export record and a checklist. At each step the grader recomputes the result from the previous step's file and checks line by line.