TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

Export the claims data without exposing anyone

Continue in TT Lab

Goal

You export synthetic indemnity-insurance claims data for analysis. You strip the direct identifiers, generalize dates, collapse small regions and rare values, raise k to 5 or more, and then try re-identification yourself and leave an export record.

Why it matters

Even if you erase the name and member number, if date of birth, sex, and postal code remain attached, most people can find their own row. Until you count how many people a quasi-identifier combination points to (k), nothing has been proven. The only ways to raise k are to blur values or remove rows, and if you reverse the order, the rows you throw away snowball. So you generalize first and suppress last. As a model for the checklist we use the 18 items of the US HIPAA's 45 CFR 164.514 Safe Harbor, but this regulation does not apply as it is to Korean insurers. The domestic standard is the Personal Information Protection Act, and health information is handled more strictly as sensitive information.

Steps

  1. Create and run /root/deid/gen_claims.py to create /root/deid/claims.csv (700 or more claims, 300 or more members) and /root/deid/zip_pop.csv. Include 5 or more claims from subscribers over 90, 3 or more diagnosis codes with frequency under 5, and 2 or more regions with a population of 20000 or fewer.
  2. Strip the direct identifiers with /root/deid/strip.py to create /root/deid/stripped.csv. Attach rec_id from R-000001 in the original order.
  3. Count the k of the quasi-identifier combination with /root/deid/kcount.py and write it to /root/deid/k_before.json (qi, rows, groups, k_min, unique_rows, lt5_rows).
  4. Generalize dates with /root/deid/generalize.py to create /root/deid/generalized.csv. Keep only the year of the date of birth, and if the age measured with the reference year 2026 exceeds 89, erase even the year and set age_over_89 to Y. For the claim date keep only the year, and drop the admission date.
  5. Create /root/deid/suppressed.csv with /root/deid/suppress.py. Keep only the first three digits of the postal code, but change regions with a population of 20000 or fewer to 000, change diagnosis codes with frequency under 5 to RARE, and drop the city name.
  6. Suppress the rows whose combination is under 5 cases with /root/deid/kanon.py to create /root/deid/export.csv and /root/deid/k_after.json. At least 70% of the rows must remain.
  7. Try re-identification with /root/deid/reidentify.py to create /root/deid/attack.json. The target is the row with the smallest rec_id in the exported copy.
  8. Leave /root/deid/export_record.json and /root/deid/checklist.md. Measure the hashes and counts from the actual files, and write the basis and scope of application along with them in the checklist.

Notes

Generate the synthetic claims snapshot and the regional population table

Create and run /root/deid/gen_claims.py to create /root/deid/claims.csv and /root/deid/zip_pop.csv.

Make one person claim several times, and spread the dates of birth down to the day. That way rows whose combination is only themselves arise. You also need subscribers over 90, sparsely populated regions, and diagnosis codes that appear only rarely, for there to be something to deal with in the later steps.

Strip the direct identifiers

Create /root/deid/stripped.csv with /root/deid/strip.py. Drop the name, member number, claim number, and hospital name, and attach rec_id.

The file to read is /root/deid/claims.csv made in step 1. Do not throw away a single row. rec_id must be attached in the original order so that you can point to the same row later. The reason to drop the hospital name is that a small hospital narrows down where the person lives.

Count how many people one combination points to

Count the k of the birth_date, sex, zip5 combination with /root/deid/kcount.py and write it to /root/deid/k_before.json.

The file to read is /root/deid/stripped.csv. Count the rows per combination and the minimum is k. Also count how many rows are alone (k of 1) and how many rows are in groups of fewer than five. These numbers are the starting point of the next steps.

Reduce dates to years and group the elderly

Create /root/deid/generalized.csv with /root/deid/generalize.py. If the age is over 90, erase even the birth year.

The file to read is /root/deid/stripped.csv. The Safe Harbor rule is not to leave any element of a date other than the year. Age is taken as 2026 minus the birth year. Very old ages narrow down a person from the year alone, so group them into one cell.

Collapse small regions and rare diagnosis codes

Create /root/deid/suppressed.csv with /root/deid/suppress.py. Regions with a population of 20000 or fewer become 000, and diagnosis codes with frequency under 5 become RARE.

The files to read are /root/deid/generalized.csv and /root/deid/zip_pop.csv. You keep the first three digits of the postal code only when the region has enough people. Decide by reading the population table. Count diagnosis code frequencies within the data being exported this time.

Raise k to 5 and measure what was lost

Suppress the rows whose combination is under 5 cases with /root/deid/kanon.py to create /root/deid/export.csv and /root/deid/k_after.json.

The file to read is /root/deid/suppressed.csv. Suppression is the last resort. If the earlier generalization is sufficient, few rows are thrown away. If the share that remains is under 70%, do not increase suppression; look at the earlier steps again.

Try re-identification yourself

Create /root/deid/attack.json with /root/deid/reidentify.py. The target is the row with the smallest rec_id in the exported copy.

The files to read are /root/deid/export.csv and /root/deid/stripped.csv. What the attacker knows is only the generalized values (birth year, sex, region, claim year). If you compare it with how many rows the same person narrows to in the version with only the direct identifiers erased, you can see in numbers what generalization actually blocked.

Leave an export record and a checklist

Leave /root/deid/export_record.json and /root/deid/checklist.md. Measure the hashes and counts from the actual files.

Measure the hashes and counts from /root/deid/claims.csv and /root/deid/export.csv. The record contains the hashes of the original and the exported copy, the rules applied, the purpose and approver, and the retention period. In the checklist, along with the statutory article that is the basis, also write whether that standard applies to us as it is. If you copy original dates into the list, the list becomes yet another leaked copy.