TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

Purging Personal Data Past Its Retention Period

Continue in TT Lab

Goal

You keep the retention period table and retention orders as data and delete personal information past its period, split into a rehearsal and an actual purge. You leave purge evidence, and check even whether the deleted people's information truly cannot be read and what remains in the backup.

Why it matters

With personal information, not deleting is an incident and deleting wrongly is an incident. Once the retention period passes you must destroy without delay, but a case under a retention order because of a dispute or a regulator's request must not be deleted even after the period passes. If you leave the place where the two rules collide to human judgment, it gets handled differently each time, so you turn the judgment into data that can be recomputed. In particular, the whole result changes with a single basis date. Depending on whether you count a claim's 3 years from the accident date or the closing date, dozens of records that must still be retained may be deleted, and that incident shows itself only after the erasure. The grader does not trust the list you submit. It recomputes directly from the contract, claim, and rules tables to compare, and after the purge it sweeps whether the contact details recovered from the backup are still in files outside the backup.

Steps

  1. Create and run /root/retention/gen_retention.py to create /root/retention/ins.db, the documents, and the backup. It has 120 contracts, 240 claims, 360 lines of personal information, 4 rule lines, and 12 retention orders.
  2. Write the cases that would be purged more if the basis date were switched to the old version rules to /root/retention/basis_gap.csv, and five numbers to /root/retention/basis.txt.
  3. Write the cases whose retention period has passed as of the reference date to /root/retention/due.csv. Do not look at retention orders yet.
  4. Write the cases removed by retention orders to /root/retention/hold_excluded.csv and the final targets to /root/retention/due_final.csv.
  5. Output the rehearsal plan to /root/retention/purge_plan.json. Nothing is deleted in this step.
  6. Create and run /root/retention/purge.py to delete as planned and leave the history in purge_log.
  7. Output the purge evidence to /root/retention/purge_evidence.json.
  8. Write the copies remaining in the backup to /root/retention/backup_residual.csv and a report to /root/retention/retention_report.md.

Notes

Generate the contract and claim snapshot

Create and run /root/retention/gen_retention.py to create /root/retention/ins.db, /root/retention/docs/, and /root/retention/backup/2026-08-31/docs/. It has 120 contracts, 240 claims, 360 lines of personal information, 4 rule lines, and 12 retention orders.

Six tables. About one in nine claims should have an empty closing date (cases not yet ended), and the rules table holds both the current rules (policy, claim) and the old version rules (policy_alt, claim_alt). Make one document per claim and copy it as is into the backup directory.

What changes if you change the basis date

Write the cases that are purge targets under the old version rules (policy_alt, claim_alt) but not yet under the current rules to /root/retention/basis_gap.csv as subject_type,subject_id,primary_expiry,alt_expiry, and write due_policy=, due_claim=, due_policy_alt=, due_claim_alt=, extra_under_alt= to /root/retention/basis.txt.

The expiry date is the basis date plus the number of years. If you use timedelta(days=365*n) it slips in leap years, so use date.replace(year=...). A case with an empty basis date has no expiry date, so leave the primary_expiry cell blank.

Pull out the cases whose period has passed

Write the cases whose retention period has passed as of the reference date (sys_param.asof) to /root/retention/due.csv as subject_type,subject_id,rule_kind,basis_date,expiry_date, in ascending order of subject_type and subject_id. Do not look at retention orders yet.

The current rules are 5 years from contract_end for contracts and 3 years from claim_closed for claims. A claim with an empty closing date is not a target, since its prescription has not started. If you fill an empty date with a very old date, claims in progress are erased entirely.

Take out the cases under a retention order

Write the cases among the purge targets caught by legal_hold to /root/retention/hold_excluded.csv as subject_type,subject_id,hold_id,reason, and the remaining final targets to /root/retention/due_final.csv with the same header as due.csv.

Apply retention orders after the period judgment. If you change the order, cases that have an order but whose period has not yet passed also go into the exclusion list, and the numbers get blurred. For reason, use the wording written in the legal_hold table as it is.

Produce the plan before deleting

In /root/retention/purge_plan.json, produce a rehearsal plan containing mode (dry-run), executed (false), asof, counts (policy, claim, files), subjects (subject_type, subject_id, expiry_date), and files (paths of the documents to delete). Nothing is deleted in this step.

If you split plan and execution with a single flag in the same code, someday that flag gets set wrong. If you output the plan as a file and make the execution read that file, there is a place for a person to check. Document paths exist only for claims.

Actually delete as planned

Create and run /root/retention/purge.py to delete the subject rows and document files of the targets written in the plan and leave the history in purge_log. Do not touch the backup.

If you delete trusting only the plan file, you miss a retention order placed after the plan was made. Judge once more right before execution and stop if it differs from the plan. The table rows and the files must be handled together, and purged_at is a date and time.

Produce purge evidence and check that it cannot be revived

Write asof, purged (policy, claim), files_removed, held_kept, pii_rows_left_for_purged, doc_files_left_for_purged, and backup_copies_left to /root/retention/purge_evidence.json. The numbers must be values obtained by directly counting the current state.

Evidence is a count, not wording. The two numbers that must not remain (personal information rows, document files) must be 0, and the number of backup copies will not be 0. That difference is the topic of step 8.

Reveal the copies left in the backup and write the report

Write the cases that were deleted but still have a copy in the backup to /root/retention/backup_residual.csv as subject_type,subject_id,backup_path, and write five sections in /root/retention/retention_report.md: ## 무엇을 지웠나, ## 기산일을 어떻게 정했나, ## 보존 명령, ## 남은 사본, ## 재발 방지 (the five Korean section titles mean: what was deleted, how the basis date was decided, retention orders, remaining copies, and preventing recurrence).

Write only the backup paths where a file actually exists. In the report, include, with numbers, the number purged, the statutory article that is the basis for the 3 years of claims, which day was taken as the basis date, the number kept because of retention orders, and the story of the copies left in the backup.