When You End Up Seeing a Diagnosis Code Mid-Investigation
In one line
Insurance data contains health information, and it is sensitive information that is subject to stronger regulation than ordinary personal information, so looking at it in an investigation and moving it into a document are completely different acts.
Why this was needed
You query the claims table. You are measuring elapsed time to investigate a payment delay, and columns like this show up on the screen together.
claim_no policy_no accident_date diagnosis_code claimed_amount
C0034 P0034 2026-04-05 C34.9 5000000
C34.9 is a KCD code. It is a malignant neoplasm of the lung, that is, lung cancer. What we meant to look at was elapsed time, but right next to that column is what illness this person has.
One careless act here becomes an incident. It is capturing the screen and pasting it into the company messenger. Copying the whole table over to explain the cause is the same. At that moment one lung cancer patient's diagnosis information leaves the client's system.
How it works
Personal information has grades.
일반 개인정보 이름, 연락처, 주소 같은 것
민감정보 건강, 유전, 사상·신념, 노동조합 가입, 성생활, 범죄경력, 생체인식
고유식별정보 주민등록번호, 여권번호, 운전면허번호, 외국인등록번호
Health information is sensitive information. The Personal Information Protection Act treats sensitive information separately, and it cannot be processed under consent obtained bundled together with ordinary personal information. You must have obtained separate consent or have a basis in law. An insurer has that basis, but that basis belongs to the insurer, not to us.
So the line to keep in the field is clear.
조사에 필요해서 화면으로 보는 것 업무 범위 안. 접근 로그가 남는다
결과를 문서·메신저·이메일로 옮기는 것 반출. 별개의 판단이 필요하다
The second is the problem. The first is usually under control and leaves logs, but the second happens without any resistance. One line of a query result pasted into Slack, one screenshot put into a report, one CSV downloaded locally for debugging.
That is why clients block log export. The first time you run into it, it feels like "they will not let me work," but insurer application logs often have diagnosis codes and insured-person identifiers mixed in during claims processing. A single log file is a sensitive-information file. The export ban is not an obstruction; it is that side being normal.
There are three rules to keep when writing a report.
사람이 아니라 건으로 지목한다 청구번호·계약번호로 말한다. 이름·식별자는 쓰지 않는다
진단은 쓰지 않는다 사유 코드까지가 한계다. 병명·KCD 코드는 옮기지 않는다
무엇을 안 담았는지 밝힌다 읽는 사람이 조사 범위를 알 수 있어야 한다
The third will seem unfamiliar, but this is what really builds trust. If there is one line saying "diagnosis information was looked up but not included in the report," the client's security officer can tell what we knew and what we left out. Without that sentence, that person has to guess what we carried out.
Also keep pseudonymization and anonymization distinct. Pseudonymization is a state in which, if combined with other information, an individual can be recognized again, so it is still personal information. Anonymization is a state that cannot be restored. "It is fine because we replaced it with a number" is usually pseudonymization, and even pseudonymized data needs a procedure to be exported.
What it looks like in the field
First, a debugging dump is the most common path to an incident. Because it cannot be reproduced, you take part of the production data and run it locally. There is no malice and it is for getting work done, but that file stays on the laptop and goes up to the cloud as a backup. If reproduction is needed, it should be done inside the client's environment, or you should negotiate a way of receiving only the structure with the values removed.
Second, pasting into AI tools has become a new path in the last few years. You paste an error message, and there really are cases where the stack trace contains claims data. You need the habit of looking at what is in it before pasting.
Third, requesting the minimum scope when receiving access permissions protects you later. If you have no read permission on the table with diagnosis codes, there is neither a chance of seeing it by mistake nor of moving it. If you take broad permissions for convenience, then when an incident happens those permissions become the scope of suspicion.
When designing screens and queries that handle sensitive information
In the insurance domain, developers get to see medical history, diagnosis names, and beneficiary relationships. This is one level above ordinary personal information, so you separate what can be seen and what cannot be seen in the design, not in the code.
Screens are masked by default. Let the people who need it unmask it when they need to, and record the act of unmasking itself. If "who looked up what and when" is not left, you cannot narrow down the scope even when an incident happens.
Protect the lookup log at the same grade as the data. Because the lookup records contain patient identifiers and diagnosis codes, if that log leaks, it is no different from the original leaking.
When exporting to an analytics system, write the purpose first. There are almost no cases where a resident registration number is needed to produce statistics. Age band, sex, and region are enough, and if linking is needed, use an irreversible pseudonymous identifier. If you use the same identifier in several datasets, they can be combined and re-identified, so make a different value for each dataset.
Do not copy production data into the development environment. "Just for a moment" stays for months, and the development environment has weaker controls than production. What you need is usually data of the same shape, so build synthetic data. If you really need the real thing, only the minimum number of records, with a deadline, under separate approval.
Block the paths through which query results leave the screen. Excel downloads, raw queries in admin tools, and messenger pasting are the real leak paths. Put a record-count cap on bulk lookups and require approval when it is exceeded.
Set a retention period for each kind of data. Policy-related documents have a statutory retention period, and lookup logs may be shorter than that. If you do not set it, everything stays forever, and that is the biggest risk.
What to check in the next quiz
You check what sensitive information is, why lookup and export are different judgments, how to point at a person in a report, and why pseudonymization does not mean safety.