Producing the One Page That May Leave
Goal
You finish the investigation on customer data mixed with personal data, and make one sheet that may go outside. You leave each of the four principles (minimum collection, pseudonymization, aggregated export, record keeping) as a file.
Environment
You work under /root/pii. python3, openssl, and jq are available. You create the original yourself — what follows is preparation, not a task.
mkdir -p /root/pii && cd /root/pii
python3 - <<'PY'
import csv, random
random.seed(7)
names = ['김민준','이서연','박지호','최수아','정하윤','강도윤','조서준','윤지우']
codes = ['TIMEOUT','DECLINED','INVALID_CARD','NETWORK','OK']
rows = []
for i in range(1, 201):
n = random.choice(names)
rows.append({
'order_id': 'ORD-%04d' % i,
'created_at': '2026-09-%02dT%02d:%02d:00Z' % (
random.randint(1, 7), random.randint(0, 23), random.randint(0, 59)),
'name': n,
'email': 'user%03d@example.com' % i,
'phone': '010-%04d-%04d' % (random.randint(1000, 9999), random.randint(1000, 9999)),
'address': '서울시 어딘가 %d길 %d' % (random.randint(1, 90), random.randint(1, 300)),
'card_last4': '%04d' % random.randint(1000, 9999),
'amount': random.randint(1000, 90000),
'status': random.choice(['FAILED', 'FAILED', 'PAID']),
'error_code': random.choice(codes),
})
with open('orders.csv', 'w', newline='', encoding='utf-8') as f:
w = csv.DictWriter(f, fieldnames=list(rows[0].keys()))
w.writeheader(); w.writerows(rows)
print(len(rows), '행')
PY
There is one purpose of the investigation — to find out for what reasons payment failures occur. Names and addresses are not needed for that purpose.
What you will build
slim.csv 조사에 필요한 열만 남긴 사본
pseudo.sh 같은 사람을 묶어야 할 때 쓰는 가명값 생성기
masked.csv 육안 확인용 부분 마스킹본
report.csv 밖으로 나가는 유일한 파일 — 사유별 집계
scan.sh 반출 전 검사기
access.log 무엇을 왜 봤는지
handoff.md 무엇을 내보내고 무엇을 안 내보내는지, 그 근거
How it is graded
In steps 3 and 6, the grader runs your scripts directly.
pseudo.sh 같은 값(대소문자·공백만 다름) → 같은 결과여야 한다
키가 다르면 → 다른 결과여야 한다
출력 길이 → 32자 이상이어야 한다
scan.sh 이메일이 든 파일 → 막아야 한다
전화번호가 든 파일 → 막아야 한다
집계만 든 파일 → 통과시켜야 한다
Steps
- Create the original (the preparation block above).
slim.csv— keep only the columns needed for the investigation. The number of rows stays the same.pseudo.sh— takes one value and outputs one line of pseudonymized value. The key is received through thePII_KEYenvironment variable.masked.csv— mask phone numbers in the form010-****-1234, and hide emails too.report.csv— counts by reason. The total must equal the number of rows in the original.scan.sh— takes one file and ends with a nonzero code if there is personal data.access.log— every time you query, one line with the time, the target, and the purpose.handoff.md— what you send out and what you do not, and why you chose the masking methods as you did.
Notes
There are three spots where people commonly get caught in step 3. If you hash without normalization, the same person is not grouped when only the case differs. If you hash without a key, phone numbers and emails can be brute-forced by generating candidates. And if you truncate to 8 characters, collisions occur. The grader checks the three separately, so you can tell right away which one you tripped on.
Do not write personal data in the access.log of step 7 either. An audit log is also a leak path.
The original mixed with personal data
Create the original (the preparation block above).
Run the preparation block in the instructions as it is. The columns needed for the investigation and the personal-data columns must be there together so that it is practice in picking them out.
Fetch only the columns you need
slim.csv — keep only the columns needed for the investigation. The number of rows stays the same.
Instead of SELECT *, pick only the columns you will use in the investigation. Names and addresses are not needed to look at payment failure reasons. Do not reduce the rows; reduce only the columns — you must not make it unusable by trimming too far.
When you need to group the same person
pseudo.sh — takes one value and outputs one line of pseudonymized value. The key is received through the PII_KEY
environment variable.
Use openssl dgst -sha256 -hmac "$PII_KEY" -r. You must convert to lowercase and remove whitespace (normalization) before hashing so that the same person is grouped to the same value. Do not truncate the output.
Partial masking for visual checking
masked.csv — mask phone numbers in the form 010-****-1234, and hide emails too.
For a phone number, keep only the last four digits (010-****-1234). It cannot be reversed but can be inferred, so use it only when a person has to cross-check by eye.
Only the aggregate goes outside
report.csv — counts by reason. The total must equal the number of rows in the original.
Instead of the original 200 rows, it is one table of 'counts by reason'. Only if the total equals the number of rows in the original can you say nothing is missing.
The pre-export scanner
scan.sh — takes one file and ends with a nonzero code if there is personal data.
It takes one file and ends with a nonzero code if there is personal data. Blocking everything is also a failure — a file containing only aggregates must pass. If it gives false positives, nobody will use it.
Record what you looked at and why
access.log — every time you query, one line with the time, the target, and the purpose.
It is the only means of defending yourself when the question 'who looked at this?' comes later. If the purpose is missing, it does not work as a defense. And do not write personal data in this file either.
What you send out and what you do not
handoff.md — what you send out and what you do not, and why you chose the masking methods as you did.
Masking methods are divided by whether they can be reversed. Write why you used each method where you did. Personal data must not go into this document itself either.