Having Permission Does Not Mean You Should Look
In one line
An FDE is in a position where their hands can reach the customer's real data. What you must learn before technique is the habit of seeing only the minimum, moving only the minimum, and leaving a trace.
Why this was needed
While investigating, you naturally end up typing a command like this.
SELECT * FROM orders WHERE status = 'FAILED' LIMIT 100;
This does not contain only the order number. Names, phone numbers, addresses, and sometimes payment information come out with it. And the result stays in the terminal scrollback, in the captured image, in the Slack message, and in a temporary file on your laptop.
Accidents arise not from malice but from convenience. "I pulled it out as a CSV and emailed it to see it quickly" is the sentence that appears most often in real incident reports.
Four principles
1. Fetch only the columns you need
Instead of SELECT *, specify only the columns needed for the investigation. There are almost no cases where names and addresses are needed for the root cause analysis.
-- 나쁨
SELECT * FROM orders WHERE status='FAILED';
-- 좋음
SELECT id, created_at, status, error_code FROM orders WHERE status='FAILED';
2. Pseudonymize identifiers by purpose If linking to the same person is truly necessary, apply HMAC-SHA-256 with a per-purpose secret key to the normalized identifier, and keep an output of sufficient length.
uid = HMAC_SHA256(secret_from_kms, normalize(email))
This value is not an anonymized value but a pseudonymized identifier. If the secret key or the surrounding data is exposed, it can be linked back, so you control access to it at the same level as the original. You share the same purpose key only when linking between systems has been approved, and for analyses that must not be linkable to each other, you use different keys per purpose.
3. Do not send it outside You handle data where it is. The moment you download it locally or send it by messenger, it is out of your control. If unavoidable, you move it only in aggregated form — instead of 100 original rows, one table of "counts by failure reason".
4. Leave a record Record what you looked up and when. It is the only means of defending yourself when the question "who looked at this?" comes later. Most systems have an audit log, and if there is none, your work notes play that role.
Masking is divided by whether it can be reversed
| Method | Reversibility | When |
|---|---|---|
| Deletion | Impossible | Columns not needed for the investigation |
| A sufficiently long HMAC with a per-purpose key | Hard without the key | When you need to group the same person within an approved scope |
Partial masking (010-****-1234) |
Impossible, but can be inferred | When visual checking is needed |
| Tokenization (mapping kept) | Possible | When you will need the original later |
| Encryption | Possible | When moving or storing is unavoidable |
A plain hash or a public salt alone is not safe. Phone numbers and emails are easy to brute-force by generating candidates, and if you truncate the output, as with 8 characters, the risk of collisions also grows. A different salt for each record cannot group the same person to the same value, so when a repeatable link is needed, you use an access-controlled per-purpose secret key and a sufficiently long HMAC.
When you build test data
Copying production data for use in tests is the most common and the most dangerous practice. If it is truly necessary, build it in a way that keeps the structure and changes the values — names from a dictionary, emails with the domain set to example.com, and amounts with noise added while keeping the distribution. This way the properties needed for reproduction (length distribution, duplicates, missing values) remain and the personal data is gone.
What it looks like in the field
- A log screenshot pasted in the outage channel has a customer email as is → leaked to everyone in the channel.
- An investigation CSV left on the laptop desktop and the project ends → an accident years later.
- The result of
SELECT *pasted as is into a ticket → kept permanently in the ticket system.
Logs and error reports are the most common leak path
Personal-data incidents do not arise only from a database being breached. They leak more often from the records that a properly working system faithfully leaves.
The habit of logging the whole request is the most dangerous. If a line put in for debugging stays in production, resident registration numbers and card numbers flow into the log store. Logs are usually not encrypted, have long retention periods, and have broad read permissions.
# 이렇게 두면 어느 날 반드시 샌다
log.info("요청: %s", request.json())
# 필요한 것만, 식별자는 해시로
log.info("결제 요청 user=%s amount=%s", hash_id(user.id), amount)
Error tracking tools send local variables along. Sending the variables of the frame where the exception occurred is that tool's strength, but if a password or token is in that frame, it goes along as is. You must put a place to filter before sending.
def before_send(event, hint):
for k in ("password", "token", "authorization", "ssn", "card"):
_scrub(event, k)
return event
Do not put personal data in URLs. The query string stays in access logs, proxy logs, browser history, and the Referer header. Even the same value, if you put it in the body, does not stay in these four places.
The third leak path is people. While investigating an outage, people download production data to a laptop, paste a table into the company messenger, and put a screenshot into a document. Technical controls have difficulty preventing this, so removing the reason to download in the first place is what works: a read-only lookup screen, a masked admin screen, and an analysis environment that automatically erases query results.
The records you leave will be requested someday. Manage as a list what goes into which log, and set the retention period. When a deletion request comes, if you deleted only the database and it remains in the logs and backups, you have not deleted it. Not leaving it in the first place is always cheaper than deleting it.
What you will do in the lab that follows
You receive 200 order rows mixed with personal data and go as far as building the one sheet that may be sent outside. You leave each of the four principles as a file, and at the end you build a pre-export scanner yourself.
The grader runs that scanner on a leaked file and on a clean file. A scanner that blocks everything fails, and so does one that passes everything — because if it gives false positives nobody will use it, and if it misses, it is useless.