TT Lab
Get started
Learn Learning paths Courses

The Language of Banking

Can you still follow the case after you have masked it?

Continue in TT Lab

In one line

Removing personal information from logs and exported files is not "making it invisible" but keeping the connectivity the investigation needs while making people unrecognizable. What sets that boundary is the masking rules and the deterministic pseudonym values, and whether it was done properly is proven by measuring again with a machine.

Why this was needed

In the course of an incident investigation, you end up receiving application logs whole. But the transfer API's logs have free-text sentences mixed in, and it is common for account numbers, card numbers, phone numbers, and emails to be written in those sentences as they are. It was not written by a developer with bad intent; it is one line that printed "the whole request message" to find outages quickly and has stayed for years.

To hand these logs to an analyst or a vendor, you have to strip out the personal information. The Personal Information Protection Act requires personal information to be destroyed without delay once its purpose has been achieved (Article 21), and what an investigation needs is not a person's identity but connectivity such as "how many times did it fail on the same account." But if you cover account numbers entirely with asterisks, that connectivity disappears too. You can no longer tell whether the same person failed three times or three people failed once each. It must be covered, yet it must be possible to link them. Satisfying these two requirements at the same time is the whole of this topic.

How it works

Scrubbing splits into three stages. Find, sort out, and replace.

Finding is regular expressions. With the re module, you sweep for account number shapes, 16-digit numbers, mobile phone shapes, and email shapes. The problem is that a regular expression looks only at "shape." In bank logs, 16-digit numbers are scattered around besides card numbers. Settlement sequence numbers, message trace numbers, and batch job identifiers are all 16 digits. If you cover those with asterisks too, the identifiers the investigation needs disappear.

Sorting out comes in here. A card number has a check digit attached. Starting from the second digit from the right end, you double every other digit, and if the doubled value is 10 or more, you add its digits (or subtract 9). If the sum is a multiple of 10, it passes. It is commonly called the Luhn algorithm, and ISO/IEC 7812-1, the standard for card number systems, specifies this check digit method (the text of the standard is paid, so this article does not link to it). The probability that a random 16-digit number passes this check by chance is one in ten, so attaching the check digit reduces false positives tenfold. It also means they do not disappear completely.

Replacing does two things together. One is masking that keeps only the last four digits so a person can confirm by eye, and the other is a pseudonym token that lets a machine link records. You make the token with the hmac module. You must not hash the value itself. Account numbers have few possibilities, so a plain SHA-256 is reversed by a dictionary attack. Only an HMAC that uses a secret key cannot be reversed by a party that does not know the key. You make the key with the secrets module or openssl rand, and do not keep it in the same place as the logs.

You put the purpose into the token's message too. This is to make acct|91231457820 and card|91231457820 different values. If you do this, even if the token table of one side leaks, you cannot match it against the other.

원본  "계좌 912-31-457820 카드 9012345678901234 일련번호 9900112233445566"
          │                    │                      │
          │                    │                      └ 검사숫자 실패 → 카드 아님 → 그대로 둔다
          │                    └ 검사숫자 통과 → ************1234
          └ ***-**-**7820  +  tokens.acct = ["acct_1f3c…"]  (같은 계좌면 언제나 같은 값)

What it looks like in the field

First, a scrubber that gives different results when run twice. Pipelines retry. If you scrub a line that is already covered and it counts the asterisks again in ***-**-**7820, or recomputes the token and gets a different value, reconciliation breaks. You leave a marker of already-processed, and make the rules not apply again to their own output.

Second, confusing covering with deleting. Masking that keeps the last four digits cannot be reversed, but for someone who already has a list of account numbers, it is a clue that narrows the candidates. In a file that handles only a few accounts, the last four digits alone can even identify one person. So in an exported file, you do not use the masked value itself as the investigation key; you use the token.

Third, pseudonymization without key management. If the key is kept in the same directory as the logs, embedded in code, or nobody knows when the key changed, the meaning of the pseudonym values differs every time. If you write a key identifier (the first digits of the key's hash) in the report, you can later say "which key were these tokens made with."

What really matters in practice

What you will do in the next lab

You build 640 lines of a synthetic transfer log yourself, find candidates, sort cards from sequence numbers with check digits, and build an export copy with masking and pseudonym tokens attached. Then, without ever using the original account numbers, you trace the incident customer's case with tokens alone, run twice to check that the results are the same and that the remaining leaks are 0, and finish with a policy and a report. The grader recomputes the tokens with your key and directly runs your programs with numbers it makes itself.