TT Lab
Get started
Learn Learning paths Courses

Air-Gapped Sites — Defence and Government

Making a people-stained log fit to leave the network

Continue in TT Lab

Goal

Pick out the identifiers from an operational log stained with people and organizations, apply key-bound deterministic pseudonymization and masking separately, reduce the remaining re-identification risk by generalization, and then export the export package after a machine self-check.

Why it matters

When an outage occurs in an air-gapped network that you cannot solve there, a moment comes when you have to send logs outside. Covering employee IDs with asterisks at that point passes review but kills the analysis. That is because the facts outage analysis uses are usually not the values themselves but the relationship that the values are equal. So what you need is a function that always changes the same value into the same pseudonym but cannot be turned back, and a keyed hash fills that role. The hard part of this lab is not the pseudonym function but counting what counts as an identifier. Labeled fields are easy, and values embedded like sentences in the message body and values hidden behind encoding are what leak the export.

Steps

  1. With python3, create six data files in /root/deid/data. Use the generation script as is.
  2. Write the classification, the number of distinct values, and the handling for each field in /root/deid/classify.csv.
  3. Write the values the regular expression found in the fields and the values it missed in /root/deid/scan.json.
  4. Build a key-bound pseudonym function and leave /root/deid/vault/map.csv and /root/deid/vault/proof.json.
  5. Apply pseudonymization and masking separately to create /root/deid/out/access.deid.log and /root/deid/out/app.deid.log.
  6. Count the rows that become unique by the quasi-identifier combination in /root/deid/risk.json, and leave the generalized analysis table in /root/deid/out/events.csv.
  7. Write the result of checking the export against the original identifier list in /root/deid/selfcheck.json.
  8. In /root/deid/release, build an export package containing only the pseudonymized data, the methodology document, and integrity.

Notes

Create the data that will be the basis of de-identification

With python3, create six files in /root/deid/data: access.log, app.log, roster.csv, policy.json, key.hex, and asof.txt. Use the generation script that uses no random numbers, as it is.

There is no sample to download in an air-gapped network, so create the data yourself first. Only if it uses no random numbers does everyone get the same data no matter who runs it how many times, and you can check one another's judgments against each other. The grader converts the data to a canonical form and checks fingerprints, so if you edit the data by hand, all the later steps get blocked.

Separate what each field is

In /root/deid/classify.csv, put field,category,distinct_values,action on the first line and write one line for each of the fields in policy.json. distinct_values is the number of distinct values that field has in the data, and action is the handling that action_by_category sets for that category.

The classification is set by the organization, so it is already written in policy.json. Your job is to attach the handling that fits that classification and to actually count the data to fill in the number of distinct values. A column with few distinct values cannot point to a person by itself, but combined with other columns it can — that is a quasi-identifier.

Count together what the regular expression found and what it missed

In /root/deid/scan.json, write four items: field_values (the number of distinct field values by kind), message_only (values that are only inside msg and in no field), encoded (values that come out when you base64-decode payload), and note_only (values only inside note). The last three items are sorted lists.

Find labeled places with 종류=값 (the placeholders are the kind and the value). What comes next is the core of this step — sweep separately between the double quotes inside msg and note, and decode payload first before sweeping it. No regular expression can find an encoded value in its plain-text state.

Build a key-bound deterministic pseudonym

In /root/deid/vault/map.csv, put kind,value,pseudonym on the first line and write one line for each identifier found in the data, and in /root/deid/vault/proof.json, write six items: values, pseudonyms, collisions, stable, key_sensitive, and alt_key_overlap.

Build the pseudonym with HMAC-SHA256. If you feed in the kind and the value joined by a vertical bar, the pseudonyms separate even if an employee ID and an address happen to be the same string. There are three tests — is computing the same value twice the same, do different values never become the same pseudonym, and does computing with the other key in policy.json give a pseudonym set that does not overlap at all. This vault stays inside the network.

Apply pseudonymization and masking separately

Create /root/deid/out/access.deid.log and /root/deid/out/app.deid.log. Replace free-text fields wholesale with the mask_token from policy.json, replace encoded fields with the pseudonym of the decoded value, and replace all remaining patterns with pseudonyms. The number of lines and the line order must be the same as the original.

The order changes the result. If you do not mask first, values inside the free text are turned into pseudonyms and left there, and then what should have been deleted goes out undeleted. For an encoded field, decode it, decide which kind it is, and then attach the pseudonym of that kind — it has to become the same pseudonym as where the same person appeared in plain text.

Count the risk that remains even after changing every direct identifier

In /root/deid/risk.json, write six numbers — rows, k, groups_before, unique_before, groups_after, and unique_after — and in /root/deid/out/events.csv, leave a generalized analysis table with pseudo_emp,dept,grade,hour on the first line, in the line order of the access log.

The quasi-identifier combination is written in quasi in policy.json. Department and grade come from the roster, and the time comes from the access log. First group by minute-level time and count the rows where fewer than k rows share the same combination, then drop the time to hour-level and count again. The amount it decreased is the safety bought by generalization, and the precision lost is its price.

Check by machine whether the original remains in the export

In /root/deid/selfcheck.json, write four numbers: files_checked, forbidden_values, hits, and pseudonyms_found. forbidden_values is the total number of identifiers that were in the data, and hits is how many of them were found in the export, which must be 0.

The forbidden list is wider than the pseudonym vault. Values removed by masking have no pseudonym, but they must not appear in the export either. Put even the values inside free text into the list. The check target is the files in the out directory, and the forbidden list and the selfcheck result are kept outside it — the list itself is the most dangerous file.

Keep the vault inside and send out only the package

In /root/deid/release, create the three files access.deid.log, app.deid.log, and events.csv, plus /root/deid/release/method.md and /root/deid/release/SHA256SUMS. The methodology document has five sections with the headings ## 무엇을 바꿨나, ## 가명은 어떻게 만들었나, ## 무엇을 지웠나, ## 남은 위험, and ## 무결성 (the Korean headings mean, in order, "What was changed," "How the pseudonyms were made," "What was deleted," "Remaining risk," and "Integrity"), and the mapping table and key do not go into the package.

The receiving side sees only pseudonyms. So if an explanation of what was changed by which rules does not go along, that data becomes unreadable data. Compute the integrity hash after the last time you edit the files. And look once more at a listing of what went into the package — the accident of including the vault is the most common.