Air-Gapped Sites — Defence and Government
Erase it and the analysis dies — the craft of changing it instead
In one line
You have to export logs, but people and organizations are stained into them. If you delete, analysis becomes impossible, and if you export as is, export review fails. What sits between the two is deterministic pseudonymization — the same value always becomes the same pseudonym, so correlation analysis survives, and the key that turns it back into the original stays only inside.
Why this was needed
When an outage occurs at an air-gapped site that you cannot solve there, you end up sending logs outside. Whether it is the headquarters development team or the manufacturer, the people who know that code are outside the network. But operational logs contain employee IDs, accounts, internal addresses, device identifiers, and remarks that people wrote by hand, all as they are. Export review rejects on the basis of those.
The first reaction is usually masking. You replace all the employee IDs with asterisks and cut off the addresses. Then the review passes, but the analysis does not. The fact used most often in outage analysis is something like "the same user failed the same request five times within 3 minutes," and that fact comes not from the values themselves but from the relationship that the values are equal. If you turn everything into the same asterisks, the five lines become strangers to one another, and if you turn everything into different random numbers, it looks as if five people failed once each. Either way, the cause cannot be found.
So what you need is a function that changes values so that the same value stays the same, different values stay different, and it cannot be turned back. The standard tool that satisfies all three conditions at once is a keyed hash, that is, HMAC. The structure is defined by RFC 2104, and US federal standards fixed the same thing as FIPS 198-1. In Python, a single line with the standard library hmac is enough.
How it works
First, separate masking from pseudonymization. Masking removes the value and pseudonymization replaces the value. What you must remove is places where you cannot count what is in them. A remarks field that people write freely can contain phone numbers and names, and whatever regular expression you write, the next person writes it in a new form. It is more honest to delete such a field wholesale. Conversely, fields that a machine stamps out in a fixed template are ones where you can know what comes in, so you pseudonymize them.
Second, hang the pseudonym on a key. You must not make it with just sha256(사번) (the placeholder is the employee ID). Employee IDs have a narrow format, so there are only tens of thousands of possible values, and if the receiving side hashes them all and builds a table, it reverses in a few seconds. The same is true of phone numbers and resident-number delimiters. If you mix in a key, someone who does not know the key cannot build that table. The key must not go out with the data, and if the key changes the pseudonyms change wholesale, so record inside when you used which key.
Third, separately count the places the regular expression misses. Places with a label attached, like emp=E24-0101, are easy. The hard ones are three — values embedded like a sentence inside a log message body, values that a person typed into a free-text field, and values hidden behind an encoding such as base64. The first two never come out from a field-by-field sweep, and the third never comes out from any regular expression. You have to decode first and then search. The moment you say "we handled everything with regular expressions" without counting these three, the export leaks.
Fourth, look at the risk that remains even after deleting. Even if you change every direct identifier, if you look at department, grade, and access time together, rows remain that narrow down to one person. Such columns are called quasi-identifiers, and a state in which fewer than k rows share the same combination is treated as a risk. The way to reduce it is not to delete but to make it coarser. If you drop minute-level times to hour-level, rows with the same combination cluster together and the number of unique rows falls. What you lose is precision and what you gain is exportability, and writing this trade in numbers is the output of this step. The US standards agency's de-identification overview NIST IR 8053 and NIST SP 800-188 discuss this trade at length.
Fifth, do not put the mapping table in the export. The table that turns pseudonyms back into the originals is indispensable for investigation, but if it goes out with the data, it is the same as not having pseudonymized. Keep the table in the inside vault and send only pseudonyms outside. General log management is summarized in NIST SP 800-92.
What it looks like in the field
Once I had an export review rejected twice. The first was for the expected reason — the employee IDs were still there. I fixed it and resubmitted, and it was rejected a second time, and the reason was the remarks field. There was one line where the person in charge had written a contact number after "called," and that line did not match any regular expression we had built. Since then we made it a rule to delete free-text fields wholesale. If it was truly needed for analysis, a person reread it and copied only the needed sentences by hand.
Another thing I remember is encoding. In an authentication failure log, a token fragment was carried in base64, and when I decoded it, it was an account address. It passed every check that searches the export for account strings. That was because the searching side searched only for the plain text. It is the reason that when you build a self-check, you should look not only at "is the value that was in the original absent from the export" but also at "is it absent in encoded form too."
What you will do in the next lab
You create 96 lines of access log, 16 lines of app log, and a roster of 18 people, and separate which fields are direct identifiers and which are quasi-identifiers. You count separately the values the regular expression found and the ones it missed, implement a key-bound deterministic pseudonymization function, and show by test whether the same value becomes the same pseudonym and whether it changes when you change the key. Next you apply pseudonymization and masking separately, count the rows that become unique by the quasi-identifier combination and reduce them by generalization, and then export the export package after a machine self-check.