FDE Capstone: The Warehouse Got the Same Order Three Times
Logs that cannot leave: a support bundle the customer runs themselves
In one line
What you need for a customer who cannot send logs outside is not "please send us the logs" but a collector that the customer can run on their own server, review with their own eyes, and then approve. Only if it gathers by an allowlist, records the number of redactions, and bundles so that the same input gives the same bytes can the security owner approve it.
Why this was needed
On the third day of the outage, the customer's operations team said, "Our internal policy does not let the logs leave. If you organize what you need, we'll pull it for you." The FDE sent a few file names, and the compressed file that came back contained the original configuration file. The DB password and the payment API key were right there. After learning that the file had gone out, the customer's security team said it would review every following bundle in advance, and the reviews started taking two days each.
The problem was not the people but the procedure. If a person picks what to include each time, they pick differently each time, and if you do not write down what was redacted, the reviewer has no choice but to read the file from start to finish. For a reviewer to approve quickly, three things must be visible. What went in (the allowlist), what was redacted and how many times (the manifest), and whether the file you just received is the very file that was reviewed (the hash). A support bundle collector is what makes these three identically every time, with a script instead of a person.
How it works
1. Gather by an allowlist. "Everything except the files that contain secrets" is a denylist, and a denylist lets through files it does not know. You decide only what goes in by name, such as the version (VERSION), the configuration (under etc), the current logs (*.log directly under logs), and the service's environment. The rotated app.log.1, the customer data directory, and symbolic links are not in the rules, so they naturally drop out. If you follow a link, a file outside the server (for example, the host's authentication file) can come along with it.
2. Read the service's environment, not the shell's. The environment variables of the shell that ran the collector have nothing to do with the outage. On Linux, /proc/<pid>/environ, which proc_pid_environ(5) describes, holds the environment at the time the process started, separated by NUL bytes. As the same documentation points out, values the process changed itself after starting are not reflected. The order can differ from run to run, so you turn it into lines and then sort by name.
3. Redact narrowly and precisely. If the rule is wide, even settings such as password_min_length: 12 or token_ttl_sec: 3600 disappear and the bundle becomes useless, and if it is narrow, secrets leak. That is why you define them by shape, like "values whose key name ends with password, secret, token, or api_key", "the value after Bearer", "an email with a domain", and "a private key block from BEGIN to END". The order matters too. A private key is multi-line, so you have to replace the whole block before the line-by-line rules. In the redacted place you leave the kind, as in [REDACTED:token]. The reviewer can count the markers and check them against the manifest's numbers, and the FDE can narrow down the cause from the mere fact that "a token was here".
4. Truncate after redacting. Logs are large, so you set an upper limit per file and keep only the recent lines. If you truncate first and redact later, the BEGIN line of a private key block that straddles the truncation boundary is gone, so it does not match the rule and only the body is left. Not cutting in the middle of a line is for the same reason.
5. Only deterministic values in the manifest. For each file, you write the path, size, sha256, the number of redactions, and whether it was truncated. You are tempted to put in a generation time, but the moment you do, a bundle made from the same input is different every time.
6. The same bytes for the same input. The archives documentation of reproducible-builds.org summarizes that tar records the modification time, file order, owner name and number, and permissions, which makes the result fluctuate, and recommends --sort=name (from 1.28), --mtime, and --owner=0 --group=0 --numeric-owner for GNU tar. If you bundle with Python, you can set mtime, uid, gid, uname, gname, and mode directly on the TarInfo of the tarfile documentation, and there is one more layer. The gzip documentation says that if you do not give the mtime of GzipFile, the current time goes into the header, and if you need a result that does not depend on time, you should give mtime=0.
Below is the result of waiting 1 second on this lab image, changing only the mtime of the source, bundling again, and comparing the sha256 (measured: GNU tar 1.35, gzip 1.12, Python 3.12.3).
| Bundling method | sha256 of two bundlings |
|---|---|
tar -czf with no options |
Differs |
tar --sort=name --mtime=@0 --owner=0 --group=0 --numeric-owner -czf |
Same |
Adding a directory with Python tarfile.open(..., "w:gz") |
Differs |
TarInfo time fixed but no mtime given to GzipFile |
Differs |
TarInfo time, owner, and permissions fixed + GzipFile(mtime=0) |
Same |
In the measurement, GNU tar's -z put 0 in the gzip header time. On the other hand, Python's w:gz put the current time in the header, so the hash differed even when the file contents were the same.
support-bundle/
VERSION
env.txt ← run/environ 을 줄로 바꿔 정렬
etc/... ← 가린 설정
logs/app.log ← 가린 뒤 최근 줄만
manifest.json ← path·size·sha256·redactions·truncated
What it looks like in the field
The conversation with the customer's security owner changes. Before, you had no answer to "How can you guarantee that this file has no sensitive information?" With a collector, you can say "The allowlist is these eight lines, the redaction rules are these four, and in this bundle we redacted 33 tokens and 45 emails. The sha256 of the file you received is this value." The reviewer counts the markers, compares the hash, and approves.
Each customer also has additional things to redact. Things like employee numbers, internal hostnames, and internal ticket numbers. If you hard-code these into the collector code, you end up managing a different version for each customer, and the customer cannot fix the rules themselves either. It lasts longer to receive them as a rules file, apply them, and show their kinds together in the manifest's counts.
What really matters in practice
- An allowlist, not a denylist. You do not include files you do not know.
- Over-redaction is also a failure. Check together, in line with the number of redactions, whether non-secret settings were left.
- Redact and then truncate, and truncate at a line boundary.
- Only if the bundle is deterministic does the statement that the reviewed file and the received file are the same hold. You write the time outside the bundle (in the delivery record).
- Receive customer rules as a file, not as code.
What you will do in the next lab
You decide the collection scope on the copy of the customer server given as material, and grow the redaction filter and the collector in turn. You add the manifest, the log size cap, a deterministic tar.gz, and the customer rules file one step at a time, and at the end you make the delivery bundle and the hash and redaction report from the material. The grader runs your collector directly on a server tree where it has planted new secrets each time.