Air-Gapped Sites — Defence and Government
Building the table that links one requirement to one file
Goal
Build an evidence index that binds which files are the evidence for each requirement, seal it with hashes, count the breaks with a verifier, and close the audit response package with copies that have gone through export review.
Why it matters
The items that fail an audit are usually not work that was not done but work that could not be shown. If no path is laid between a line of a requirement and a place in a file, you cannot answer on the spot, and an item you could not answer is recorded the same as work not done. That path cannot be laid the day before the audit — evidence files get overwritten and disappear over time. So the index contains not only the path but the hash of that file as it was then, and when you mask values on the way out, the hash changes along with them. The hard part of this lab is not the rules but this consistency.
Steps
- With
python3, create the requirements, 14 evidence files, the policy, and the previous index in/root/audit/data. Use the generation script as is. - Read the headers of the evidence files and write the evidence list per requirement in
/root/audit/index.json. - Attach a SHA-256 to each piece of evidence and seal it as
/root/audit/index-sealed.json. - Verify the
/root/audit/data/snapshot.jsonsubmitted at the previous inspection and write the result in/root/audit/verify.json. - Divide the evidence into design and operational and write the verdict per requirement in
/root/audit/strength.csv. - Run the same collection again to create
/root/audit/index-rerun.jsonand write the determinism in/root/audit/determinism.json. - Create masked copies with the export review rules under
/root/audit/submit/evidence/and leave/root/audit/redaction.csv. - Seal
/root/audit/submit/index.jsonwith the masked copies' hashes and close the package with/root/audit/submit/verify.jsonand/root/audit/submit/INDEX.md.
Notes
- The reference date is in
/root/audit/data/asof.txt. If you use today's date fromdate, the same data gives a different answer every day. - The headers of the evidence files come in two formats.
.jsonhas them in top-level keys (evidence_id,covers,kind,collected_at), and the others in leading comment lines (like# evidence-id:). - The
pathin the index is a path relative to/root/audit. Only the index inside the package is relative to/root/audit/submit. - The evidence strength classification and the export review rules are in
/root/audit/data/policy.json. Do not hard-code the rules; read them from that file. - The canonical form for step 6 is the SHA-256 of the string made by
json.dumps(..., ensure_ascii=False, sort_keys=True, separators=(",", ":"))on the index withgenerated_atremoved. - Common mistake 1: editing the original evidence files in step 7. What is masked must be the copies, and the originals must remain as they are.
- Common mistake 2: carrying over the original hashes as they are in step 8. The masked copies must carry the masked copies' hashes for the receiving side's verification to pass.
- Standards documents: NIST SP 800-171 Rev 3, NIST SP 800-53 Rev 5, NIST SP 800-92. The REQ numbers and RED numbers in this lab are not control numbers from the standards but synthetic numbers used only within the data.
Create the requirements and evidence files
With python3, create requirements.json, policy.json, snapshot.json, and asof.txt in /root/audit/data, and 14 evidence files under evidence/. Use the generation script that uses no random numbers, as it is.
There is no sample to download in an air-gapped network, so create the data yourself first. Only if it uses no random numbers does everyone get the same data no matter who runs it how many times, and you can check one another's judgments against each other. The grader converts the data to a canonical form and checks fingerprints, so if you edit the data by hand, all the later steps get blocked.
Build an index that links evidence to each requirement
In /root/audit/index.json, put three keys, generated_at, asof, and entries, where entries holds, for each requirement id, a list of evidence items (evidence_id, path, kind, collected_at). Include requirements with no evidence as empty lists too.
The evidence files state in their headers which requirements they cover. Sweep the directory, read the headers, and invert them, and you have the index. Do not forget that there are two formats. Sort the items by evidence id, and write path as a path relative to /root/audit.
Pin the evidence hashes into the index
In /root/audit/index-sealed.json, copy the index as it is but add sha256 to each evidence item. Leave generated_at as the value from the step 2 index.
Sealing is not recollecting but adding a fingerprint to the current index. So the collection time must not change. Write the hash as the SHA-256 over all the bytes of the file, in lowercase hexadecimal.
Run the previous inspection's index through the verifier
In /root/audit/verify.json, write the result of verifying /root/audit/data/snapshot.json. In each of the three keys missing_file, hash_mismatch, and uncovered, put a count and a list (evidence_ids or req_ids).
The three are different failures. A file that is pointed to but does not exist, a file that exists but whose hash is off, and a requirement that has no evidence items at all. The third has to count even the cases where the index did not list the requirement at all, so you have to loop from the requirement list side. A requirement that has one item and that item is broken does not count as no evidence.
Separate the strength of evidence and find the gaps
In /root/audit/strength.csv, put req_id,needs,has_design,has_operational,verdict on the first line and write the 12 requirements one per line. needs is 설계 or 설계+운영 (the Korean words for "design" and "design plus operation"), whether it is held is 예/아니오 (the Korean words for "yes" and "no"), and the verdict is one of 충족·설계근거없음·운영근거없음·근거없음 (the Korean words for "satisfied," "no design evidence," "no operational evidence," and "no evidence").
Configuration and policy documents show that it is set up that way, and logs and inspection results show that it actually worked that way. Which kind is which is in the policy file. If there is no evidence at all, "no evidence" comes first, and after that you check whether the needed side is empty, starting with design.
See whether collecting again gives the same index
Create /root/audit/index-rerun.json anew by the same method, and in /root/audit/determinism.json write four keys: generated_at_differs, normalized_sha256_first, normalized_sha256_second, and same_without_time.
Remove only the collection time from the two indexes, serialize them in canonical form, and compare the hashes. The two hashes must be the same for this collection to be deterministic. Conversely, the collection times must differ for it to count as a rerun, so write the time down to microseconds.
Make masked copies through export review
Apply the policy's review rules in order of rule number to make copies of the 14 pieces of evidence under /root/audit/submit/evidence/ with the same relative paths, and write only the files that triggered a rule in /root/audit/redaction.csv as evidence_id,rules,hits,sha256_after.
Files that triggered nothing also need copies for the package to be complete. Write rules as the triggered rule numbers joined with plus signs, and hits is the total number of places changed. sha256_after is the hash of the masked copy — do not touch the originals.
Close the package and self-check with the verifier
Seal /root/audit/submit/index.json based on the masked copies (paths relative to /root/audit/submit), leave the result of running the same verifier in /root/audit/submit/verify.json, and leave a table of contents that lays out the evidence count and verdict per requirement as a table in /root/audit/submit/INDEX.md.
The index inside the package must use the masked copies' hashes for the receiving side's verification to pass. Missing files and hash mismatches become 0, but requirements with no evidence at all remain as they are — do not hide that gap; write it in the table of contents as a verdict. The table in the table of contents has three columns: requirement id, evidence count, and the verdict from step 5.