TT Lab
Get started
Learn Learning paths Courses

In Front of an Unfamiliar System

Collect in order of volatility, then seal it

Continue in TT Lab

Goal

You harden into a file a plan for collecting evidence in the volatility order of RFC 3227, and build a collector that follows that order as is. You write the collection time, the collector, the command, and the sha256 in the manifest, and after sealing the manifest you build a tool that verifies whether the seal is broken, and draw a line at the items that contain customer personal data.

Why it matters

When you go on site, your hands go to the logs first. During the 10 minutes you spend downloading the logs, the process list changes, connections close, and temporary files are deleted. The logs will be in the same place tomorrow, but those things exist only now, and once they are gone you cannot answer the question "was that process running at the time?". Section 2.1 of RFC 3227 says to proceed from the most volatile to the least volatile and gives an example order in seven groups. What this order gives is not only a priority but the basis for your judgment — to the question of why you grabbed that first, you can answer with a class rather than a hunch. Collecting is not enough. A few weeks later the question comes: "is this file really from that server on that day?" For each item you write when, by whom, and with what it was collected, along with a content fingerprint, and you also seal the manifest itself. This is because if you only write the file hashes, all it takes is editing the manifest. Conversely, if you try to collect everything, you end up carrying out the customer's personal data wholesale. Draw the line, but do not leave things out quietly; leave the fact that you left them out, and the reason, in the record. The grader does not trust the wording you write down. It actually runs your collector with its own plan and its own scope file and checks the order, the times, and the hashes, and it secretly edits one file in the bundle that gets made and checks whether your verifier names it.

Steps

  1. Create and run /root/evidence/gen_scene.py to produce /root/evidence/host/. It holds 2 configuration files, 2 logs, 1 copy of the central log, 6 temporary files, and the production DB (customer 120 rows, charge 360 rows).
  2. In /root/evidence/plan.json, write eight items in order, starting with the most volatile. For each item write id, volatility_class, command, and why.
  3. Create /root/evidence/collect.py so that it collects in the plan's order as is and leaves a manifest.
  4. Make it write the sha256 and the byte count for each item, and the collector in the manifest. Add --collector.
  5. Make it seal the manifest itself and leave MANIFEST.sha256.
  6. Create /root/evidence/verify.py to verify whether the seal is broken, and make it print the name of the broken file.
  7. In /root/evidence/scope.json, write the items you decided not to collect and the reasons, and add --scope so that the excluded items remain in the record.
  8. Build the real bundle in /root/evidence/bundle/, verify it, and then write /root/evidence/evidence_report.md in four sections.

Notes

id What it is Where to read it
proc_table The list of processes running now The process table of this Pod
net_state The routing table and open connections /proc/net/route, /proc/net/tcp
kernel_stats Kernel statistics and memory /proc/stat, /proc/meminfo
tmp_files What remains on the temporary filesystem host/tmp
app_logs Application logs host/var/log
app_db The production database host/var/lib/app.db
remote_logs A copy received from the central log server host/remote
host_config Configuration and physical layout host/etc

Get the site in hand

Create and run /root/evidence/gen_scene.py to produce /root/evidence/host/. It holds 2 files in etc, 2 in var/log, 1 in remote, 6 in tmp, and var/lib/app.db (customer 120 rows, charge 360 rows).

In a lab Pod you cannot touch a real customer server, so you only build the disk side. You do not need to imitate the highly volatile things — because /proc of this Pod is real. After you build it, skim it once with tree.

Build the plan in volatility order

In /root/evidence/plan.json, write eight items in order, starting with the most volatile. For each item include id, the RFC 3227 volatility_class (1..7), the command you will actually run, and why, which states why it is in that position.

Use the seven groups of RFC 3227 Section 2.1 as the classes as they are. The process table, the kernel statistics, and the routing table are in the same group, then the temporary filesystem, then the disk, then the remote logging data, then the physical layout. command must actually run and produce something.

Collect according to the plan

Create /root/evidence/collect.py so that it collects in the plan's order as is and leaves <묶음>/manifest.json (the placeholder is the bundle). For each item write id, order, volatility_class, command, collected_at, file, exit_code, and status.

Do not re-sort the plan by class — the moment you sort, the judgment hides inside the code. Run the commands with bash and put the standard output as is into <id>.txt. Only if you write the time separately for each item is the order proven.

Write when, who, and with what

Add --collector <이름> (the placeholder is the name), and make it write the sha256 and bytes for each collected item, and the collector in the manifest.

A few weeks later the question comes: 'is this file really from that server on that day?' What lets you answer then is the collection time, the collector, the command that was actually run, and the content fingerprint. Compute the sha256 by reading the whole file.

Seal the manifest too

After the manifest is fully written, compute the sha256 of that file and leave it in <묶음>/MANIFEST.sha256 (the placeholder is the bundle) as one line, <해시> manifest.json (the placeholder is the hash).

If you only write the file hashes in the manifest, all it takes is editing the manifest. A seal does not prevent forgery, but it catches changes that happen without anyone knowing — cases where someone opened it in an editor and saved it, or it was truncated during a copy, actually happen far more often.

Name what broke

Create /root/evidence/verify.py. With --bundle <묶음> (the placeholder is the bundle), compare the seal and the file hashes, end with 0 if intact and with a nonzero value if broken, and print to standard output a line containing the name of the broken file.

A verifier that prints only 'the seal is broken' keeps a person standing on the spot. You have to name what broke so that the next action exists. Catch both the case where the manifest changed and the case where a collected file changed, and for an excluded item it is normal for there to be no file.

Leave what you decided not to collect in the record

In /root/evidence/scope.json, write policy and excluded to exclude app_db, and add --scope <파일> (the placeholder is the file) to collect.py. For an excluded item you do not create a file; it remains in the manifest with status excluded and a reason.

The production DB of a payment server contains names, emails, and phone numbers. The moment you dump that and carry it out, we are no longer the people investigating but a new risk. What matters is not leaving it out quietly — the fact that you left it out has to remain in the record so that you know later what to request.

Build the bundle and report

Build the real bundle in /root/evidence/bundle/ (giving both the plan and the scope), verify it with verify.py, and then write /root/evidence/evidence_report.md in four sections: ## 무엇을 어떤 차례로 모았나 ## 무엇을 모으지 않았나 ## 봉인과 검증 ## 남은 위험 (the Korean headings mean "What was collected, and in what order", "What was not collected", "The seal and the verification", and "The remaining risks").

Do not write the report by hand; generate it from the manifest. The item names and classes, the excluded items and reasons, the seal value, and the verification command must all be included so that the next person can verify it again as is.