In Front of an Unfamiliar System
Collect what disappears first
In one line
Collect evidence starting with what disappears fastest. You have to harden the order of collection and the line of what you decided not to collect into code, not a document, so that the next person can do the same thing again.
Why this was needed
When you go into an outage site, your hands go to the log files first. Logs are visible, they can be read right away when opened, and they are easy to copy. But during the 10 minutes you spend downloading the logs, the process list changes, open connections close, kernel statistics are updated, and temporary files are deleted. The logs will be in the same place tomorrow, but those things exist only now.
The difference of those 10 minutes decides the conclusion later. If you cannot answer the question "was that process running at the time?", you cannot rule out one hypothesis, and a hypothesis you could not rule out stays with the investigation scope left wide.
How it works
This problem was sorted out long ago. Section 2.1 of RFC 3227 "Guidelines for Evidence Collection and Archiving" says to proceed "from the most volatile to the least volatile", and gives an example order for a typical system in seven groups.
1 레지스터, 캐시
2 라우팅 테이블, ARP 캐시, 프로세스 표, 커널 통계, 메모리
3 임시 파일 시스템
4 디스크
5 그 시스템과 관련된 원격 로깅·모니터링 자료
6 물리 구성, 네트워크 토폴로지
7 보관 매체
What this order gives is not only a priority. It gives the basis for your judgment. To the question of why you grabbed the process list first, you can answer not "a hunch" but "it is group 2 and the disk is group 4". So make the collection plan a file that records a class for each item, and have the collector follow the order in that file as is, without sorting. The moment you sort, the judgment hides inside the code.
Section 2.2 of the same RFC also lists what to avoid. Things like not rebooting or shutting down the system, not trusting programs that could be evidence, and not using tools that alter the data.
How do you prove you collected it
Collecting is not enough. A few weeks later the question comes: "is this file really from that server on that day?" So for each collected item you record three things together.
- When — the collection time. Write it separately for each item. A single time for the whole bundle cannot prove the order.
- Who — the collector. There has to be someone to ask later.
- With what — the command that was actually run. Not "I got the logs" but that command.
On top of that you add a content fingerprint. If you compute a sha256 for each file and write it in the manifest, you can later do the same computation again and know whether even one character has changed. There is a reason to go one step further here — if you only write the file hashes in the manifest, all it takes is editing the manifest. So you leave the hash of the manifest itself in a separate file. This is the seal.
A seal does not prevent forgery. There are plenty of people who can edit both the manifest and the seal file. What the seal prevents is changes that happen without anyone knowing — someone opened it in an editor and saved it, it was truncated during a copy, or the archive was corrupted. In practice, accidents like these happen far more often than malicious forgery.
What it looks like in the field
First, if you try to collect everything, you end up carrying out the customer's data wholesale. The production database of a payment server contains names, emails, and phone numbers. The moment you dump that and put it on your laptop, we are no longer the people investigating but a new risk. So in the collection plan you also write the items you decided not to collect and the reason in the same file.
What matters is not quietly leaving it out but leaving the fact that you left it out in the record. If the manifest says "this item was excluded, and the reason is this", then when that data becomes necessary later you know right away what to request. If you leave it out quietly, nobody knows it existed.
Second, there are cases where the collection command changes the data. Things like opening a log file so that an editor creates a lock file, or connecting to a database and causing it to update its statistics. If there is a way to open it read-only, use it, and if not, write down what you changed.
Third, write times in a standard notation. If the server's time zone and our laptop's time zone differ, a whole hour ends up off when you line up the timeline later. Write UTC in the RFC 3339 notation, and also put the server's time zone setting in as a separate collection item.
Fourth, leave the verifier along with it. If you leave only the seal value, the next person has to compute hashes by hand and compare. If you leave a verifier with it, it takes one line, and above all it names what broke. A verifier that only prints "the seal is broken" keeps a person standing on the spot.
What really matters in practice
- Collect in the order things disappear. Write the order in the plan file, and the collector follows that order as is.
- Write when, who, and with what for each item. Not once per bundle.
- Seal the manifest too. With file hashes alone, editing the manifest is the end of it.
- Leave what you decided to omit in the record. If you leave it out quietly, nobody knows it existed.
What you will do in the next lab
You hold one payment server of (fictional) Seojin Chemical. The highly volatile things you actually read from /proc of the lab Pod, and the things left on disk you read from the place that was set up. You give each item an RFC 3227 class to build the collection plan, build a collector that collects in that order, write the collection time, the collector, the command, and the sha256 in the manifest, and seal the manifest. Then you build a tool that verifies whether the seal is broken — the grader secretly edits one file in your bundle and checks whether your verifier names it. Finally you draw a line at the items that contain customer personal data and check that the fact of leaving them out remains in the record.