TT Lab
Get started
Learn Learning paths Courses

Cron Ran curl at Three in the Morning

Reconstruct one hour of that night

Continue in TT Lab

Goal

Go through an investigation once around with four kinds of data from one incident. You count, pick, preserve, weave a timeline, take a pivot and follow it all the way to the cluster, decide the order of containment, and write a post-incident report.

Why it matters

What you do in the 30 minutes after an alert goes off decides how the incident ends. If in your haste you delete the Pod and reboot the node, the intrusion path disappears with it, and conversely, while you gather evidence perfectly, another namespace gets breached with the leaked credential.

So response is a question not of "what is right" but of "what comes first," and that order has to be decided before an incident happens. Gather the most volatile things first, do not touch the originals, unify times in UTC, and do not postpone revoking credentials — these rules were all born after someone lost something once.

The data in this lab is synthetic. The formats, however, are the same as the real ones. It uses Falco's JSON output, the raw records of the kernel audit, the fields of journalctl -o json, and Kubernetes audit events as they are.

Steps

  1. Count the lines of the four kinds of data and save them to /root/ir/01-sources.tsv.
  2. Save the rule name, time, and PID of the earliest Falco alert to /root/ir/02-first.json.
  3. Copy the four kinds of data to /root/ir/evidence/ and leave a manifest.sha256.
  4. Write a timeline that puts the four sources on one time axis, at least 8 rows, to /root/ir/04-timeline.tsv.
  5. Write four pivots (pid, cron-file, c2-host, k8s-identity) with values and rationale to /root/ir/05-pivot.tsv.
  6. Write the timing and rationale of six containment actions to /root/ir/06-containment.tsv.
  7. Write what the stolen identity did in the cluster to /root/ir/07-k8s.tsv.
  8. Write the post-incident report to /root/ir/08-postmortem.md.

Notes

Count what exists

Count the lines of the four files falco.jsonl, audit.log, cron.jsonl, and kube-audit.jsonl in /opt/lab/fixtures/detection/incident/ and save them to /root/ir/01-sources.tsv, one line each, as 파일이름<탭>줄수 (file name, a tab, and the line count).

wc -l is enough. If you skip this step in an investigation, later you get "oh, that log existed too."

Pick the first alert

Find the earliest alert in falco.jsonl and save its rule name, time, and PID to /root/ir/02-first.json as {"rule": …, "time": …, "pid": …}.

You can pick the earliest with jq -s 'min_by(.time)'. The PID is inside output_fields. Use the time as written in the file.

Preserve without touching the originals

Copy the four kinds of data to /root/ir/evidence/, and create /root/ir/evidence/manifest.sha256 that verifies with sha256sum -c inside that directory.

sha256sum <파일들> > manifest.sha256 is written with relative paths (the placeholder stands for the files). You verify inside that directory with sha256sum -c manifest.sha256.

Put it on one time axis

Write the records of the four sources to /root/ir/04-timeline.tsv, at least 8 rows, in four columns 시각<탭>출처<탭>행위자<탭>무슨 일 (time, source, actor, what happened, separated by tabs), in ascending time order. All four sources must appear at least once.

Unify the source names as falco, auditd, cron, and kube-audit. Convert the epoch seconds of the raw kernel audit with date -u -d @<초> +%Y-%m-%dT%H:%M:%SZ (the placeholder is the number of seconds). The __REALTIME_TIMESTAMP of journald is in microseconds.

Take a pivot and jump

Write four lines to /root/ir/05-pivot.tsv as 축<탭>값<탭>근거 (pivot, value, rationale, separated by tabs). The pivot names are pid, cron-file, c2-host, and k8s-identity, and for the rationale write, in at least 10 characters, which file and what in it you looked at to judge so.

pid is that of the process executed in a temporary directory. cron-file is the scheduled job newly created that night, and it is in the PATH record of the raw kernel audit. k8s-identity is the user.username of the audit log.

What comes first

Write the six containment actions to /root/ir/06-containment.tsv as 조치<탭>시점<탭>근거 (action, timing, rationale, separated by tabs). The timing is one of 지금, 나중, and 안함 (now, later, and never), and the rationale is at least 20 characters.

Gather the most volatile things first. And there is one thing that stays valid even if you turn off the node — that cannot be postponed.

What the stolen identity did in the cluster

Find every request made by system:node:edge-07 in kube-audit.jsonl and write them to /root/ir/07-k8s.tsv in five columns, verb<탭>리소스<탭>네임스페이스<탭>이름<탭>응답코드 (verb, resource, namespace, name, response code, separated by tabs). If there is no namespace or name, write -.

Filter with jq -r 'select(.user.username == …)'. A rejected request (403) is an action too — what was attempted tells you the intent.

A write-up that changes next week

Write six sections, ## 요약, ## 타임라인, ## 영향, ## 원인, ## 탐지 공백, and ## 재발 방지 (in order: summary, timeline, impact, cause, detection gap, and prevention of recurrence), in /root/ir/08-postmortem.md. The body must include the command-and-control address, the namespace where the credential was used, and the path of the persistence mechanism that was caught.

## 탐지 공백 (detection gap) is the most important section in this report. Whether there was no rule, whether the alert went where nobody looks, or whether there was no log at all leads to completely different remedies. A report that ends with a person's name invites the same accident next week.