TT Lab
Get started
Learn Learning paths Courses

Cron Ran curl at Three in the Morning

What comes first after the alert

Continue in TT Lab

In one line

Response is a fight over order. Containment erases evidence, and gathering evidence enlarges the damage. So you decide in advance what to do now and what to do later, and during an incident you follow that table.

Why you need this

What happens in the 30 minutes after an alert goes off is usually this. Someone deletes the Pod. Someone reboots the node. Someone deletes the cron file. An hour later, when asked "how did they get in," nobody knows. They have erased everything that held the intrusion path.

The opposite direction is a failure too. While you refrain from touching anything because you want to gather evidence perfectly, another namespace gets breached with that credential.

So the response procedure is a question not of "what is right" but of "what comes first." This is also why NIST SP 800-61 is cited for so long — containment, eradication, and recovery come after detection and analysis, and it tells you to decide the criteria for judgment in between before an incident.

How it works

The order used in practice can be reduced to four steps.

1. 범위 정하기   무엇이 있는지 세고, 각 자료가 언제부터 언제까지인지 적는다
2. 증거 보존     원본은 건드리지 않고 사본 + 해시. 휘발성이 큰 것부터
3. 타임라인      출처가 다른 기록을 하나의 시간축(UTC)에 올린다
4. 격리·근절     표에 따라 조치. 각 조치에 근거 한 줄

The order of volatility is the key. Memory disappears the moment you kill the process, the process list and open sockets disappear on reboot, and files on disk mostly remain. So the answer to the question "should we kill the process now" is almost always "after we capture the memory."

Conversely, revoking credentials cannot be postponed. Even if you turn off one node, a certificate that leaked from that node continues to be used anywhere in the cluster. In Kubernetes, a node's identity is system:node:<이름> (the placeholder is the node name), and the Node authorization mode decides what that identity can read. The Secrets used by the Pods scheduled on that node are included, so one node's credential is the entire set of secrets of the workloads on that node.

When building a timeline, it is better to nail down the format in advance. Times in UTC ISO 8601, source names in fixed words, and the actor and what happened each in its own column. If you do that, it can be merged even when several people write it in parts, and the story shows up from sorting alone.

The containment table is not written anew for every incident; you make it in advance and only fill in the values of the day. For example, it looks like this.

조치           시점    근거
메모리 수집    지금    프로세스를 건드리면 사라진다. 되돌릴 수 없다
네트워크 차단  지금    외부로 나가는 것만 끊는다. 프로세스는 살려 둔다
프로세스 종료  나중    메모리를 뜬 뒤. 지금 죽이면 그 증거가 함께 죽는다
예약 작업 제거 나중    지속화 수단은 원본을 보존한 뒤에 치운다
자격증명 폐기  지금    노드를 꺼도 유출된 인증서는 클러스터에서 계속 쓰인다
노드 재설치    나중    디스크 이미지를 뜬 뒤. 마지막 단계다

With a table, there is less to decide at three in the morning. The less there is to decide, the fewer the mistakes.

And the way work actually proceeds in an investigation is the pivot. You take one value — a PID, a file path, a destination address, an identity — and look for the same value in other data. A single PID in the host log leads to a kernel audit record, from there the path of the certificate that was read comes out, and the identity of that certificate matches user.username in the API server audit log. That connection turns "the node was breached" into "the secrets in the payments namespace were read."

What it looks like in the field

The section most often missing from a post-incident report is the detection gap. Everyone writes what happened and how it was fixed, but nobody writes "why did nobody know for three hours." Yet it is that section that reduces the next incident. Whether there was no rule, whether there was a rule but the alert went to a channel nobody watches, or whether there was no log at all leads to completely different remedies.

The second thing that collapses most often is the integrity of the evidence. If you open the original file in an editor to look at it and press the save button, from that moment the file is no longer evidence. The 30 seconds it takes to make a copy and leave a hash prevents that.

The third is a report that ends with a person's name. A report that ends with "the person in charge made a mistake" invites the same accident next week. What needs fixing is usually on the side of "why was that mistake possible" — that anyone could write to that scheduled-job directory, that the node certificate was on the same filesystem as the workloads, that nobody was on the alert channel.

The fourth is timeline time zones getting mixed. Falco alerts are left in UTC, the journal uses the host's local time when displaying on screen, and the raw audit log is epoch seconds. If you paste these three together as they are, a three-hour incident looks like six hours or the order is reversed. So the first column of the timeline is always unified to UTC, and the original format is appended in the last column if needed.

What you will do in the next lab

In the lab you receive four kinds of data from one real incident — Falco alerts, the raw kernel audit, cron's journal, and the Kubernetes audit log — and go through an investigation once around. You count the data, pick the first alert, preserve the evidence with copies and hashes, put the four sources on one timeline, take pivots to jump from the host to the cluster, decide the order of six containment actions with their rationale, and finally write a post-incident report. It deals only with files, so it runs in the lab Pod.