TT Lab
Get started
Learn Learning paths Courses

Cron Ran curl at Three in the Morning

Where does 'who did what' actually land?

Continue in TT Lab

In one line

Detection does not start with a rule. It starts with whether that event is recorded anywhere in the first place. On a Linux host there are three places for that, and what the three see does not overlap — kernel audit (auditd) sees syscalls, journald sees what programs say about themselves, and the Kubernetes audit log sees requests that arrived at the API server.

Why you need this

Suppose a scheduled job sent a request outside at three in the morning. To prove that sentence when you come in to work, you need to know four things. What was executed, who scheduled it, when that schedule was created, and where that request went.

The wall people hit most often here is "we are collecting all the logs, yet there is no answer." That is because collecting logs and leaving records that can answer a question are different things. /var/log/syslog keeps what cron executed, but it does not have who created that cron file and when. Conversely, kernel audit records exactly the moment a file is created, but it does not look at paths you have not set a watch rule on at all.

So the first job of detection engineering is not writing rules but surveying the sources. You write the questions down first, and then check one by one whether records that can answer them are currently being left. If they are not, getting them left is the first task, and only after that do you earn the right to write rules.

How it works

The kernel audit subsystem accepts two kinds of rules. One is a file watch and the other is a syscall filter.

-w /etc/cron.d -p wa -k cron_change
   └ 경로      └ 권한  └ 검색용 키
   이 디렉터리에 쓰기(w)나 속성 변경(a)이 일어나면 남긴다

-a always,exit -F arch=b64 -S execve -F exe=/usr/bin/curl -k outbound_exec
   └ 언제        └ 아키텍처   └ 시스콜  └ 추가 필터        └ 키
   64비트 execve 중 실행 파일이 curl 인 것만 남긴다

-w is really a shorthand for a syscall rule, and the key attached with -k later becomes the search key for ausearch -k. The rule syntax and filter fields are in auditctl(8), and the format for writing in a rules file is in audit.rules(7). A rule typed in by hand with auditctl disappears on reboot, so to make it survive you write it in /etc/audit/rules.d/ and let augenrules(8) merge and load it. When reading, ausearch(1) slices by key, user, and time, and aureport(8) produces aggregates.

journald is of a completely different nature. It is not something the kernel observes; it writes down what programs say about themselves. So the single (root) CMD (…) line cron leaves is there because cron is kind, not because the kernel guarantees it. In exchange, journald also carries structured fields — things like _PID, _UID, _SYSTEMD_UNIT, and SYSLOG_IDENTIFIER, and the list is in systemd.journal-fields(7). If you extract with journalctl -o json, these fields come out as they are and a machine can read them (journalctl(1)).

Kubernetes is yet another layer. What happens in front of the API server is not left in the host logs. The audit log is off by default and is turned on only when you give a policy with --audit-policy-file. A policy picks a level for each rule: None leaves nothing, Metadata leaves only the requester, time, resource, and verb, Request leaves the request body too, and RequestResponse leaves the response body as well. There are four stages too — RequestReceived, ResponseStarted, ResponseComplete, and Panic. This choice is directly both the cost and the resolution of the evidence.

What it looks like in the field

The first is the case of "there is a rule, but in the wrong place." It is common for a host to watch /etc/cron.d but not /usr/local/bin. An attacker does not need to create a new scheduled job — changing one line in the contents of a script that is already scheduled is enough. At that moment the audit log is silent. Silence does not mean safe; it means nobody is looking at that place.

The second is the case of "there is a record, but it cannot be read." Kernel audit leaves numbers that are hard for people to read. The uid is a number and the syscall is a number. ausearch -i turns them into names, and I have seen several teams that, not knowing this option and staring at the raw form, reached the conclusion "our audit log is useless."

The third is on the Kubernetes side. If a node is breached, that node's kubelet certificate goes out with it. The host log shows only that a process read the file, and what was done in the cluster with that certificate exists only in the API server audit log. If you cannot join the two records, the investigation ends at "one file was read."

What you will do in the next lab

In the lab you go directly into a VM, put in audit rules, cause an event on purpose, and pull out the record with ausearch. Then you find the same event again on the journald side and compare what each of the two records has. Finally, you pick a place that is not watched, see for yourself that no record is left, and add a rule that covers that place with a new key. A Pod has no kernel privileges at all, so this lab runs in a VM.