TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

It is in the logs but the user cannot see it - making your operator speak

Continue in TT Lab

Goal

Create events yourself with the two event APIs to check the required fields and aggregation, see how they disappear from the tools when you break the convention, and then build a table that reconstructs an incident by reading only the events.

Why it matters

There are only three places where a user can learn what an Operator has done — the controller log, the resource's status, and events. The log is seen only by cluster operators, and status says only the current state and does not hold the process. The only place users actually read is the Events section at the bottom of kubectl describe. That is why half of building an Operator is deciding what to leave as events. But events are not logs. Default retention is one hour, a repeat of the same reason is merged into an increase of count rather than a new object, and there are two APIs but one storage. If you do not know these properties and expect an audit trail from events, there is nothing left when you actually investigate an incident.

Steps

  1. In /root/op-events/pipeline-crd.yaml, write the CRD pipelines.ev.labhub.io — group ev.labhub.io, kind Pipeline, plural pipelines, a single version v1, with only spec.stage (string) in the schema. Create the namespace op-events, apply Pipeline build-1 (stage build) with /root/op-events/pipeline-build1.yaml, and then save its metadata.uid to /root/op-events/pipeline-uid.txt on one line.
  2. In /root/op-events/event-started.yaml, write a v1 Event build-1-started — involvedObject is Pipeline build-1 (apiVersion, kind, name, namespace, and the actual uid), reason is ReconcileStarted, type is Normal, message is one sentence, firstTimestamp and lastTimestamp are the current time, count is 1, and source.component is pipeline-operator. After applying, save the output of kubectl -n op-events events to /root/op-events/events-list.txt.
  3. In /root/op-events/event-bad.yaml, write an Event build-1-orphan but leave out involvedObject entirely (reason NoTarget, type Normal, message one sentence). Try to apply it and collect the result in /root/op-events/invalid-event.txt — the first line is apply-rc=<종료 코드> (the exit code), and below it you paste the sentence the server produced exactly as it was.
  4. In /root/op-events/event-weird.yaml, write an Event build-1-weird — the target is Pipeline build-1, reason is StageUnknown, and write type as the out-of-convention value Critical (fill in the time fields in the same format as in step 2). After applying, run kubectl -n op-events events --types=Critical and kubectl -n op-events events --types=Warning in turn and collect the results in /root/op-events/type-report.txt — it must include the output of both commands and each exit code (critical-rc=, warning-rc=).
  5. Assume the step 2 event happened four more times and update the aggregation fields — use kubectl -n op-events patch event build-1-started --type=merge to change count to 5 and lastTimestamp to the current time. Then save the output of kubectl -n op-events events to /root/op-events/count-report.txt.
  6. In /root/op-events/event-done.yaml, write an events.k8s.io/v1 Event build-1-done — the target is regarding (Pipeline build-1, including the actual uid), reason is ReconcileSucceeded, note is one sentence, type is Normal, eventTime is the current time (6 digits after the decimal point), reportingController is pipeline-operator, reportingInstance is pipeline-operator-0, and action is Reconcile. After applying, pull a list with each of the two APIs and save them to /root/op-events/both-apis.txt — the lines starting with core: and the lines starting with new: must be the name lists from kubectl get events and kubectl get events.events.k8s.io respectively.
  7. In /root/op-events/hungry-pod.yaml, write a Pod hungry — one container (app, image busybox:1.36) that requests requests.cpu of "64". Apply it, wait until the scheduler leaves an event, and then save the output of kubectl -n op-events get events --field-selector reason=FailedScheduling -o wide to /root/op-events/scheduler-event.txt.
  8. Create /root/op-events/timeline.sh — print every event in op-events as one line each in the form <type> <reason> <대상종류>/<대상이름> (type, reason, target kind, and target name), sorted and with duplicates removed. Do not include values that change every time you look, such as age, time, or count. Save the output to /root/op-events/timeline.txt, and write in /root/op-events/incident.txt, in sentences a person can read, what happened from that table alone — it must mention both the reason the scheduler left and the reason the Operator left, and must point out that an out-of-convention type is mixed in.

Notes

Create the target to attach events to

In /root/op-events/pipeline-crd.yaml, write the CRD pipelines.ev.labhub.io — group ev.labhub.io, kind Pipeline, plural pipelines, a single version v1, with only spec.stage (string) in the schema. Create the namespace op-events, apply Pipeline build-1 (stage build) with /root/op-events/pipeline-build1.yaml, and then save its metadata.uid to /root/op-events/pipeline-uid.txt on one line.

An event is always "a story about some object." So you must write not only the target's kind, name, and namespace but also the uid, so that it is not mixed up with the story of a different object recreated with the same name. You will use the uid as it is in the next steps.

The Operator leaves its first words

In /root/op-events/event-started.yaml, write a v1 Event build-1-started — involvedObject is Pipeline build-1 (apiVersion, kind, name, namespace, and the actual uid), reason is ReconcileStarted, type is Normal, message is one sentence, firstTimestamp and lastTimestamp are the current time, count is 1, and source.component is pipeline-operator. After applying, save the output of kubectl -n op-events events to /root/op-events/events-list.txt.

reason is not a sentence for humans to read but a key that machines count. That is why the convention is a single short CamelCase word, and when the same reason repeats, it is merged into one rather than becoming a new event. Put the detailed explanation in message. Even if the target is a custom resource, it is attached at the bottom of kubectl describe just the same.

You cannot create an event with no target

In /root/op-events/event-bad.yaml, write an Event build-1-orphan but leave out involvedObject entirely (reason NoTarget, type Normal, message one sentence). Try to apply it and collect the result in /root/op-events/invalid-event.txt — the first line is apply-rc=<종료 코드> (the exit code), and below it you paste the sentence the server produced exactly as it was.

An event has meaning only when it has a target. Without a target it cannot attach to any object's describe, and only the reason floats around in the list. That is why the API rejects it outright, and if you read exactly which field the error points to, you can also learn the namespace relationship between the event and its target.

An out-of-convention type disappears from the tools

In /root/op-events/event-weird.yaml, write an Event build-1-weird — the target is Pipeline build-1, reason is StageUnknown, and write type as the out-of-convention value Critical (fill in the time fields in the same format as in step 2). After applying, run kubectl -n op-events events --types=Critical and kubectl -n op-events events --types=Warning in turn and collect the results in /root/op-events/type-report.txt — it must include the output of both commands and each exit code (critical-rc=, warning-rc=).

The API server does not check the type value. So the object is created without a problem. The problem comes after that: the tools that read events are built on the assumption that there are only two, Normal and Warning. This is why you must follow the convention even though it is not enforced.

A repeat of the same reason is not a new event

Assume the step 2 event happened four more times and update the aggregation fields — use kubectl -n op-events patch event build-1-started --type=merge to change count to 5 and lastTimestamp to the current time. Then save the output of kubectl -n op-events events to /root/op-events/count-report.txt.

When the event recorder meets the same target, the same reason, and the same message again, it does not create a new object and only edits these two fields. See for yourself how the LAST SEEN column of the list view changes then — the notation in parentheses is the answer to this step. It is also the reason etcd does not blow up even when the reconcile loop runs several times per second.

There are two APIs but one storage

In /root/op-events/event-done.yaml, write an events.k8s.io/v1 Event build-1-done — the target is regarding (Pipeline build-1, including the actual uid), reason is ReconcileSucceeded, note is one sentence, type is Normal, eventTime is the current time (6 digits after the decimal point), reportingController is pipeline-operator, reportingInstance is pipeline-operator-0, and action is Reconcile. After applying, pull a list with each of the two APIs and save them to /root/op-events/both-apis.txt — the lines starting with core: and the lines starting with new: must be the name lists from kubectl get events and kubectl get events.events.k8s.io respectively.

The new API has different field names — involvedObject became regarding, message became note, and source was split into reportingController and reportingInstance. But the place they are stored is the same, so they are visible through the old API as well. If you check which fields are empty when viewing through the old API, the difference between the two schemas stands out.

Receive an event left by a real controller

In /root/op-events/hungry-pod.yaml, write a Pod hungry — one container (app, image busybox:1.36) that requests requests.cpu of "64". Apply it, wait until the scheduler leaves an event, and then save the output of kubectl -n op-events get events --field-selector reason=FailedScheduling -o wide to /root/op-events/scheduler-event.txt.

This Pod cannot fit on any node. The scheduler writes that fact not only in the Pod's status but also as an event, and the message contains "how many nodes cannot, and why" as it is. This is the sentence people actually read when looking for the cause. The event takes a few seconds to appear, so wait with a conditional loop.

Reconstruct an incident from events alone

Create /root/op-events/timeline.sh — print every event in op-events as one line each in the form <type> <reason> <대상종류>/<대상이름> (type, reason, target kind, and target name), sorted and with duplicates removed. Do not include values that change every time you look, such as age, time, or count. Save the output to /root/op-events/timeline.txt, and write in /root/op-events/incident.txt, in sentences a person can read, what happened from that table alone — it must mention both the reason the scheduler left and the reason the Operator left, and must point out that an out-of-convention type is mixed in.

The reason events are valuable in an incident investigation is that "who made what judgment about what" remains one line at a time. But default retention is one hour, so if you go in late, there is nothing. That is why what must be kept for a long time should be not events but conditions in status or records sent to external storage.