TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

Events Are Not Logs

Continue in TT Lab

In one sentence

Events are the only place where an Operator speaks to the user, but they are a summarizing mechanism with short retention in which the same reason is merged into one. Only when you write to fit this property do they become valuable.

Why this place was needed

While building an Operator, you run into a situation like this. The controller knows that "the image tag does not exist, so reconciliation was halted." But the only place that fact lives is the controller Pod's log. The developer who created the custom resource has neither permission to see that log nor the knowledge that such a thing exists. On that person's screen there is just one resource where nothing is happening.

You might think you can just write status, but status is the current state. It can say "it is not Ready right now," but it cannot hold "it failed once 10 minutes ago and then recovered." To leave the process, you need a different place with a time axis, and that is the Event.

How it works

An event is an API object and, like other objects, is stored in a namespace. But there are two APIs.

v1 (core) events.k8s.io/v1
Target involvedObject regarding
Body message note
Reporter source.component reportingController, reportingInstance
Time firstTimestamp, lastTimestamp eventTime
Repetition count series.count

Only the names differ, and the storage is the same. If you query an event created with the new API through the old API, it comes out as it is, but a field with no counterpart in the old schema looks empty. The new API exists because the schema was tidied up to reduce the burden event writes put on the API server in large clusters, and the old API remains for compatibility.

reason is the most important field in this object. It is because it is not a sentence for humans to read but a key that machines count. So the convention is a single short CamelCase word (FailedScheduling, Pulled, BackOff), and if you put a long sentence here, the same incident is recorded each time with a different reason and the aggregation collapses. The details belong in message.

A repeat of the same reason is not a new event. When the event recorder meets the same target, the same reason, and the same message again, it does not create another object; it raises count and updates lastTimestamp. What appears in the list view as (x12 over 3m) is the result. This aggregation is why etcd does not overflow with events even when the reconcile loop runs several times per second.

type is not checked by the API. Even if you write Critical, the object is created. But the tools that read events assume there are only two, Normal and Warning — kubectl events --types= does not accept any other value at all. It is a typical place where you must follow the convention even though it is not enforced.

And events disappear. The default of kube-apiserver's --event-ttl is 1 hour. If you start investigating after an hour, there are no events for that incident.

What you see in the field

First, an Operator with no reason on the user's screen. The controller log has a clear error, but the CR's describe shows nothing. The user files an inquiry, and the operator grabs the log and pastes it in. What three event lines could have settled is being handled with human round trips. Even just leaving the start, success, and failure of reconciliation greatly reduces inquiries.

Second, the incident of using events as an audit record. You designed on the belief that "to see who changed what and when, just look at the events," and then when the incident investigation begins, an hour has passed and there is nothing. What must be kept for a long time should be the conditions in status or records sent to external storage. Events are a window for seeing what is happening now.

Third, the scheduler's sentence is the diagnosis. When a Pod does not start, what people actually read is a line like 0/3 nodes are available: 3 Insufficient cpu. It is also a model of a good event message — it has numbers, it has the name of what is lacking, and it lets you guess what to do next. An Operator's events are also worth aiming at this level.

Limits of this lab environment

The Pods in the kwok cluster are fake, so you cannot see the events the kubelet leaves (Pulling, Started, BackOff). Instead, the scheduler really runs, so FailedScheduling is actually recorded. And for a custom resource that has no controller, you create and put in the events yourself in this lab — it amounts to doing by hand, once, what an Operator would do.

What you will do in the next lab

You create events for a custom resource with each of the two APIs, and check how an event that omits its target is rejected and how a type outside the convention disappears from the tools. You raise the aggregation fields yourself to see how the notation in the list view changes, collect even the real events the scheduler leaves, and finally build a table that reconstructs an incident by reading only the events.