TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

The Moment There Are Two Leaders

Continue in TT Lab

In one sentence

Leader election is an agreement in which "who is the leader right now" is written on one small object (a Lease) and renewed periodically. Whether it has expired is not written on the object and is computed by each instance with its own clock, so in principle there is an interval in which there are two leaders.

Why this mechanism was needed

If you run only one Operator, nobody reconciles while that Pod is dead. So you run two or more. But the reconcile loop is code that changes cluster state. If two run at the same time, they try to create the same resource twice, erase each other's results, and revert the scale value back and forth.

The simplest solution is "let only one work." Several instances are running, but only one of them reconciles and the rest wait on standby. The problem is how to decide on that one, and if you set up a separate consensus system, that itself becomes another point of failure. Kubernetes used what was already there — the API server's optimistic concurrency. When many try to update one object at the same time, only the one whose resourceVersion matches succeeds. Leader election is a very thin agreement layered on top of that property.

How it works

The vessel for the agreement is the Lease of coordination.k8s.io/v1. It has five fields in all.

Field Meaning
holderIdentity The name of the instance that claims to be the leader now
leaseDurationSeconds If this much time passes after renewals stop, it is treated as expired
acquireTime When the current holder acquired the lease
renewTime When the holder last renewed
leaseTransitions The number of times the holder has changed

acquireTime and renewTime are MicroTime, not an ordinary Time, so they need six digits after the decimal point. If the format differs, it is rejected at the parsing stage, not at the value stage.

What the leader does is only periodically rewrite renewTime. It touches neither the holder nor the transition count. What a candidate does is periodically read the lease and check whether renewTime + leaseDurationSeconds is in the past relative to now. If it is in the past, it treats the lease as expired and takes it over under its own name, writes a new acquireTime, and raises leaseTransitions by one. That is why the transition count is an operational metric. If the holder keeps changing, that number keeps growing.

Three flags of kube-controller-manager set this rhythm — --leader-elect-lease-duration (default 15 seconds), --leader-elect-renew-deadline (default 10 seconds), and --leader-elect-retry-period (default 2 seconds). The relationship matters. The renew deadline must be shorter than the lease duration. If the leader cannot renew within the renew deadline, it steps down from the leader position on its own, and if this deadline is longer than the lease duration, an interval arises in which "the leader still believes it is the leader, but another candidate has already seen it as expired and taken it." The retry period is set much shorter so that several attempts can be made within that deadline.

The permissions are surprisingly small. Three verbs, get, create, and update, on leases in the coordination.k8s.io group are enough. There is nothing to delete and no need for a watch — candidates check expiry not by watching but by periodic queries.

What you see in the field

First, the interval with two leaders cannot be eliminated. A leader whose network briefly dropped believes it is still the leader and keeps reconciling, and in the meantime another instance takes over the lease and starts reconciling. Eliminating this interval would require stronger guarantees from a distributed system, and Kubernetes leader election does not promise that. That is why the reconcile loop itself must be idempotent. Leader election is a device that lowers the probability, not one that guarantees mutual exclusion.

Second, an incident caused by clock skew. If the nodes' clocks are off by a few seconds, the expiry judgment comes out differently on each node. The symptom is "the leader keeps changing," and all you see on screen is leaseTransitions rising by the minute. When investigating, looking only at the lease object does not reach the cause — you must look at the node clocks and API server latency together.

Third, the traces two leaders leave. If two reconcile loops try to create the same resource, one fails with a name conflict, and if they try to write the same field with server-side apply, a field manager conflict occurs. The latter is especially useful — because the API server tells you by name that "this field belongs to another manager," a single conflict message reveals who was writing together. If you build an Operator with server-side apply, this incident does not pass quietly.

Limits of this lab environment

In a lab Pod, you cannot bring up two real controllers to make them compete. This is because kwok's Pods are fake and containers do not run. So renewal, expiry, and takeover are imitated by hand with patches, and clock skew is reproduced by writing renewTime in the past. On the other hand, the API server is real, so MicroTime format rejection, RBAC decisions, and server-side apply's field manager conflicts can all be checked as real behavior.

What you will do in the next lab

You create a lease and renew it, and build a script that judges expired leases versus live leases. You imitate a takeover with a patch to confirm the transition count rises, and see how a renewal is blocked when you break the MicroTime format. After setting up by hand the minimum permissions needed for leader election, you receive the conflict the API server leaves when two instances try to write the same resource, and finally you bundle it all into an operations checklist.