TT Lab
Get started
Learn Learning paths Courses

There Were Two Leaders

Lose the quorum and it does not slow down, it stops

Continue in TT Lab

One-line summary

etcd is a store that implements Raft directly, and the terms and quorum you built by hand earlier show up under the same names. The most commonly misunderstood part is reading: if you lose the majority, not only writes but default reads stop too.

Why this was needed

A toy implementation is good for learning but is not grounds for trust. Whether the rules you built are toy rules or real rules can be confirmed only by seeing the same numbers and the same words come out of a real system. etcd is the best target for that check. Every Kubernetes object lives in it, and its documentation describes itself as a "general-purpose foundation for large distributed systems" that stores metadata in a consistent and fault-tolerant way (etcd: why etcd).

How it works

If you start three members with the same --initial-cluster string, they recognize each other as one cluster. The static bootstrap example in the documentation gives a name, a peer address, a client address, and a cluster token together. As the documentation states, the reason for the token is to give each cluster a unique token so that different clusters do not get mixed up (etcd clustering).

Once it is up, you see the values you built in the earlier modules.

ENDPOINT          ID                IS LEADER   RAFT TERM   RAFT INDEX
127.0.0.1:22379   99f0eb44090117f8  false       2           14
127.0.0.1:22479   e64076ee26a8ab0d  false       2           14
127.0.0.1:22579   8e05ae0f7f2c67be  true        2           14

RAFT TERM is the term and RAFT INDEX is the position number in the log. If you move the leader with etcdctl move-leader, the term goes up by one; it is the same thing as the new leader raising the term and being elected earlier.

The quorum table is in the documentation as is. One member has a majority of 1 and tolerates 0 failures; 3 members: 2 and 1; 5 members: 3 and 2; 7 members: 4 and 3. The documentation recommends keeping the cluster at seven members or fewer, and states that an odd-sized cluster tolerates the same number of failures as an even-sized one while needing fewer nodes (etcd FAQ).

The most important part is reading. etcd's default read is a linearizable read, and the documentation defines linearizability as every operation applied by concurrently running processes appearing to take effect instantly at some point between its invocation and its response. It then goes on to say that linearizability comes at a cost, because a linearizable request must go through the Raft consensus process (etcd API guarantees). That is why a cluster that has lost its majority does not merely get slower; even reads stop. The same document explains that if you set the request's consistency mode to serializable, you can access data that may be stale relative to the quorum, but you get rid of the performance burden that comes from linearizable access depending on live consensus. In other words, if you can accept a stale value, that path exists.

By this point you can see that the rules you used in the previous three labs and the sentences in this documentation say the same thing. The "more than half" that we wrote in has_majority is the documentation's (n/2)+1, what commit_index was counting is RAFT INDEX here, and the new leader that was elected by raising the term is the RAFT TERM after move-leader. Only the names differ.

What it looks like in practice

A large share of the cases where a Kubernetes cluster "can't do anything" are this. If two of the three control plane nodes go down at the same time, kube-apiserver cannot even serve reads even though it is alive, because etcd has lost the majority and cannot confirm reads. No matter how much you look into the one remaining machine, the cause is not in that machine.

The etcd documentation distinguishes that a temporary loss of quorum automatically and safely resumes when the network returns, while a permanent loss of quorum is fatal. The meaning of that distinction in operations is one thing: when you have lost quorum, do not hastily delete or add members; first check whether there are members that can come back.

One more thing. The reason the ports used in this lab are not the standard 2379 is that another lab control plane in the same Pod is already using that port. The very fact that moving the ports alone gives you one more cluster also means that etcd tells clusters apart only by their bundle of addresses and token.

What you will do in the next lab

You start three etcd members in the same Pod with static bootstrap. You move the leader and watch the term rise, write with one member stopped, and try both writing and reading with two stopped. At the end, you line up the rules you built by hand in the previous three labs with the numbers you saw here.