The same thing happens in a real implementation
Goal
You confirm that the rules you wrote by hand in the previous three labs hold in a real implementation too. You start three etcd members in the same Pod, move the leader, kill one, and then kill two.
Why it matters
A toy implementation is good for learning but is not grounds for trust. Only when you see the same numbers and the same words (term, index, quorum) come out of the etcd that all of Kubernetes's state sits on do you confirm that the rules you built earlier are not toy rules.
The most important step is the sixth. Everyone expects writes to be blocked when two of the three stop, but reads are blocked too. This is because etcd's default read is a linearizable read, so the leader answers only after a majority confirms "I am still the leader." A cluster that has lost its majority does not get slower; it stops. This is the real reason to use an odd number of nodes and to calculate quorum.
Another lab's etcd is running separately on 2379 in this Pod. So this lab uses ports starting from 22379.
Steps
- Create
/root/etcd/start.shand start a, b, and c. - Write the current leader and term in
/root/etcd/leader.json. - On
/raft/lab, writehello-raft, read it from a different member, and record it in/root/etcd/kv.json. - Move the leader with
move-leaderand record it in/root/etcd/move.json. - Try writing with one follower stopped and record it in
/root/etcd/one-down.json. - Try writing and reading with b and c stopped and record it in
/root/etcd/quorum-loss.json. - Bring everything back and record it in
/root/etcd/recover.json. - Write the verdict from lining it up with the previous labs in
/root/etcd/verdict.json.
Notes
- Endpoint bundle:
--endpoints=127.0.0.1:22379,127.0.0.1:22479,127.0.0.1:22579 - Status table:
etcdctl --endpoints=... endpoint status --write-out=table - Health check:
etcdctl --endpoints=... endpoint health - Stop:
pkill -9 -f /root/etcd/b· Revive:bash /root/etcd/start.sh - If you stop gracefully (without
-9), the leader looks for somewhere to hand over and never manages to shut down. This is a place to simulate a failure, so stop it by force. - Common mistake 1: the
--initial-clusterstrings of the three members differing. If even one character differs, they do not form a cluster. - Common mistake 2: sending
move-leaderto a follower endpoint. You must send it to the current leader. - When there is no response, add
--command-timeout=4s. With the default it waits a long time.
Start three members with static bootstrap
Create /root/etcd/start.sh and start the three members a, b, and c on client ports 22379, 22479, and 22579 and peer ports 22380, 22480, and 22580. It must be safe to run multiple times.
Static bootstrap is nothing more than the three members having the same --initial-cluster string. If even one character differs, they do not see each other as the same cluster. In this Pod another etcd started by kwok is using 2379, so avoid that port. Put the data directory at /root/etcd/<이름> (the member name), and block a second start with pgrep when it is already running.
Read out who the leader is
Use endpoint status to check the current leader and the raft term, and write them in /root/etcd/leader.json as leader and raft_term.
etcdctl --endpoints=... endpoint status --write-out=table shows IS LEADER, RAFT TERM, and RAFT INDEX one line at a time. These are the same values you built by hand in the previous labs: the term goes up each time an election is counted, and the index is the position number in the log.
The same value whichever one you ask
On the key /raft/lab, write hello-raft, read it from the two members other than the one you wrote to, and then write key, value, and read_from in /root/etcd/kv.json.
A write goes through to the leader whichever endpoint you send it to. Reads are also linearizable reads by default, so even if you ask a follower, it goes through the leader's confirmation, and that is why you get the same value whichever one you ask. In step 6 you will see how this property breaks.
Moving the leader raises the term
Hand the leader over to another member with move-leader, and write the name and term before and after in /root/etcd/move.json as before, after, before_term, and after_term.
etcdctl move-leader <멤버 id> (the member id goes in the placeholder) must be sent to the endpoint that is currently the leader. If you send it elsewhere, it is rejected. The member id appears in hexadecimal in the ID column of endpoint status. After handing over, look at the term again: when the leader changes, the term goes up with it. It is the same thing as the new leader raising the term and being elected in the previous lab.
Writes still work when one is down
Try writing a value with one follower stopped, and write stopped, write_ok, and healthy in /root/etcd/one-down.json.
Choose what to stop by the data directory path, as in pkill -9 -f /root/etcd/b; going by name alone would also catch other etcd processes running in this Pod. Two of the three is a majority, so writes go through as before. In healthy, write the number of live members.
With two down, reads stop too
With b and c stopped, try both writing and reading on /raft/ghost, and write stopped, write_ok, read_ok, and error in /root/etcd/quorum-loss.json. When you have confirmed it, start them again.
Writes being blocked is as expected. What is surprising is that reads are blocked too: etcd's default read is a linearizable read, so the leader can answer only after a majority confirms "I am still the leader." In error, write the message etcdctl produced as it is. When you revive them, even if --initial-cluster-state new is attached, etcd uses the data directory if it exists.
What remains after they come back
Revive all three members and read /raft/lab and /raft/ghost each, then write healthy, survived, and ghost_present in /root/etcd/recover.json.
A value finalized while a majority existed survives, and a value attempted without a majority leaves no trace. It is the same as ghost disappearing from the old leader's log in the previous lab, and the difference is that here the client never received ok in the first place, so this side is safer.
Is it the same as the rules you built by hand
In /root/etcd/verdict.json, write quorum_of_3, tolerated_failures, write_without_quorum, linearizable_read_without_quorum, and raft_index_seen.
This is the step where you line up what you wrote as rules in the previous three labs with what you saw in this lab. What is the majority of three nodes, and so how many failures can it tolerate? What happens to writes and reads when there is no majority? In raft_index_seen, write the RAFT INDEX from endpoint status as is; this value does not decrease.