Isolate, write to both sides, then heal
Goal
You detach one of the three nodes from the others and bring it back, and you see for yourself the state of having two leaders at the same moment and the fate of the values written to each side.
Why it matters
Incidents in distributed systems usually come not from "failure" but from incomplete information. A leader that has been cut off does not know it has been cut off. All it knows is "responses haven't been coming lately," and a Raft leader has no mechanism for stepping down on its own. So it keeps acting like a leader, accepts client values, and returns ok. At the same moment, on the other side, a majority gathers and elects a new leader with a higher term.
At this point there really are two leaders. What Raft prevents is not having two leaders but having two leaders in the same term, and the leader on the minority side cannot get confirmation from a majority, so it cannot finalize anything. The moment things recover, its unfinalized entries quietly disappear. This lab leaves you with one question: what happens to the client that received ok in the meantime?
Steps
- Put the three files in
/root/raft/, start three nodes, and write the result in/root/raft/leader.json. - In
/root/raft/net.py, writeshould_deliver(src, dst, cut). - Write the value
alpha, let it be finalized, and record it in/root/raft/before.json. - Make the leader and the other two cut each other off, and record it in
/root/raft/isolated.json. - Wait for a new leader to be elected and record it in
/root/raft/two-leaders.json. - Write
ghostto the old leader and record it in/root/raft/minority.json. - Write
realto the new leader and record it in/root/raft/majority.json. - Remove the cuts and record the result in
/root/raft/verdict.json.
Notes
- Cut:
curl -s -XPOST -d '{"peers":[2,3]}' 127.0.0.1:5001/cut - Uncut:
curl -s -XPOST -d '{"peers":[]}' 127.0.0.1:5001/cut - Viewing the log:
curl -s 127.0.0.1:5001/log - Common mistake 1: applying the cut on only one side. If the opposite direction is still alive, it is not a partition.
- Common mistake 2: reading in a hurry before the new leader has been elected. An election needs as much time as the timeout.
Bring up the three nodes
Place node.py, rules-for-partition.py, and log-for-partition.py from /opt/lab/raft/ as node.py, raftrules.py, and raftlog.py in /root/raft/, start three nodes, and write the elected leader to /root/raft/leader.json.
The rules you wrote in the previous two labs are provided as finished versions. The only code you write in this lab is one partition switch, and the rest is time spent experimenting by hand. In leader.json, write leader and term.
Reject messages instead of cutting the wire
In /root/raft/net.py, write should_deliver(src, dst, cut). Messages from peers listed in cut are neither sent nor received.
This Pod does not have permission to set up a firewall. So instead of "cutting the wire," you make a node reject requests from specific peers. To the peer this is indistinguishable from no response coming back, so from the consensus algorithm's point of view it is the same as a real partition. dst is the node itself, and cut is the list of peers it has cut off. Apply it with curl -s -XPOST -d '{"peers":[2,3]}' 127.0.0.1:5001/cut and confirm it with cut in /status.
Finalize one value before the partition
Write the value alpha to the leader, confirm that commit is 1 on all three nodes, and then write value, commit, and committed_on in /root/raft/before.json.
To tell later what gets discarded and what remains, there has to be one value finalized before the partition. committed_on is the number of nodes that committed that value. The write is curl -s -XPOST -d '{"value":"alpha"}' 127.0.0.1:<리더포트>/client, with the leader's port in place of the placeholder.
Isolate the leader
Make the leader and the other two cut each other off, and right after the cut, write the old leader's state in /root/raft/isolated.json as old_leader, state_after_cut, and cut.
You have to apply it on both sides for a real partition; if you apply it on only one side, one direction still gets through. Watch whether the cut-off leader steps down by itself. A Raft leader has no mechanism for demoting itself. It keeps believing it is the leader until it sees a larger term.
There are two leaders at the same moment
Wait for a new leader to be elected on the majority side, then write the numbers and terms of the two leaders in /root/raft/two-leaders.json as old_leader, old_term, new_leader, and new_term.
This is the title of the course. If you put the /status of the three nodes side by side, two nodes answer that they are the leader at the same time. But their terms differ. What Raft prevents is not "having two leaders" but "having two leaders in the same term," and that distinction is what the safety of this algorithm stands on.
The value written to the minority side
Write the value ghost to the cut-off old leader, and write accepted, committed, log_len, and commit in /root/raft/minority.json.
The cut-off leader does accept the value. It writes it to its own log and returns ok to the client. But no confirmation from a majority arrives, so commit stays stuck at 1. You can see the distance between what was received and what was finalized right here; if a client that does not know about this distance reads ok as success, an incident follows.
The value written to the majority side
Write the value real to the new leader, and write index, committed, commit, and same_index_as_ghost in /root/raft/majority.json.
The key is that the two values receive the same position number. ghost and real are fighting over position 2 of the log, and only one can remain in a position. Which one will remain is already decided: the one that got confirmation from a majority.
Recovery: what gets discarded
Remove all the cuts on the three nodes and confirm that the old leader's log gets cleaned up, then write survived, discarded, old_leader_state_after, ghost_was_committed, and acknowledged_to_client in /root/raft/verdict.json.
Removing the cut is curl -s -XPOST -d '{"peers":[]}' 127.0.0.1:<포트>/cut, with the node's port in place of the placeholder. The old leader steps down the moment it sees a heartbeat with a larger term, and after that position 2 of its log is overwritten with the new leader's entry. The last two items are the lesson of this lab: that value was never committed, yet the client received ok.