Write the leader election rules yourself
Goal
You implement, yourself, the process by which three processes exchange votes and elect one leader. When you finish, you will have experienced by hand the term, the vote, the majority, and why the timeout is randomized.
Why it matters
What makes consensus algorithms hard is not the math but the boundaries. When the terms are equal versus greater, when a vote has not been cast yet versus already cast, when the votes are exactly half versus more than half: if you set any of these boundaries off by one, the system runs fine day to day and then creates two leaders on the day the network wobbles. So here the wiring (HTTP server, timers, retransmission) is provided in advance and you write only the functions that make decisions. The decisions are the algorithm, and the rest is plumbing.
The last step shows by measurement why the Raft paper introduced the randomized timeout. If you set the timers of the three nodes to the same value, all three become candidates at once and each gives itself one vote, and nobody reaches a majority, so the term just keeps rising forever.
Steps
- In
/root/raft/cluster.json, write the ids and ports of the three nodes. - Copy
/opt/lab/raft/node.pyto/root/raft/node.py, start one, and check/status. - In
/root/raft/raftrules.py, writeelection_timeout_ms(node_id, fixed_ms). - Add
on_timeout(state)to the same file. - Add
on_request_vote(state, req)to the same file. - Add
on_append_entries(state, req)to the same file. - Add
has_majority(votes, total)to the same file, then start the three nodes and watch a leader get elected. - Start the three nodes with
--timeout-ms 900and write the result to/root/raft/split-vote.json.
Notes
- How to start a node:
cd /root/raft && python3 node.py --id 1 --port 5001 --peers 2:5002,3:5003 & - View the status:
curl -s http://127.0.0.1:5001/status - Stop everything:
pkill -f 'node.py --id' - The wiring calls only the functions that exist in the rules file. A function you have not written yet will not crash the process.
- Common mistake 1: receiving a request with a stale term and lowering your own term to match it. A term only goes up.
- Common mistake 2: writing the majority as
votes >= total / 2. With 4 nodes, 2 votes would count as a majority and you would get two leaders.
A seating chart for three nodes
In /root/raft/cluster.json, write ids 1, 2, and 3 with ports 5001, 5002, and 5003.
The nodes are not containers but processes that differ only in port. First write down the seating chart you will use later when you start all three. Under a list named nodes, put three objects that each have an id and a port.
Put the wiring in place and start it
Copy /opt/lab/raft/node.py to /root/raft/node.py, start one, and check that /status responds.
The wiring, such as the HTTP server and the timers, is already built. All you fill in is the "decisions." Start it with python3 /root/raft/node.py --id 1 --port 5001 --peers '' and call curl 127.0.0.1:5001/status from another window. The rules file does not exist yet, so it is normal for the node to make no decisions and just sit still.
Election timeout
In /root/raft/raftrules.py, write election_timeout_ms(node_id, fixed_ms). If fixed_ms is not 0, return that value as is; if it is 0, return a random value from 800 to 1500.
random.randint(a, b) includes both endpoints. You will see for yourself in the last step why it has to be random. Be sure to leave the path that returns fixed_ms unchanged; without it, the experiment in the last step cannot work.
On timeout, become a candidate
In raftrules.py, add on_timeout(state). It returns a dictionary that raises the term by 1, sets state to candidate, and sets voted_for to its own id.
state is a dictionary with {'id', 'term', 'state', 'voted_for', 'leader', 'log', 'commit'}. Return only the keys you want to change and the wiring will apply them. A candidate voting for itself is Raft's first vote; if you leave this out, nobody reaches a majority.
One vote per term
In raftrules.py, add on_request_vote(state, req). It returns a dictionary containing term, voted_for, state, and granted.
Check three things in order. (1) If the req's term is greater than your term, raise your term, step down, and clear the vote record for this term. (2) If the term is smaller than yours, reject; you must not lower your term in that case. (3) In the same term, grant the vote only if you have not yet given it to anyone or you gave it to the same candidate. If you reject a retransmission from the same candidate, losing a single packet stops the election.
The receiving side of the heartbeat
In raftrules.py, add on_append_entries(state, req). It returns a dictionary containing term, state, leader, voted_for, and ok.
The leader keeps talking even when it has nothing to say, because that is the only signal meaning "I am still alive." If the req's term is smaller than yours, return ok as False and leave your term unchanged. Otherwise, go back to follower, record the leader, and set ok to True. If you are a candidate and see a leader in the same term, you must step down.
Count the majority and become leader
In raftrules.py, add has_majority(votes, total), then start three nodes on 5001, 5002, and 5003 and confirm that one leader is elected.
A majority is "more than" half, not "at least" half. With 4 nodes, 2 votes are not a majority. Integer division makes it easy to get this boundary wrong, so multiply both sides by 2 before comparing. After starting the three nodes, run curl 127.0.0.1:5001/status against all three ports and look at state, term, and leader together.
If the timers are the same, nobody wins
Start all three nodes with --timeout-ms 900, look at the state after 9 seconds, and write fixed_timeout_ms, leader, max_term, and states in /root/raft/split-vote.json.
When you pass --timeout-ms, the wiring aligns the wake-up times of the three nodes to the wall clock. This is a device for seeing what happens when the timers are truly the same. If all three become candidates at the same moment, each gives one vote to itself, nobody reaches a majority, and only the term rises. Write leader as null.