PostgreSQL Replication and Promotion
Stand Up Replication and Promote It
Goal
Build a primary and a standby inside a single Pod, attach replication, and even promote the standby. At the end you will see the timeline fork with your own eyes.
Why it matters
"We built HA" usually means we have never once done a failover. The configuration is right, but when you actually promote, the standby has not caught up, or the promotion works but the application is still looking at the old address, or you try to reattach the old primary and ruin the data.
This lab walks through that whole loop ahead of time.
Environment
This Pod runs as the postgres account (lab Pods have no capabilities, so you
cannot switch users with su). So you can type commands directly, without su or sudo.
psql 주 서버(5432)에 바로 붙습니다
psql -h 127.0.0.1 -p 5433 대기 서버
Be sure to add -h 127.0.0.1 for the standby. This environment moves the Unix socket
directory under the home directory (the container runtime overlays /run with a tmpfs,
so the default path disappears).
Steps
- Make the primary replication-capable
- Replication role and slot
- Base backup →
/var/lib/postgresql/standby - Start the standby on 5433
- Check replication →
/root/db/05-repl.txt - Check read-only →
/root/db/06-readonly.txt - Promote
- Compare timelines →
/root/db/08-timeline.txt
Notes
- If an error occurs, check
/var/log/postgresql.log(primary) and/var/log/labhub/standby.log(standby). pg_stat_replicationis visible only on the primary. On the standby it is normal for it to be empty.
Make the primary replication-capable
Make the primary replication-capable
Change wal_level, max_wal_senders, max_replication_slots, and hot_standby with alter system set, then restart. Restart with pg_ctl -D /var/lib/postgresql/data -w -t 40 restart -l /var/log/postgresql.log. show wal_level must return replica.
Replication role and slot
Replication role and slot
Create the role with create role rep with replication login password 'rep', and create the slot with select pg_create_physical_replication_slot('s1'). Without a slot, if the standby disconnects briefly and comes back, the WAL it needs is already gone and replication breaks.
Base backup
Base backup → /var/lib/postgresql/standby
PGPASSWORD=rep pg_basebackup -h 127.0.0.1 -U rep -D /var/lib/postgresql/standby -X stream -S s1 -R. -R automatically creates standby.signal and primary_conninfo — that one file is the marker that says "you are a standby."
Standby on 5433
Start the standby on 5433
After echo "port=5433" >> /var/lib/postgresql/standby/postgresql.auto.conf, run pg_ctl -D /var/lib/postgresql/standby -l /var/log/labhub/standby.log -w -t 40 start. To check, psql -h 127.0.0.1 -p 5433 -tAc "select pg_is_in_recovery()" must return t.\n\nBe sure to add -h 127.0.0.1. This environment has moved the Unix socket directory, so if you change only the port, the socket path cannot be found.
Is it really replicating
Check replication → /root/db/05-repl.txt
On the primary, run create table if not exists ha(id int) and insert into ha values (42), then after 2 seconds check that the same row is visible on the standby. Then save the result of select state, replay_lag from pg_stat_replication to /root/db/05-repl.txt. The state must be streaming.
Why the standby refuses writes
Check read-only → /root/db/06-readonly.txt
Try insert into ha values (1) on the standby and save the error message as is to /root/db/06-readonly.txt. The standby cannot produce WAL on its own, so it cannot accept writes — it is structure, not configuration.
Promote
Promote
Promote with pg_ctl -D /var/lib/postgresql/standby promote. pg_is_in_recovery() changes to f and writes become possible. After promotion, try an insert once.
The timeline forks
Compare timelines → /root/db/08-timeline.txt
When you promote, the new server forks onto timeline 2. Compare the WAL file names of the two servers and write them on two lines in /root/db/08-timeline.txt:\npsql -h 127.0.0.1 -p 5432 -tAc "select pg_walfile_name(pg_current_wal_lsn())" and the one for 5433. The first 8 digits of the file name are the timeline (00000001 / 00000002).\n\nThis is why you cannot simply reattach the old primary — the two servers started writing different histories from the same point. To reattach it, you have to rewind it with pg_rewind to the point where they forked.