TT Lab
Get started
Learn Learning paths Courses

PostgreSQL Replication and Promotion

Stand Up Replication and Promote It

Continue in TT Lab

Goal

Build a primary and a standby inside a single Pod, attach replication, and even promote the standby. At the end you will see the timeline fork with your own eyes.

Why it matters

"We built HA" usually means we have never once done a failover. The configuration is right, but when you actually promote, the standby has not caught up, or the promotion works but the application is still looking at the old address, or you try to reattach the old primary and ruin the data.

This lab walks through that whole loop ahead of time.

Environment

This Pod runs as the postgres account (lab Pods have no capabilities, so you cannot switch users with su). So you can type commands directly, without su or sudo.

psql                      주 서버(5432)에 바로 붙습니다
psql -h 127.0.0.1 -p 5433 대기 서버

Be sure to add -h 127.0.0.1 for the standby. This environment moves the Unix socket directory under the home directory (the container runtime overlays /run with a tmpfs, so the default path disappears).

Steps

  1. Make the primary replication-capable
  2. Replication role and slot
  3. Base backup → /var/lib/postgresql/standby
  4. Start the standby on 5433
  5. Check replication → /root/db/05-repl.txt
  6. Check read-only → /root/db/06-readonly.txt
  7. Promote
  8. Compare timelines → /root/db/08-timeline.txt

Notes

Make the primary replication-capable

Make the primary replication-capable

Change wal_level, max_wal_senders, max_replication_slots, and hot_standby with alter system set, then restart. Restart with pg_ctl -D /var/lib/postgresql/data -w -t 40 restart -l /var/log/postgresql.log. show wal_level must return replica.

Replication role and slot

Replication role and slot

Create the role with create role rep with replication login password 'rep', and create the slot with select pg_create_physical_replication_slot('s1'). Without a slot, if the standby disconnects briefly and comes back, the WAL it needs is already gone and replication breaks.

Base backup

Base backup → /var/lib/postgresql/standby

PGPASSWORD=rep pg_basebackup -h 127.0.0.1 -U rep -D /var/lib/postgresql/standby -X stream -S s1 -R. -R automatically creates standby.signal and primary_conninfo — that one file is the marker that says "you are a standby."

Standby on 5433

Start the standby on 5433

After echo "port=5433" >> /var/lib/postgresql/standby/postgresql.auto.conf, run pg_ctl -D /var/lib/postgresql/standby -l /var/log/labhub/standby.log -w -t 40 start. To check, psql -h 127.0.0.1 -p 5433 -tAc "select pg_is_in_recovery()" must return t.\n\nBe sure to add -h 127.0.0.1. This environment has moved the Unix socket directory, so if you change only the port, the socket path cannot be found.

Is it really replicating

Check replication → /root/db/05-repl.txt

On the primary, run create table if not exists ha(id int) and insert into ha values (42), then after 2 seconds check that the same row is visible on the standby. Then save the result of select state, replay_lag from pg_stat_replication to /root/db/05-repl.txt. The state must be streaming.

Why the standby refuses writes

Check read-only → /root/db/06-readonly.txt

Try insert into ha values (1) on the standby and save the error message as is to /root/db/06-readonly.txt. The standby cannot produce WAL on its own, so it cannot accept writes — it is structure, not configuration.

Promote

Promote

Promote with pg_ctl -D /var/lib/postgresql/standby promote. pg_is_in_recovery() changes to f and writes become possible. After promotion, try an insert once.

The timeline forks

Compare timelines → /root/db/08-timeline.txt

When you promote, the new server forks onto timeline 2. Compare the WAL file names of the two servers and write them on two lines in /root/db/08-timeline.txt:\npsql -h 127.0.0.1 -p 5432 -tAc "select pg_walfile_name(pg_current_wal_lsn())" and the one for 5433. The first 8 digits of the file name are the timeline (00000001 / 00000002).\n\nThis is why you cannot simply reattach the old primary — the two servers started writing different histories from the same point. To reattach it, you have to rewind it with pg_rewind to the point where they forked.