A multi-region architecture keeps your service available when an entire data center or region goes offline, not just a single server. It is also the most expensive and complex form of high availability, so the design should start from how much downtime and data loss you can actually accept. This guide explains the decisions involved, compares the common patterns, and walks through a concrete active-passive setup on Ubuntu 24.04: PostgreSQL replicated to a second region, DNS-based failover, a promotion procedure that avoids split-brain, and a failback.
Prerequisites
The concepts apply to any stack. For the hands-on examples you need:
- Two Ubuntu 24.04 LTS servers in different regions, for example two CubePath servers in different locations, each with a non-root
sudouser. The examples call themdb-a(primary, region A) anddb-b(standby, region B). - PostgreSQL 16 from the Ubuntu repositories on both (
sudo apt install postgresql). - An encrypted path between regions. PostgreSQL replication below uses TLS, but a private tunnel such as WireGuard is recommended for everything else that crosses regions.
- A DNS provider that supports low TTLs and, ideally, health-checked failover records.
Start with RTO and RPO
Two numbers drive every other decision:
- RTO (recovery time objective): how long the service may be down after a regional failure.
- RPO (recovery point objective): how much recent data you may lose, measured in time.
| Target | What it implies |
|---|---|
| RTO of hours, RPO of 24 hours | Nightly backups copied to another region. Rebuild on demand. No multi-region infrastructure needed. |
| RTO under 1 hour, RPO of minutes | Warm standby: asynchronous replication to a second region, manual or scripted promotion. |
| RTO of minutes, RPO of seconds | Hot standby with automated health checks and DNS failover. Asynchronous replication with monitored lag. |
| RTO near zero, RPO zero | Active-active with synchronous or consensus-based replication. Every write pays the inter-region round trip. |
Be honest here. Most applications land in the second or third row, and that is what the rest of this guide builds.
Active-passive versus active-active
| Active-passive | Active-active | |
|---|---|---|
| Traffic | One region serves, the other waits | Both regions serve users |
| Writes | Single primary database | Writes in both regions, or routed to one |
| Replication | Asynchronous is enough | Synchronous, conflict resolution or data partitioning |
| Failover | Promote the standby, move DNS | Stop sending traffic to the failed region |
| Complexity | Moderate | High: conflicts, latency on every write |
| Cost | Standby can be smaller | Full capacity in both regions |
Active-active looks attractive, but multi-primary databases across regions bring write conflicts and latency that most applications are not built for. A common middle ground is active-active for the stateless tiers (web and API servers read from a local replica) with a single write primary.
Make the application tier stateless
Failover only works if any application server in any region can serve any request. Before replicating anything, move state out of the application servers:
- Sessions: store them in the database or a shared cache, not on local disk.
- Uploaded files: use S3-compatible object storage with replication to the second region, or at least sync them to the standby region on a schedule.
- Configuration and deploys: deploy both regions from the same pipeline and the same artifact, so the standby is never running old code.
- Scheduled jobs: run cron jobs in the active region only. A job that runs in both regions at the same time will duplicate emails, invoices or cleanups.
Choose synchronous or asynchronous replication
The speed of light sets the cost of synchronous replication. Between Europe and the US East Coast the round trip is roughly 80 to 100 ms, so each synchronous commit waits at least that long.
| Asynchronous | Synchronous | |
|---|---|---|
| Write latency | Local only | Local plus inter-region round trip |
| Data loss on failover | Whatever had not replicated yet (usually under a second) | None |
| If the link to the standby fails | Primary keeps working | Writes block until the standby returns or is removed |
| Typical use across regions | Default choice | Only for small, critical write volumes |
For a cross-region standby, asynchronous replication with lag monitoring is the usual answer. The next steps set it up with PostgreSQL.
Step 1 - Preparing the primary in region A
On db-a, create a role for replication. Replace your_strong_password with a long random password:
sudo -u postgres psql -c "CREATE ROLE replicator WITH REPLICATION LOGIN PASSWORD 'your_strong_password';"
Edit the main configuration file:
sudo nano /etc/postgresql/16/main/postgresql.conf
Set these values:
listen_addresses = '*'
wal_level = replica
max_wal_senders = 10
max_replication_slots = 10
max_slot_wal_keep_size = 10GB
wal_log_hints = on
max_slot_wal_keep_size stops a disconnected standby from filling the primary's disk with retained WAL. wal_log_hints allows pg_rewind to turn this server back into a standby after a failover (Step 5).
Allow the standby to connect for replication only, over TLS, in pg_hba.conf:
sudo nano /etc/postgresql/16/main/pg_hba.conf
Add at the end, replacing db_b_ip:
hostssl replication replicator db_b_ip/32 scram-sha-256
Ubuntu enables TLS in PostgreSQL by default with a self-signed certificate. Restart and allow the standby through the firewall:
sudo systemctl restart postgresql
sudo ufw allow from db_b_ip to any port 5432 proto tcp
Step 2 - Cloning the standby in region B
On db-b, stop PostgreSQL and move the empty default cluster aside:
sudo systemctl stop postgresql
sudo mv /var/lib/postgresql/16/main /var/lib/postgresql/16/main.orig
Copy the primary with pg_basebackup. -R writes the standby configuration, and -C -S creates a replication slot on the primary so it keeps the WAL this standby still needs:
sudo -u postgres pg_basebackup -h db_a_ip -U replicator -D /var/lib/postgresql/16/main -R -X stream -C -S standby_region_b -P
Enter the replicator password when asked. Then apply the same postgresql.conf and pg_hba.conf changes you made on db-a (with db_a_ip in pg_hba.conf), so this server is ready to act as a primary later, and start it:
sudo systemctl start postgresql
Confirm on db-b that it runs as a standby:
sudo -u postgres psql -c "SELECT pg_is_in_recovery();"
pg_is_in_recovery
-------------------
t
(1 row)
Step 3 - Monitoring replication lag
Your real RPO is the replication lag at the moment of failure, so measure it continuously. On db-a:
sudo -u postgres psql -c "SELECT client_addr, state, sync_state, replay_lag FROM pg_stat_replication;"
client_addr | state | sync_state | replay_lag
---------------+-----------+------------+-----------------
198.51.100.20 | streaming | async | 00:00:00.084312
(1 row)
Alert when state is not streaming or when replay_lag exceeds your RPO. Also watch the slot, because an inactive slot means the standby is not receiving data:
sudo -u postgres psql -c "SELECT slot_name, active, wal_status FROM pg_replication_slots;"
Feed both queries into whatever monitoring you use (the PostgreSQL exporter for Prometheus exposes the same data).
Step 4 - Routing traffic with DNS failover
DNS is the simplest way to move users between regions:
- Publish the service under one name, for example
app.your_domain, pointing to region A. - Keep the TTL low, 60 seconds is a common value, so a change reaches most clients within a minute or two. Some clients and resolvers cache longer than the TTL, so plan for a tail of stragglers.
- If your DNS provider supports health-checked failover records, configure a check against an HTTP endpoint in region A that tests the application and its database, not only that the web server answers. The provider switches the record to region B when the check fails for a sustained period.
- Put a load balancer (HAProxy or Nginx) in front of the application servers inside each region, so a single server failure is handled locally and never triggers a regional failover.
Check what resolvers return during a test:
dig +noall +answer app.your_domain @1.1.1.1
app.your_domain. 60 IN A 203.0.113.10
Anycast or a global load balancer gives faster failover than DNS, but DNS with health checks is enough for an RTO measured in minutes.
Step 5 - Failing over without split-brain
Split-brain happens when both regions believe they are the primary and accept writes. The data then diverges and cannot be merged automatically. The main defense is simple: fence the old primary before promoting the new one.
The trap is that from region B, "region A is unreachable" and "the link between A and B is broken" look identical. Automatic promotion based only on B's view of A causes split-brain the first time the inter-region link fails. Either keep promotion manual, or use a tool that decides by quorum from at least three locations (for PostgreSQL, Patroni with an etcd cluster spread over three sites).
A manual failover runbook looks like this:
- Confirm the outage from a third location (your laptop, a monitoring probe in another network), not only from region B.
- Fence region A: stop PostgreSQL there if you can reach it (
sudo systemctl stop postgresql), or block the application from reaching it. If you cannot reach it at all, make sure it cannot come back as a primary on its own, for example by powering the server off through its provider. - Promote the standby on
db-b:
sudo -u postgres psql -c "SELECT pg_promote();"
pg_promote
------------
t
(1 row)
- Verify that
db-bnow accepts writes:
sudo -u postgres psql -c "SELECT pg_is_in_recovery();"
pg_is_in_recovery
-------------------
f
(1 row)
- Point the application in region B at
db-band start the jobs that should only run in the active region. - Move DNS to region B if the health check has not already done it.
Step 6 - Failing back
After region A returns, do not start the old primary as a primary. It may contain a few transactions that never reached region B. Turn it into a standby of the new primary with pg_rewind, which uses the wal_log_hints setting from Step 1.
On db-b (the new primary), create a slot for the returning server:
sudo -u postgres psql -c "SELECT pg_create_physical_replication_slot('standby_region_a');"
On db-a, with PostgreSQL stopped, rewind the data directory against db-b. pg_rewind connects as a superuser, so add a temporary hostssl all postgres db_a_ip/32 scram-sha-256 rule to pg_hba.conf on db-b and give the postgres role a password, or use a dedicated role with the privileges described in the PostgreSQL documentation:
sudo -u postgres /usr/lib/postgresql/16/bin/pg_rewind --target-pgdata=/var/lib/postgresql/16/main --source-server="host=db_b_ip user=postgres dbname=postgres sslmode=require" --write-recovery-conf --progress
--write-recovery-conf creates standby.signal and the connection settings. Edit /var/lib/postgresql/16/main/postgresql.auto.conf on db-a so primary_conninfo uses the replicator role and add primary_slot_name = 'standby_region_a', then start it and confirm with pg_is_in_recovery() that it returns t. Remove the temporary superuser rule from db-b afterwards.
If you want region A to be primary again, repeat the failover runbook in the other direction during a maintenance window, when you can do it in a controlled way with zero replication lag.
Step 7 - Testing with regular drills
An untested failover plan will fail in ways nobody predicted. Run a drill at least twice a year:
- Announce a window and follow the runbook exactly as written, including DNS.
- Measure the real RTO (time from the start of the drill until the service answers from region B) and the data loss (compare the last transaction on A with what reached B).
- Record every step that did not match the runbook and fix the runbook, not just the servers.
- Check the things that are easy to forget: TLS certificates valid in region B, cron jobs, outgoing email, firewall rules for third-party APIs that allowlist your IPs.
Which architecture should you choose?
- Backups in another region if an RTO of several hours is acceptable. It is cheap, simple, and protects against the worst case.
- Active-passive with asynchronous replication (this guide) for most production applications. It gives an RTO of minutes and an RPO of seconds with manageable complexity.
- Active-active only when the business needs zero downtime and the application is designed for it, with data partitioned by region or a database built for multi-region consensus.
Conclusion
A multi-region design starts from RTO and RPO, keeps the application tier stateless, replicates data asynchronously to a standby region, and moves traffic with health-checked DNS. The hard part is not the replication but the failover decision: fence before you promote, keep automatic promotion behind a real quorum, and practice the runbook. Next steps: add alerting on pg_stat_replication lag, replicate object storage and backups to region B, and schedule your first failover drill.
