A disaster recovery (DR) plan answers one question: if this server disappears right now, how do you get the service back, how long does it take and how much data is lost? A backup job alone is not a plan until you know what it covers, where the copies live and that a restore actually works. In this tutorial you will define recovery targets, inventory what needs protecting, set up encrypted offsite backups with restic on Ubuntu 24.04, automate them with a systemd timer, test a restore and write a runbook your team can follow under pressure.

Prerequisites

  • One or more servers running Ubuntu 24.04 LTS, for example CubePath VPS or dedicated servers, and a non-root user with sudo privileges.
  • A backup destination in a different location from the server: a second server in another data center reachable over SSH, or an S3-compatible object storage bucket.
  • If you run MySQL/MariaDB or PostgreSQL, credentials for a user that can dump all databases.

Step 1 - Defining RTO and RPO

Two numbers drive every decision in the plan:

  • Recovery Time Objective (RTO): the maximum time a service can be down. An RTO of 4 hours means the service must be back within 4 hours of the incident.
  • Recovery Point Objective (RPO): the maximum amount of data, measured in time, you can afford to lose. An RPO of 1 hour means backups or replication must run at least every hour.

Set them per service, together with whoever owns the business side. A typical table looks like this:

ServiceTierRTORPOProtection
Customer databaseCritical1 h5 minReplica in another data center + hourly dumps
Web applicationHigh4 h24 hCode in Git, config backed up daily, rebuild from automation
Internal wikiMedium24 h24 hDaily backup
Build cacheLow1 weekn/aNo backup, rebuild

Low RPO values cannot be met with nightly backups alone; they need replication (see the guide on cross-datacenter replication). Low RTO values need either a standby server or a fully automated rebuild. Writing the targets down first avoids spending money on protection nobody needs, and exposes the services that are under-protected.

Step 2 - Inventorying what to protect

For each server, list what you would need to rebuild it from scratch. Typically:

  • Data: databases, uploaded files, mail spools, anything users create.
  • Configuration: /etc, systemd units, crontabs, TLS certificates and keys, application .env files.
  • Things that are not on the server: DNS records, firewall rules at the provider, SSH keys, secrets in a vault, license keys.
  • How to rebuild: OS version, installed packages, and ideally an Ansible playbook or similar that reinstalls everything.

The list of installed packages lets you reinstall the same software on a new server. You can see it with dpkg --get-selections; the backup script in Step 5 saves it automatically next to the database dumps.

Also record the risks you are protecting against, because each needs a different copy: hardware or disk failure (another server), data center outage (another location), accidental deletion or ransomware (versioned backups with retention that an attacker on the server cannot delete), and a bad deploy (fast rollback).

The classic rule of thumb is 3-2-1: three copies of the data, on two different types of storage, with one copy offsite.

Step 3 - Dumping databases consistently

Copying the files of a running database produces a corrupt backup. Dump it with the database's own tool into a local directory, and let the file backup in the next step pick the dump up.

Create the directory:

sudo mkdir -p /var/backups/db
sudo chmod 700 /var/backups/db

For MySQL on Ubuntu, root can connect through the local socket without a password. --single-transaction takes a consistent snapshot of InnoDB tables without locking them:

sudo mysqldump --all-databases --single-transaction --routines --triggers --events \
  | gzip | sudo tee /var/backups/db/mysql.sql.gz > /dev/null

For PostgreSQL, dump each database in the custom format, which pg_restore can restore selectively, plus the global objects such as roles:

sudo -u postgres pg_dump -Fc your_database | sudo tee /var/backups/db/your_database.dump > /dev/null
sudo -u postgres pg_dumpall --globals-only | sudo tee /var/backups/db/globals.sql > /dev/null

Check that the files exist and are not empty:

sudo ls -lh /var/backups/db/
-rw-r--r-- 1 root root 48M Sep 25 02:00 mysql.sql.gz

You will run these commands automatically in Step 5.

Step 4 - Setting up encrypted offsite backups with restic

restic creates encrypted, deduplicated and versioned snapshots, and supports SFTP, S3-compatible storage and more as destinations. Install it from the Ubuntu repositories:

sudo apt update
sudo apt install restic

This guide uses an SFTP destination: a user backup on a separate server, backup.example.com, in another data center. First create an SSH key for root on the source server and authorize it on the backup server:

sudo ssh-keygen -t ed25519 -N "" -f /root/.ssh/id_ed25519_backup
sudo ssh-copy-id -i /root/.ssh/id_ed25519_backup.pub [email protected]

Tell SSH to use that key for the backup host:

sudo nano /root/.ssh/config
Host backup.example.com
    User backup
    IdentityFile /root/.ssh/id_ed25519_backup

Store the repository location and a strong encryption password in root-only files. Replace your_strong_password with a long random value, and keep a copy of it in your password manager: without it the backups cannot be restored.

sudo mkdir -p /etc/restic
echo 'your_strong_password' | sudo tee /etc/restic/password > /dev/null
sudo nano /etc/restic/env
RESTIC_REPOSITORY=sftp:backup.example.com:/srv/restic/web01
RESTIC_PASSWORD_FILE=/etc/restic/password
sudo chmod 600 /etc/restic/password /etc/restic/env

Initialize the repository:

sudo bash -c 'set -a; . /etc/restic/env; restic init'
created restic repository 1a2b3c4d5e at sftp:backup.example.com:/srv/restic/web01

Run a first backup of configuration, data and database dumps. Adjust the paths to your inventory from Step 2:

sudo bash -c 'set -a; . /etc/restic/env; restic backup /etc /var/www /var/backups/db'

List the snapshots to confirm it worked:

sudo bash -c 'set -a; . /etc/restic/env; restic snapshots'
ID        Time                 Host   Tags  Paths
---------------------------------------------------------------
4f8e2c1a  2026-09-25 02:00:11  web01        /etc
                                            /var/backups/db
                                            /var/www

Step 5 - Automating backups with a systemd timer

Put the dump and backup commands in one script so they run in order and fail loudly:

sudo nano /usr/local/sbin/dr-backup
#!/usr/bin/env bash
set -euo pipefail

set -a
. /etc/restic/env
set +a

# 1. Database dump (keep only the one you use)
mysqldump --all-databases --single-transaction --routines --triggers --events \
  | gzip > /var/backups/db/mysql.sql.gz

# 2. Package list
dpkg --get-selections > /var/backups/db/packages.list

# 3. Offsite snapshot
restic backup --tag scheduled /etc /var/www /var/backups/db

# 4. Retention: 7 daily, 4 weekly, 6 monthly snapshots
restic forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune

# 5. Verify repository structure and a sample of the data
restic check --read-data-subset=5%

Make it executable and run it once by hand:

sudo chmod 700 /usr/local/sbin/dr-backup
sudo /usr/local/sbin/dr-backup

Create the service unit:

sudo nano /etc/systemd/system/dr-backup.service
[Unit]
Description=Database dump and offsite restic backup
Wants=network-online.target
After=network-online.target

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/dr-backup

And the timer that runs it every night at 02:00:

sudo nano /etc/systemd/system/dr-backup.timer
[Unit]
Description=Run dr-backup nightly

[Timer]
OnCalendar=*-*-* 02:00:00
RandomizedDelaySec=15m
Persistent=true

[Install]
WantedBy=timers.target

Enable the timer and check the next run:

sudo systemctl daemon-reload
sudo systemctl enable --now dr-backup.timer
systemctl list-timers dr-backup.timer
NEXT                        LEFT     LAST PASSED UNIT            ACTIVATES
Fri 2026-09-26 02:07:41 UTC 13h left -    -      dr-backup.timer dr-backup.service

After the first scheduled run, check the result in the journal:

journalctl -u dr-backup.service --since today

If your RPO is shorter than 24 hours, change OnCalendar accordingly, for example hourly. To get alerted when a backup fails, add OnFailure= to the service pointing to a unit that sends a notification, or have your monitoring system check the age of the latest snapshot.

Step 6 - Testing a restore

A backup that has never been restored is only a hope. Test at two levels.

Monthly file-level test: restore the latest snapshot into a temporary directory on the server and compare a known file:

sudo bash -c 'set -a; . /etc/restic/env; restic restore latest --target /tmp/restore-test'
sudo diff -r /etc/ssh /tmp/restore-test/etc/ssh && echo "restore OK"
sudo rm -rf /tmp/restore-test

Quarterly full drill: create a fresh server, then follow the runbook from Step 7 exactly as written, without shortcuts from memory:

  1. Install Ubuntu 24.04 and restic, and copy /etc/restic/env and the password from your password manager.
  2. Restore the snapshot to a temporary path with restic restore latest --target /restore.
  3. Reinstall packages from the saved packages.list (sudo dpkg --set-selections < packages.list && sudo apt-get dselect-upgrade).
  4. Copy back configuration and application files, and import the database dump (gunzip < mysql.sql.gz | sudo mysql, or pg_restore -d your_database your_database.dump).
  5. Start the services, point a test hostname at the new server and check that the application works.

Time every step. If the total exceeds the RTO from Step 1, you need more automation or a standby server. Write down every command that was missing or wrong, and fix the runbook.

Step 7 - Writing the runbook

The runbook is the document someone opens during an outage, possibly at night and possibly not the person who built the system. Keep it short, specific and stored outside the infrastructure it describes (for example in a shared drive plus a printed or offline copy). A useful structure:

# DR runbook: web01 (customer portal)

## Targets
RTO 4 h, RPO 24 h. Owner: ops on-call.

## Where things are
- Backups: restic, sftp:backup.example.com:/srv/restic/web01
- Restic password: password manager, entry "restic web01"
- DNS: provider panel, zone example.com, record portal A
- Automation: git repo infra/ansible, playbook web.yml

## Declare the incident
Who decides, who is notified, where status is communicated.

## Recovery steps
1. Create a new server (Ubuntu 24.04, 4 GB RAM) in another location.
2. ... exact commands ...
9. Update DNS record portal to the new IP, TTL 300.

## Verify
- https://portal.example.com loads, login works, latest order visible.

## Contacts
Provider support, database owner, business owner.

## Last tested
2026-09-15, 2 h 40 min, issues found: step 4 missing TLS key path (fixed).

Keep DNS TTLs for critical records low (300 seconds is common) so the switch in the final step propagates quickly.

Troubleshooting

  • Fatal: unable to open config file or SSH errors on restic init: test the connection first with sudo ssh backup.example.com, and check that the directory in RESTIC_REPOSITORY is writable by the backup user.
  • repository is already locked: a previous run was interrupted. Make sure no backup is running, then run restic unlock.
  • The backup timer never fires: confirm it is enabled with systemctl list-timers, and check journalctl -u dr-backup.service for errors from the last run.
  • mysqldump: Got error: 1045: Access denied: run the script as root, or create a /root/.my.cnf with a dedicated backup user and restrict it with chmod 600.

Conclusion

You now have written recovery targets, an inventory of what matters, automated encrypted offsite backups with retention and verification, a tested restore procedure and a runbook that anyone on the team can follow. Review the plan every time the infrastructure changes and repeat the full drill at least twice a year. As next steps, add database replication to another data center for services with a low RPO, and automate server provisioning with Ansible to shorten your RTO.