A disaster recovery (DR) plan answers one question: if this server disappears right now, how do you get the service back, how long does it take and how much data is lost? A backup job alone is not a plan until you know what it covers, where the copies live and that a restore actually works. In this tutorial you will define recovery targets, inventory what needs protecting, set up encrypted offsite backups with restic on Ubuntu 24.04, automate them with a systemd timer, test a restore and write a runbook your team can follow under pressure.
Prerequisites
- One or more servers running Ubuntu 24.04 LTS, for example CubePath VPS or dedicated servers, and a non-root user with
sudoprivileges. - A backup destination in a different location from the server: a second server in another data center reachable over SSH, or an S3-compatible object storage bucket.
- If you run MySQL/MariaDB or PostgreSQL, credentials for a user that can dump all databases.
Step 1 - Defining RTO and RPO
Two numbers drive every decision in the plan:
- Recovery Time Objective (RTO): the maximum time a service can be down. An RTO of 4 hours means the service must be back within 4 hours of the incident.
- Recovery Point Objective (RPO): the maximum amount of data, measured in time, you can afford to lose. An RPO of 1 hour means backups or replication must run at least every hour.
Set them per service, together with whoever owns the business side. A typical table looks like this:
| Service | Tier | RTO | RPO | Protection |
|---|---|---|---|---|
| Customer database | Critical | 1 h | 5 min | Replica in another data center + hourly dumps |
| Web application | High | 4 h | 24 h | Code in Git, config backed up daily, rebuild from automation |
| Internal wiki | Medium | 24 h | 24 h | Daily backup |
| Build cache | Low | 1 week | n/a | No backup, rebuild |
Low RPO values cannot be met with nightly backups alone; they need replication (see the guide on cross-datacenter replication). Low RTO values need either a standby server or a fully automated rebuild. Writing the targets down first avoids spending money on protection nobody needs, and exposes the services that are under-protected.
Step 2 - Inventorying what to protect
For each server, list what you would need to rebuild it from scratch. Typically:
- Data: databases, uploaded files, mail spools, anything users create.
- Configuration:
/etc, systemd units, crontabs, TLS certificates and keys, application.envfiles. - Things that are not on the server: DNS records, firewall rules at the provider, SSH keys, secrets in a vault, license keys.
- How to rebuild: OS version, installed packages, and ideally an Ansible playbook or similar that reinstalls everything.
The list of installed packages lets you reinstall the same software on a new server. You can see it with dpkg --get-selections; the backup script in Step 5 saves it automatically next to the database dumps.
Also record the risks you are protecting against, because each needs a different copy: hardware or disk failure (another server), data center outage (another location), accidental deletion or ransomware (versioned backups with retention that an attacker on the server cannot delete), and a bad deploy (fast rollback).
The classic rule of thumb is 3-2-1: three copies of the data, on two different types of storage, with one copy offsite.
Step 3 - Dumping databases consistently
Copying the files of a running database produces a corrupt backup. Dump it with the database's own tool into a local directory, and let the file backup in the next step pick the dump up.
Create the directory:
sudo mkdir -p /var/backups/db
sudo chmod 700 /var/backups/db
For MySQL on Ubuntu, root can connect through the local socket without a password. --single-transaction takes a consistent snapshot of InnoDB tables without locking them:
sudo mysqldump --all-databases --single-transaction --routines --triggers --events \
| gzip | sudo tee /var/backups/db/mysql.sql.gz > /dev/null
For PostgreSQL, dump each database in the custom format, which pg_restore can restore selectively, plus the global objects such as roles:
sudo -u postgres pg_dump -Fc your_database | sudo tee /var/backups/db/your_database.dump > /dev/null
sudo -u postgres pg_dumpall --globals-only | sudo tee /var/backups/db/globals.sql > /dev/null
Check that the files exist and are not empty:
sudo ls -lh /var/backups/db/
-rw-r--r-- 1 root root 48M Sep 25 02:00 mysql.sql.gz
You will run these commands automatically in Step 5.
Step 4 - Setting up encrypted offsite backups with restic
restic creates encrypted, deduplicated and versioned snapshots, and supports SFTP, S3-compatible storage and more as destinations. Install it from the Ubuntu repositories:
sudo apt update
sudo apt install restic
This guide uses an SFTP destination: a user backup on a separate server, backup.example.com, in another data center. First create an SSH key for root on the source server and authorize it on the backup server:
sudo ssh-keygen -t ed25519 -N "" -f /root/.ssh/id_ed25519_backup
sudo ssh-copy-id -i /root/.ssh/id_ed25519_backup.pub [email protected]
Tell SSH to use that key for the backup host:
sudo nano /root/.ssh/config
Host backup.example.com
User backup
IdentityFile /root/.ssh/id_ed25519_backup
Store the repository location and a strong encryption password in root-only files. Replace your_strong_password with a long random value, and keep a copy of it in your password manager: without it the backups cannot be restored.
sudo mkdir -p /etc/restic
echo 'your_strong_password' | sudo tee /etc/restic/password > /dev/null
sudo nano /etc/restic/env
RESTIC_REPOSITORY=sftp:backup.example.com:/srv/restic/web01
RESTIC_PASSWORD_FILE=/etc/restic/password
sudo chmod 600 /etc/restic/password /etc/restic/env
Initialize the repository:
sudo bash -c 'set -a; . /etc/restic/env; restic init'
created restic repository 1a2b3c4d5e at sftp:backup.example.com:/srv/restic/web01
Run a first backup of configuration, data and database dumps. Adjust the paths to your inventory from Step 2:
sudo bash -c 'set -a; . /etc/restic/env; restic backup /etc /var/www /var/backups/db'
List the snapshots to confirm it worked:
sudo bash -c 'set -a; . /etc/restic/env; restic snapshots'
ID Time Host Tags Paths
---------------------------------------------------------------
4f8e2c1a 2026-09-25 02:00:11 web01 /etc
/var/backups/db
/var/www
ImportantIf an attacker gets root on the server, they can use the same credentials to delete the backups. For protection against ransomware, use a destination that enforces retention on its side, such as an S3 bucket with object lock or a backup server that takes its own snapshots of
/srv/restic.
Step 5 - Automating backups with a systemd timer
Put the dump and backup commands in one script so they run in order and fail loudly:
sudo nano /usr/local/sbin/dr-backup
#!/usr/bin/env bash
set -euo pipefail
set -a
. /etc/restic/env
set +a
# 1. Database dump (keep only the one you use)
mysqldump --all-databases --single-transaction --routines --triggers --events \
| gzip > /var/backups/db/mysql.sql.gz
# 2. Package list
dpkg --get-selections > /var/backups/db/packages.list
# 3. Offsite snapshot
restic backup --tag scheduled /etc /var/www /var/backups/db
# 4. Retention: 7 daily, 4 weekly, 6 monthly snapshots
restic forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune
# 5. Verify repository structure and a sample of the data
restic check --read-data-subset=5%
Make it executable and run it once by hand:
sudo chmod 700 /usr/local/sbin/dr-backup
sudo /usr/local/sbin/dr-backup
Create the service unit:
sudo nano /etc/systemd/system/dr-backup.service
[Unit]
Description=Database dump and offsite restic backup
Wants=network-online.target
After=network-online.target
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/dr-backup
And the timer that runs it every night at 02:00:
sudo nano /etc/systemd/system/dr-backup.timer
[Unit]
Description=Run dr-backup nightly
[Timer]
OnCalendar=*-*-* 02:00:00
RandomizedDelaySec=15m
Persistent=true
[Install]
WantedBy=timers.target
Enable the timer and check the next run:
sudo systemctl daemon-reload
sudo systemctl enable --now dr-backup.timer
systemctl list-timers dr-backup.timer
NEXT LEFT LAST PASSED UNIT ACTIVATES
Fri 2026-09-26 02:07:41 UTC 13h left - - dr-backup.timer dr-backup.service
After the first scheduled run, check the result in the journal:
journalctl -u dr-backup.service --since today
If your RPO is shorter than 24 hours, change OnCalendar accordingly, for example hourly. To get alerted when a backup fails, add OnFailure= to the service pointing to a unit that sends a notification, or have your monitoring system check the age of the latest snapshot.
Step 6 - Testing a restore
A backup that has never been restored is only a hope. Test at two levels.
Monthly file-level test: restore the latest snapshot into a temporary directory on the server and compare a known file:
sudo bash -c 'set -a; . /etc/restic/env; restic restore latest --target /tmp/restore-test'
sudo diff -r /etc/ssh /tmp/restore-test/etc/ssh && echo "restore OK"
sudo rm -rf /tmp/restore-test
Quarterly full drill: create a fresh server, then follow the runbook from Step 7 exactly as written, without shortcuts from memory:
- Install Ubuntu 24.04 and
restic, and copy/etc/restic/envand the password from your password manager. - Restore the snapshot to a temporary path with
restic restore latest --target /restore. - Reinstall packages from the saved
packages.list(sudo dpkg --set-selections < packages.list && sudo apt-get dselect-upgrade). - Copy back configuration and application files, and import the database dump (
gunzip < mysql.sql.gz | sudo mysql, orpg_restore -d your_database your_database.dump). - Start the services, point a test hostname at the new server and check that the application works.
Time every step. If the total exceeds the RTO from Step 1, you need more automation or a standby server. Write down every command that was missing or wrong, and fix the runbook.
Step 7 - Writing the runbook
The runbook is the document someone opens during an outage, possibly at night and possibly not the person who built the system. Keep it short, specific and stored outside the infrastructure it describes (for example in a shared drive plus a printed or offline copy). A useful structure:
# DR runbook: web01 (customer portal)
## Targets
RTO 4 h, RPO 24 h. Owner: ops on-call.
## Where things are
- Backups: restic, sftp:backup.example.com:/srv/restic/web01
- Restic password: password manager, entry "restic web01"
- DNS: provider panel, zone example.com, record portal A
- Automation: git repo infra/ansible, playbook web.yml
## Declare the incident
Who decides, who is notified, where status is communicated.
## Recovery steps
1. Create a new server (Ubuntu 24.04, 4 GB RAM) in another location.
2. ... exact commands ...
9. Update DNS record portal to the new IP, TTL 300.
## Verify
- https://portal.example.com loads, login works, latest order visible.
## Contacts
Provider support, database owner, business owner.
## Last tested
2026-09-15, 2 h 40 min, issues found: step 4 missing TLS key path (fixed).
Keep DNS TTLs for critical records low (300 seconds is common) so the switch in the final step propagates quickly.
Troubleshooting
Fatal: unable to open config fileor SSH errors onrestic init: test the connection first withsudo ssh backup.example.com, and check that the directory inRESTIC_REPOSITORYis writable by thebackupuser.repository is already locked: a previous run was interrupted. Make sure no backup is running, then runrestic unlock.- The backup timer never fires: confirm it is enabled with
systemctl list-timers, and checkjournalctl -u dr-backup.servicefor errors from the last run. mysqldump: Got error: 1045: Access denied: run the script as root, or create a/root/.my.cnfwith a dedicated backup user and restrict it withchmod 600.
Conclusion
You now have written recovery targets, an inventory of what matters, automated encrypted offsite backups with retention and verification, a tested restore procedure and a runbook that anyone on the team can follow. Review the plan every time the infrastructure changes and repeat the full drill at least twice a year. As next steps, add database replication to another data center for services with a low RPO, and automate server provisioning with Ansible to shorten your RTO.
