A backup job that finishes without errors does not prove that you can recover from it. Archives get truncated when a disk fills up, dumps silently skip a database, retention deletes the only good copy, and nobody notices until the day a restore is needed. In this tutorial you will check that your backups are recent and intact, restore files and databases into a safe location, automate a weekly restore test with a systemd timer, and run a full recovery drill on a fresh server running Ubuntu 24.04.
Prerequisites
To follow this tutorial, you will need:
- A server running Ubuntu 24.04 LTS with a non-root user that has
sudoprivileges. - Existing backups to test. The examples assume the layout used in the other CubePath backup guides:
- compressed tar archives with
.sha256files in/var/backups/tar, - dated database dump directories with a
SHA256SUMSfile in/var/backups/db, containingmysql-<name>.sql.gz(mysqldump) andpg-<name>.dump(pg_dump custom format) files.
- compressed tar archives with
- For the recovery drill in Step 5: a second, empty server, for example a small CubePath VPS that you delete afterwards.
If your backups live elsewhere or use other names, adjust the paths in the commands.
Step 1 - Deciding what a successful restore means
Before testing, write down what you are testing against. Two numbers matter:
- Recovery point objective (RPO): how much data you can afford to lose. With a nightly backup, the RPO is up to 24 hours.
- Recovery time objective (RTO): how long the service may be down while you restore. Only a real drill tells you whether you meet it.
Then list what has to work after a restore. For a typical web server, that means:
| Check | How to verify |
|---|---|
| Latest backup is recent enough | File or directory age is below the RPO |
| Files are intact | Checksums match, archives can be read to the end |
| Important files are present | A restored sample matches the live copy |
| Databases load | Dumps restore into an empty database without errors |
| Application works | The site answers from the restored server |
The rest of this tutorial turns each row into a command.
Step 2 - Checking freshness and integrity
Start with the cheapest checks: is there a recent backup, and is it undamaged? Find the newest tar archive and its age:
sudo ls -lt /var/backups/tar | head -n 4
total 124M
-rw------- 1 root root 98 Sep 25 03:07 server-20260925-030712.tar.gz.sha256
-rw------- 1 root root 41M Sep 25 03:07 server-20260925-030712.tar.gz
-rw------- 1 root root 98 Sep 24 03:11 server-20260924-031122.tar.gz.sha256
Verify its checksum. sha256sum -c looks for the file relative to the current directory, so run it from the backup directory:
sudo sh -c 'cd /var/backups/tar && sha256sum -c server-20260925-030712.tar.gz.sha256'
server-20260925-030712.tar.gz: OK
A matching checksum proves the file has not changed since it was created, but not that it was complete when it was created. Read the whole archive to confirm that tar can reach the end:
sudo tar -tzf /var/backups/tar/server-20260925-030712.tar.gz > /dev/null && echo "archive readable"
archive readable
Do the same for the database dumps. Check the checksums of the newest dump directory:
sudo ls /var/backups/db
sudo sh -c 'cd /var/backups/db/20260925-023412 && sha256sum -c SHA256SUMS'
mysql-your_database.sql.gz: OK
pg-globals.sql.gz: OK
pg-your_database.dump: OK
A complete mysqldump file always ends with a Dump completed comment, and a valid PostgreSQL custom-format dump has a readable table of contents:
sudo zcat /var/backups/db/20260925-023412/mysql-your_database.sql.gz | tail -n 1
sudo pg_restore --list /var/backups/db/20260925-023412/pg-your_database.dump | grep -c 'TABLE DATA'
-- Dump completed on 2026-09-25 2:34:15
21
The second number is the count of tables with data in the dump. Compare it with what you expect for your application.
Step 3 - Restoring files and comparing them
The fastest way to confirm that an archive matches the live data is tar --compare (-d). It reads the archive and reports every file that differs from the filesystem:
sudo tar -dzf /var/backups/tar/server-20260925-030712.tar.gz -C / 2>&1 | head -n 20
var/www/your_domain/wp-content/uploads/2026/09/banner.jpg: Mod time differs
var/www/your_domain/wp-content/uploads/2026/09/banner.jpg: Size differs
home/your_user/.bash_history: Mod time differs
Some differences are expected, because files changed after the backup. Look for unexpected ones, such as Warning: Cannot stat: No such file or directory for files that should exist, or entire directories missing.
Then perform a real restore of a sample into a staging directory, never over the live paths:
sudo mkdir -p /srv/restore-test
sudo tar -xzf /var/backups/tar/server-20260925-030712.tar.gz -C /srv/restore-test etc/nginx var/www/your_domain
Compare the restored copy with the live one. diff -r prints only differences; -q limits it to file names:
sudo diff -rq /srv/restore-test/etc/nginx /etc/nginx && echo "nginx config identical"
nginx config identical
Check that ownership and permissions came back too, since a restore with the wrong owner often breaks the application:
sudo stat -c '%U:%G %a %n' /var/www/your_domain/index.php /srv/restore-test/var/www/your_domain/index.php
www-data:www-data 644 /var/www/your_domain/index.php
www-data:www-data 644 /srv/restore-test/var/www/your_domain/index.php
Remove the staging copy when you are done:
sudo rm -rf /srv/restore-test
Step 4 - Automating a weekly restore test
Manual checks are easy to forget. The following script runs the checks from Steps 2 and 3 on the newest backups and also loads every database dump into a temporary database. If anything fails, it exits with an error, and systemd marks the run as failed. Create the script:
sudo nano /usr/local/bin/restore-test.sh
#!/usr/bin/env bash
# Verify the newest backups: freshness, checksums, readability and database restores.
set -euo pipefail
umask 077
TAR_DIR="/var/backups/tar"
DB_DIR="/var/backups/db"
MAX_AGE_MIN=$((26 * 60))
TEST_DB="restore_check"
fail() { echo "FAIL: $*" >&2; exit 1; }
cd /
# Files: newest archive must be recent, match its checksum and be readable.
latest_tar=$(find "$TAR_DIR" -maxdepth 1 -name '*.tar.gz' -mmin -"$MAX_AGE_MIN" -printf '%T@ %p\n' \
| sort -n | tail -n 1 | cut -d' ' -f2-)
[ -n "$latest_tar" ] || fail "no archive newer than 26 hours in $TAR_DIR"
(cd "$TAR_DIR" && sha256sum -c --quiet "$(basename "$latest_tar").sha256") || fail "checksum mismatch: $latest_tar"
entries=$(tar -tzf "$latest_tar" | wc -l) || fail "cannot read $latest_tar"
[ "$entries" -gt 0 ] || fail "$latest_tar is empty"
echo "OK: $latest_tar ($entries entries)"
# Databases: newest dump directory must be recent and match its checksums.
latest_db=$(find "$DB_DIR" -mindepth 1 -maxdepth 1 -type d -mmin -"$MAX_AGE_MIN" | sort | tail -n 1)
[ -n "$latest_db" ] || fail "no dump directory newer than 26 hours in $DB_DIR"
(cd "$latest_db" && sha256sum -c --quiet SHA256SUMS) || fail "checksum mismatch in $latest_db"
for dump in "$latest_db"/mysql-*.sql.gz; do
[ -e "$dump" ] || continue
mysql -e "DROP DATABASE IF EXISTS $TEST_DB; CREATE DATABASE $TEST_DB"
zcat "$dump" | mysql "$TEST_DB" || fail "restore failed: $dump"
tables=$(mysql -N -B -e "SELECT COUNT(*) FROM information_schema.tables WHERE table_schema = '$TEST_DB'")
mysql -e "DROP DATABASE $TEST_DB"
[ "$tables" -gt 0 ] || fail "$dump restored no tables"
echo "OK: $dump ($tables tables)"
done
for dump in "$latest_db"/pg-*.dump; do
[ -e "$dump" ] || continue
runuser -u postgres -- dropdb --if-exists "$TEST_DB"
runuser -u postgres -- createdb "$TEST_DB"
runuser -u postgres -- pg_restore --no-owner --exit-on-error -d "$TEST_DB" < "$dump" \
|| fail "restore failed: $dump"
tables=$(runuser -u postgres -- psql -At -d "$TEST_DB" -c \
"SELECT count(*) FROM pg_tables WHERE schemaname NOT IN ('pg_catalog', 'information_schema')")
runuser -u postgres -- dropdb "$TEST_DB"
[ "$tables" -gt 0 ] || fail "$dump restored no tables"
echo "OK: $dump ($tables tables)"
done
echo "All restore checks passed"
The script uses a database named restore_check, which it drops and recreates on every run, so make sure no real database has that name. It also treats a dump with zero tables as a failure; if you intentionally keep an empty database, exclude it from your backups. Make it executable and run it:
sudo chmod 700 /usr/local/bin/restore-test.sh
sudo /usr/local/bin/restore-test.sh
OK: /var/backups/tar/server-20260925-030712.tar.gz (18244 entries)
OK: /var/backups/db/20260925-023412/mysql-your_database.sql.gz (23 tables)
OK: /var/backups/db/20260925-023412/pg-your_database.dump (21 tables)
All restore checks passed
Restoring a large database uses CPU and disk, so schedule the test outside peak hours. Create the service:
sudo nano /etc/systemd/system/restore-test.service
[Unit]
Description=Weekly backup restore test
After=mysql.service mariadb.service postgresql.service
[Service]
Type=oneshot
ExecStart=/usr/local/bin/restore-test.sh
Nice=15
IOSchedulingClass=idle
Then the timer, which runs every Sunday at 05:30:
sudo nano /etc/systemd/system/restore-test.timer
[Unit]
Description=Run the backup restore test every week
[Timer]
OnCalendar=Sun *-*-* 05:30:00
Persistent=true
[Install]
WantedBy=timers.target
Enable it and confirm the schedule:
sudo systemctl daemon-reload
sudo systemctl enable --now restore-test.timer
systemctl list-timers restore-test.timer
NEXT LEFT LAST PASSED UNIT ACTIVATES
Sun 2026-09-27 05:30:00 UTC 1 day left - - restore-test.timer restore-test.service
After each run, systemctl status restore-test.service shows whether it passed, and a failed test also appears in systemctl --failed. Make sure whatever monitoring you use for the server alerts on failed units, otherwise a failing test goes unnoticed just like a failing backup.
Tipfor large databases, run this script on a separate restore server that downloads the backups (for example with rclone) instead of on production. That also proves the offsite copy is usable.
Step 5 - Running a full recovery drill
The automated test proves that the backups are readable. A drill proves that you can rebuild the service from them, and tells you how long it takes. Do it at least a few times a year, and after any major change to the server.
Create a fresh server with the same Ubuntu version, log in as a sudo user, and follow only your written procedure, as if the original server were gone. Time each phase with a clock or by prefixing long commands with time. A typical drill for a web server looks like this:
-
Get the backups. Install rclone, recreate its configuration from your password manager, and download the newest backups:
sudo rclone copy offsite-crypt:current/var/backups /var/backups --progress -
Install the software stack with the same packages as production, for example:
sudo apt update sudo apt install nginx php8.3-fpm php8.3-mysql mysql-server certbot python3-certbot-nginx -
Restore the files from the newest archive into place:
sudo tar -xzf /var/backups/tar/server-20260925-030712.tar.gz -C / var/www etc/nginx/sites-available etc/letsencrypt -
Restore the databases, recreating the application user first:
sudo mysql -e "CREATE DATABASE your_database" sudo mysql -e "CREATE USER 'app_user'@'localhost' IDENTIFIED BY 'your_strong_password'" sudo mysql -e "GRANT ALL PRIVILEGES ON your_database.* TO 'app_user'@'localhost'" sudo zcat /var/backups/db/20260925-023412/mysql-your_database.sql.gz | sudo mysql your_databaseUse the password from the application's configuration file, so that the restored site can connect.
-
Enable the site and start the services:
sudo ln -s /etc/nginx/sites-available/your_domain /etc/nginx/sites-enabled/ sudo nginx -t sudo systemctl reload nginx -
Test the application without changing DNS. From your own computer, send a request to the drill server with
curl --resolve, which connects to the given IP while keeping the real host name for TLS and virtual hosts:curl -sI --resolve your_domain:443:drill_server_ip https://your_domain/HTTP/2 200 server: nginx content-type: text/html; charset=UTF-8Also log in to the application and open a few pages that read from the database.
When the site works, delete the drill server. Then write down the results while they are fresh:
- Date, who ran the drill, and which backup (file names and dates) was restored.
- Time for each phase and the total, compared with your RTO.
- The age of the newest restored data, compared with your RPO.
- Every step that was missing from the procedure, every password or key you had to search for, and every error you hit.
- Changes to make before the next drill.
Update your restore procedure with anything you had to improvise. The next person to use it may be restoring a real outage.
Troubleshooting
sha256sum: WARNING: 1 computed checksum did NOT match: the file changed or was corrupted after it was created, often during a copy between servers. Copy it again from the source and compare checksums; if the original also fails, fall back to the previous backup and investigate the disk.gzip: stdin: unexpected end of filewhen reading an archive: the archive is truncated. Check for full disks at the time of the backup withjournalctl -u tar-backup.service.- MySQL restore fails with
ERROR 1227 (42000): Access denied; you need (at least one of) the SUPER or SET_USER_ID privilege(s): the dump contains views or routines with aDEFINERuser, and you are restoring as a user without enough privileges. Restore as root withsudo mysql, or create the definer user first. pg_restore: error: could not execute query: ERROR: role "app_user" does not exist: restore the roles frompg-globals.sql.gzfirst, or add--no-owner --no-privilegesfor a test restore.- The site shows the Nginx default page during the drill: the site configuration was not linked into
sites-enabled, orserver_namedoes not match the host name you used withcurl --resolve.
Conclusion
You now verify backups at three levels: integrity checks that catch corrupted and missing files, an automated weekly restore of every database, and periodic full drills that measure your real recovery time. A backup that has passed all three is one you can rely on. As next steps, you can:
- Run the restore test on a dedicated server that pulls the offsite copy, so it tests exactly what you would use after losing the main server.
- Alert on failed systemd units from your monitoring system, so a failed backup or restore test is noticed within a day.
- Keep the recovery runbook, the rclone configuration and the encryption passwords somewhere that does not depend on the server they protect.
