A backup job that finishes without errors does not prove that you can recover from it. Archives get truncated when a disk fills up, dumps silently skip a database, retention deletes the only good copy, and nobody notices until the day a restore is needed. In this tutorial you will check that your backups are recent and intact, restore files and databases into a safe location, automate a weekly restore test with a systemd timer, and run a full recovery drill on a fresh server running Ubuntu 24.04.

Prerequisites

To follow this tutorial, you will need:

  • A server running Ubuntu 24.04 LTS with a non-root user that has sudo privileges.
  • Existing backups to test. The examples assume the layout used in the other CubePath backup guides:
    • compressed tar archives with .sha256 files in /var/backups/tar,
    • dated database dump directories with a SHA256SUMS file in /var/backups/db, containing mysql-<name>.sql.gz (mysqldump) and pg-<name>.dump (pg_dump custom format) files.
  • For the recovery drill in Step 5: a second, empty server, for example a small CubePath VPS that you delete afterwards.

If your backups live elsewhere or use other names, adjust the paths in the commands.

Step 1 - Deciding what a successful restore means

Before testing, write down what you are testing against. Two numbers matter:

  • Recovery point objective (RPO): how much data you can afford to lose. With a nightly backup, the RPO is up to 24 hours.
  • Recovery time objective (RTO): how long the service may be down while you restore. Only a real drill tells you whether you meet it.

Then list what has to work after a restore. For a typical web server, that means:

CheckHow to verify
Latest backup is recent enoughFile or directory age is below the RPO
Files are intactChecksums match, archives can be read to the end
Important files are presentA restored sample matches the live copy
Databases loadDumps restore into an empty database without errors
Application worksThe site answers from the restored server

The rest of this tutorial turns each row into a command.

Step 2 - Checking freshness and integrity

Start with the cheapest checks: is there a recent backup, and is it undamaged? Find the newest tar archive and its age:

sudo ls -lt /var/backups/tar | head -n 4
total 124M
-rw------- 1 root root  98 Sep 25 03:07 server-20260925-030712.tar.gz.sha256
-rw------- 1 root root 41M Sep 25 03:07 server-20260925-030712.tar.gz
-rw------- 1 root root  98 Sep 24 03:11 server-20260924-031122.tar.gz.sha256

Verify its checksum. sha256sum -c looks for the file relative to the current directory, so run it from the backup directory:

sudo sh -c 'cd /var/backups/tar && sha256sum -c server-20260925-030712.tar.gz.sha256'
server-20260925-030712.tar.gz: OK

A matching checksum proves the file has not changed since it was created, but not that it was complete when it was created. Read the whole archive to confirm that tar can reach the end:

sudo tar -tzf /var/backups/tar/server-20260925-030712.tar.gz > /dev/null && echo "archive readable"
archive readable

Do the same for the database dumps. Check the checksums of the newest dump directory:

sudo ls /var/backups/db
sudo sh -c 'cd /var/backups/db/20260925-023412 && sha256sum -c SHA256SUMS'
mysql-your_database.sql.gz: OK
pg-globals.sql.gz: OK
pg-your_database.dump: OK

A complete mysqldump file always ends with a Dump completed comment, and a valid PostgreSQL custom-format dump has a readable table of contents:

sudo zcat /var/backups/db/20260925-023412/mysql-your_database.sql.gz | tail -n 1
sudo pg_restore --list /var/backups/db/20260925-023412/pg-your_database.dump | grep -c 'TABLE DATA'
-- Dump completed on 2026-09-25  2:34:15
21

The second number is the count of tables with data in the dump. Compare it with what you expect for your application.

Step 3 - Restoring files and comparing them

The fastest way to confirm that an archive matches the live data is tar --compare (-d). It reads the archive and reports every file that differs from the filesystem:

sudo tar -dzf /var/backups/tar/server-20260925-030712.tar.gz -C / 2>&1 | head -n 20
var/www/your_domain/wp-content/uploads/2026/09/banner.jpg: Mod time differs
var/www/your_domain/wp-content/uploads/2026/09/banner.jpg: Size differs
home/your_user/.bash_history: Mod time differs

Some differences are expected, because files changed after the backup. Look for unexpected ones, such as Warning: Cannot stat: No such file or directory for files that should exist, or entire directories missing.

Then perform a real restore of a sample into a staging directory, never over the live paths:

sudo mkdir -p /srv/restore-test
sudo tar -xzf /var/backups/tar/server-20260925-030712.tar.gz -C /srv/restore-test etc/nginx var/www/your_domain

Compare the restored copy with the live one. diff -r prints only differences; -q limits it to file names:

sudo diff -rq /srv/restore-test/etc/nginx /etc/nginx && echo "nginx config identical"
nginx config identical

Check that ownership and permissions came back too, since a restore with the wrong owner often breaks the application:

sudo stat -c '%U:%G %a %n' /var/www/your_domain/index.php /srv/restore-test/var/www/your_domain/index.php
www-data:www-data 644 /var/www/your_domain/index.php
www-data:www-data 644 /srv/restore-test/var/www/your_domain/index.php

Remove the staging copy when you are done:

sudo rm -rf /srv/restore-test

Step 4 - Automating a weekly restore test

Manual checks are easy to forget. The following script runs the checks from Steps 2 and 3 on the newest backups and also loads every database dump into a temporary database. If anything fails, it exits with an error, and systemd marks the run as failed. Create the script:

sudo nano /usr/local/bin/restore-test.sh
#!/usr/bin/env bash
# Verify the newest backups: freshness, checksums, readability and database restores.
set -euo pipefail
umask 077

TAR_DIR="/var/backups/tar"
DB_DIR="/var/backups/db"
MAX_AGE_MIN=$((26 * 60))
TEST_DB="restore_check"

fail() { echo "FAIL: $*" >&2; exit 1; }
cd /

# Files: newest archive must be recent, match its checksum and be readable.
latest_tar=$(find "$TAR_DIR" -maxdepth 1 -name '*.tar.gz' -mmin -"$MAX_AGE_MIN" -printf '%T@ %p\n' \
    | sort -n | tail -n 1 | cut -d' ' -f2-)
[ -n "$latest_tar" ] || fail "no archive newer than 26 hours in $TAR_DIR"
(cd "$TAR_DIR" && sha256sum -c --quiet "$(basename "$latest_tar").sha256") || fail "checksum mismatch: $latest_tar"
entries=$(tar -tzf "$latest_tar" | wc -l) || fail "cannot read $latest_tar"
[ "$entries" -gt 0 ] || fail "$latest_tar is empty"
echo "OK: $latest_tar ($entries entries)"

# Databases: newest dump directory must be recent and match its checksums.
latest_db=$(find "$DB_DIR" -mindepth 1 -maxdepth 1 -type d -mmin -"$MAX_AGE_MIN" | sort | tail -n 1)
[ -n "$latest_db" ] || fail "no dump directory newer than 26 hours in $DB_DIR"
(cd "$latest_db" && sha256sum -c --quiet SHA256SUMS) || fail "checksum mismatch in $latest_db"

for dump in "$latest_db"/mysql-*.sql.gz; do
    [ -e "$dump" ] || continue
    mysql -e "DROP DATABASE IF EXISTS $TEST_DB; CREATE DATABASE $TEST_DB"
    zcat "$dump" | mysql "$TEST_DB" || fail "restore failed: $dump"
    tables=$(mysql -N -B -e "SELECT COUNT(*) FROM information_schema.tables WHERE table_schema = '$TEST_DB'")
    mysql -e "DROP DATABASE $TEST_DB"
    [ "$tables" -gt 0 ] || fail "$dump restored no tables"
    echo "OK: $dump ($tables tables)"
done

for dump in "$latest_db"/pg-*.dump; do
    [ -e "$dump" ] || continue
    runuser -u postgres -- dropdb --if-exists "$TEST_DB"
    runuser -u postgres -- createdb "$TEST_DB"
    runuser -u postgres -- pg_restore --no-owner --exit-on-error -d "$TEST_DB" < "$dump" \
        || fail "restore failed: $dump"
    tables=$(runuser -u postgres -- psql -At -d "$TEST_DB" -c \
        "SELECT count(*) FROM pg_tables WHERE schemaname NOT IN ('pg_catalog', 'information_schema')")
    runuser -u postgres -- dropdb "$TEST_DB"
    [ "$tables" -gt 0 ] || fail "$dump restored no tables"
    echo "OK: $dump ($tables tables)"
done

echo "All restore checks passed"

The script uses a database named restore_check, which it drops and recreates on every run, so make sure no real database has that name. It also treats a dump with zero tables as a failure; if you intentionally keep an empty database, exclude it from your backups. Make it executable and run it:

sudo chmod 700 /usr/local/bin/restore-test.sh
sudo /usr/local/bin/restore-test.sh
OK: /var/backups/tar/server-20260925-030712.tar.gz (18244 entries)
OK: /var/backups/db/20260925-023412/mysql-your_database.sql.gz (23 tables)
OK: /var/backups/db/20260925-023412/pg-your_database.dump (21 tables)
All restore checks passed

Restoring a large database uses CPU and disk, so schedule the test outside peak hours. Create the service:

sudo nano /etc/systemd/system/restore-test.service
[Unit]
Description=Weekly backup restore test
After=mysql.service mariadb.service postgresql.service

[Service]
Type=oneshot
ExecStart=/usr/local/bin/restore-test.sh
Nice=15
IOSchedulingClass=idle

Then the timer, which runs every Sunday at 05:30:

sudo nano /etc/systemd/system/restore-test.timer
[Unit]
Description=Run the backup restore test every week

[Timer]
OnCalendar=Sun *-*-* 05:30:00
Persistent=true

[Install]
WantedBy=timers.target

Enable it and confirm the schedule:

sudo systemctl daemon-reload
sudo systemctl enable --now restore-test.timer
systemctl list-timers restore-test.timer
NEXT                        LEFT       LAST PASSED UNIT               ACTIVATES
Sun 2026-09-27 05:30:00 UTC 1 day left -    -      restore-test.timer restore-test.service

After each run, systemctl status restore-test.service shows whether it passed, and a failed test also appears in systemctl --failed. Make sure whatever monitoring you use for the server alerts on failed units, otherwise a failing test goes unnoticed just like a failing backup.

Step 5 - Running a full recovery drill

The automated test proves that the backups are readable. A drill proves that you can rebuild the service from them, and tells you how long it takes. Do it at least a few times a year, and after any major change to the server.

Create a fresh server with the same Ubuntu version, log in as a sudo user, and follow only your written procedure, as if the original server were gone. Time each phase with a clock or by prefixing long commands with time. A typical drill for a web server looks like this:

  1. Get the backups. Install rclone, recreate its configuration from your password manager, and download the newest backups:

    sudo rclone copy offsite-crypt:current/var/backups /var/backups --progress
    
  2. Install the software stack with the same packages as production, for example:

    sudo apt update
    sudo apt install nginx php8.3-fpm php8.3-mysql mysql-server certbot python3-certbot-nginx
    
  3. Restore the files from the newest archive into place:

    sudo tar -xzf /var/backups/tar/server-20260925-030712.tar.gz -C / var/www etc/nginx/sites-available etc/letsencrypt
    
  4. Restore the databases, recreating the application user first:

    sudo mysql -e "CREATE DATABASE your_database"
    sudo mysql -e "CREATE USER 'app_user'@'localhost' IDENTIFIED BY 'your_strong_password'"
    sudo mysql -e "GRANT ALL PRIVILEGES ON your_database.* TO 'app_user'@'localhost'"
    sudo zcat /var/backups/db/20260925-023412/mysql-your_database.sql.gz | sudo mysql your_database
    

    Use the password from the application's configuration file, so that the restored site can connect.

  5. Enable the site and start the services:

    sudo ln -s /etc/nginx/sites-available/your_domain /etc/nginx/sites-enabled/
    sudo nginx -t
    sudo systemctl reload nginx
    
  6. Test the application without changing DNS. From your own computer, send a request to the drill server with curl --resolve, which connects to the given IP while keeping the real host name for TLS and virtual hosts:

    curl -sI --resolve your_domain:443:drill_server_ip https://your_domain/
    
    HTTP/2 200
    server: nginx
    content-type: text/html; charset=UTF-8
    

    Also log in to the application and open a few pages that read from the database.

When the site works, delete the drill server. Then write down the results while they are fresh:

  • Date, who ran the drill, and which backup (file names and dates) was restored.
  • Time for each phase and the total, compared with your RTO.
  • The age of the newest restored data, compared with your RPO.
  • Every step that was missing from the procedure, every password or key you had to search for, and every error you hit.
  • Changes to make before the next drill.

Update your restore procedure with anything you had to improvise. The next person to use it may be restoring a real outage.

Troubleshooting

  • sha256sum: WARNING: 1 computed checksum did NOT match: the file changed or was corrupted after it was created, often during a copy between servers. Copy it again from the source and compare checksums; if the original also fails, fall back to the previous backup and investigate the disk.
  • gzip: stdin: unexpected end of file when reading an archive: the archive is truncated. Check for full disks at the time of the backup with journalctl -u tar-backup.service.
  • MySQL restore fails with ERROR 1227 (42000): Access denied; you need (at least one of) the SUPER or SET_USER_ID privilege(s): the dump contains views or routines with a DEFINER user, and you are restoring as a user without enough privileges. Restore as root with sudo mysql, or create the definer user first.
  • pg_restore: error: could not execute query: ERROR: role "app_user" does not exist: restore the roles from pg-globals.sql.gz first, or add --no-owner --no-privileges for a test restore.
  • The site shows the Nginx default page during the drill: the site configuration was not linked into sites-enabled, or server_name does not match the host name you used with curl --resolve.

Conclusion

You now verify backups at three levels: integrity checks that catch corrupted and missing files, an automated weekly restore of every database, and periodic full drills that measure your real recovery time. A backup that has passed all three is one you can rely on. As next steps, you can:

  • Run the restore test on a dedicated server that pulls the offsite copy, so it tests exactly what you would use after losing the main server.
  • Alert on failed systemd units from your monitoring system, so a failed backup or restore test is noticed within a day.
  • Keep the recovery runbook, the rclone configuration and the encryption passwords somewhere that does not depend on the server they protect.