A backup you have never restored is only a hope. Backups fail silently all the time: a job keeps running but excludes a new directory, a database dump is truncated, a repository password is lost. The only reliable check is to restore the backup regularly and verify the result. In this tutorial you will set up a dedicated test server on Ubuntu 24.04 that, every week, verifies a restic repository, restores the latest snapshot, loads the MySQL dump it contains into a scratch database, checks the data and reports to a monitoring service, alerting you if any step fails.

Prerequisites

To follow this guide you need:

  • An existing restic repository that backs up a production server (called web01 below), including the site files and a MySQL dump. This guide assumes the dump is created before each backup with a command such as mysqldump --single-transaction appdb | gzip > /var/backups/mysql/appdb.sql.gz, without --databases, so it can be loaded into a database with a different name.
  • A separate test server running Ubuntu 24.04 LTS with a non-root sudo user, for example a small CubePath VPS. It needs enough disk for a full restore and must be able to reach the repository. Do not run the test on the production server: a restore test must prove that the backup works without the original machine.
  • The repository location and password.
  • Optionally, a free account on a dead man's switch service such as Healthchecks.io, which alerts you both when the test fails and when it stops running.

Step 1 - Installing the tools

Install restic, the MySQL server that will host the scratch database, and jq to parse restic's JSON output:

sudo apt update
sudo apt install restic mysql-server jq

Check the restic version:

restic version
restic 0.16.4 compiled with go1.22.2 on linux/amd64

The test script runs as root, and on Ubuntu the MySQL root account authenticates through the Unix socket, so mysql works as root without a password.

Step 2 - Storing the repository credentials

Keep the repository settings in a root-only directory. Create it:

sudo install -d -m 0700 /etc/dr-test

Store the repository password in its own file. Replace your_repository_password with the real one:

echo 'your_repository_password' | sudo tee /etc/dr-test/restic.pass > /dev/null
sudo chmod 600 /etc/dr-test/restic.pass

Create the environment file:

sudo nano /etc/dr-test/restic.env

Add the repository location. This example uses an SFTP repository; for S3-compatible storage, use the s3: URL and add AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY:

RESTIC_REPOSITORY=sftp:backup@your_backup_host:/srv/restic/web01
RESTIC_PASSWORD_FILE=/etc/dr-test/restic.pass

Protect it:

sudo chmod 600 /etc/dr-test/restic.env

For an SFTP repository, root on the test server needs an SSH key that the backup host accepts. Create it with sudo ssh-keygen -t ed25519 and add the public key to the backup account on your_backup_host.

Test access by listing the snapshots:

sudo bash -c 'set -a; . /etc/dr-test/restic.env; restic snapshots --host web01'
ID        Time                 Host   Tags  Paths
------------------------------------------------------------------------
4f2a9c1e  2026-09-24 02:00:05  web01        /etc
                                            /var/backups/mysql
                                            /var/www
9b7d03aa  2026-09-25 02:00:04  web01        /etc
                                            /var/backups/mysql
                                            /var/www
------------------------------------------------------------------------
2 snapshots

Step 3 - Deciding what "restored correctly" means

Before writing the script, decide which checks prove the backup is usable. Good checks are specific to your application:

CheckWhy it matters
Repository integrity (restic check)Detects corrupted or missing data in the repository.
Age of the latest snapshotDetects a backup job that stopped running. This is your real recovery point.
A known file exists and is not emptyDetects excluded or missing directories.
The dump loads without errorsDetects truncated or corrupt dumps.
A key table has a plausible row countDetects empty or partial dumps that still load.
Total restore timeMeasures your real recovery time, not the one in a document.

For this example, the key file is /var/www/your_domain/index.php, and the key table is users, which should always have more than 1,000 rows. Adjust these values to your application.

Step 4 - Writing the restore test script

Create the script:

sudo nano /usr/local/sbin/dr-restore-test

Add the following content and adjust the settings block at the top:

#!/usr/bin/env bash
# Restore the latest restic snapshot and verify that it is usable.
set -euo pipefail

# Settings
BACKUP_HOST="web01"
MAX_AGE_HOURS=26
WEB_ROOT="/var/www/your_domain"
KEY_FILE="index.php"
DUMP_FILE="/var/backups/mysql/appdb.sql.gz"
CHECK_TABLE="users"
MIN_ROWS=1000
TEST_DB="dr_test"
PING_URL=""   # for example https://hc-ping.com/your_check_uuid

set -a
. /etc/dr-test/restic.env
set +a

workdir="$(mktemp -d /var/tmp/dr-test.XXXXXX)"
cleanup() {
    rm -rf "$workdir"
    mysql -e "DROP DATABASE IF EXISTS \`$TEST_DB\`" || true
}
trap cleanup EXIT

fail() {
    echo "FAIL: $*" >&2
    exit 1
}

started="$(date +%s)"

echo "Checking repository integrity (reading 5% of the data)"
restic check --read-data-subset=5%

snapshot_time="$(restic snapshots --host "$BACKUP_HOST" --json | jq -r 'sort_by(.time) | last | .time')"
[[ -n "$snapshot_time" && "$snapshot_time" != "null" ]] || fail "no snapshots found for $BACKUP_HOST"
age_hours=$(( ($(date +%s) - $(date -d "$snapshot_time" +%s)) / 3600 ))
echo "Latest snapshot: $snapshot_time (${age_hours}h old)"
(( age_hours <= MAX_AGE_HOURS )) || fail "latest snapshot is ${age_hours}h old (limit ${MAX_AGE_HOURS}h)"

echo "Restoring latest snapshot to $workdir"
restic restore latest --host "$BACKUP_HOST" --target "$workdir" \
    --include "$WEB_ROOT" --include "$DUMP_FILE"

[[ -s "$workdir$WEB_ROOT/$KEY_FILE" ]] || fail "$WEB_ROOT/$KEY_FILE missing or empty"
echo "Files OK: $(find "$workdir$WEB_ROOT" -type f | wc -l) files restored"

echo "Loading database dump into $TEST_DB"
gzip -t "$workdir$DUMP_FILE" || fail "dump file is corrupt"
mysql -e "DROP DATABASE IF EXISTS \`$TEST_DB\`; CREATE DATABASE \`$TEST_DB\`"
gunzip -c "$workdir$DUMP_FILE" | mysql "$TEST_DB"

rows="$(mysql -N -B -e "SELECT COUNT(*) FROM \`$TEST_DB\`.\`$CHECK_TABLE\`")"
(( rows >= MIN_ROWS )) || fail "table $CHECK_TABLE has $rows rows (expected at least $MIN_ROWS)"
echo "Database OK: $CHECK_TABLE has $rows rows"

duration=$(( $(date +%s) - started ))
echo "PASS: restore test finished in ${duration}s, snapshot age ${age_hours}h"
printf '%s snapshot_age_hours=%s duration_seconds=%s\n' \
    "$(date --iso-8601=seconds)" "$age_hours" "$duration" >> /var/lib/dr-test/results.log

if [[ -n "$PING_URL" ]]; then
    curl -fsS -m 10 --retry 3 "$PING_URL" > /dev/null
fi

A few design choices worth noting:

  • set -euo pipefail stops the script at the first failed command, including a failure in the middle of the gunzip | mysql pipeline, and the non-zero exit marks the systemd unit as failed.
  • The trap removes the restored files and the scratch database even when a check fails, so the next run starts clean.
  • Each result is appended to /var/lib/dr-test/results.log, which gives you a history of recovery times.

Make the script executable:

sudo chmod 750 /usr/local/sbin/dr-restore-test

Step 5 - Running the test manually

Create the results directory for this first manual run (systemd will manage it later) and run the script:

sudo install -d -m 0750 /var/lib/dr-test
sudo /usr/local/sbin/dr-restore-test
Checking repository integrity (reading 5% of the data)
...
no errors were found
Latest snapshot: 2026-09-25T02:00:04.512837194Z (9h old)
Restoring latest snapshot to /var/tmp/dr-test.Xk31pQ
...
Files OK: 18342 files restored
Loading database dump into dr_test
Database OK: users has 48213 rows
PASS: restore test finished in 412s, snapshot age 9h

If the dump fails to load with an error about a missing DEFINER user, create that user on the test server or strip definers when dumping. That kind of finding is exactly what this test is for.

Now prove that the test can fail. Temporarily set MIN_ROWS=100000000 in the script and run it again:

FAIL: table users has 48213 rows (expected at least 100000000)

Right after the failing run, check the exit code, which must not be zero. Then restore the real value in the script:

echo $?

Step 6 - Scheduling the test with a systemd timer

Create the service unit:

sudo nano /etc/systemd/system/dr-restore-test.service
[Unit]
Description=Automated disaster recovery restore test
Wants=network-online.target
After=network-online.target mysql.service
OnFailure=dr-restore-test-alert.service

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/dr-restore-test
StateDirectory=dr-test
TimeoutStartSec=3h
Nice=10
IOSchedulingClass=idle

StateDirectory=dr-test makes systemd create /var/lib/dr-test if it does not exist. OnFailure= starts an alert unit when the test fails. Create that unit, which reports the failure to your monitoring service:

sudo nano /etc/systemd/system/dr-restore-test-alert.service
[Unit]
Description=Report a failed disaster recovery restore test

[Service]
Type=oneshot
ExecStart=/usr/bin/curl -fsS -m 10 --retry 3 https://hc-ping.com/your_check_uuid/fail

If you do not use Healthchecks.io, replace the command with whatever sends an alert in your environment (a webhook, an email with a configured MTA). Now create the timer:

sudo nano /etc/systemd/system/dr-restore-test.timer
[Unit]
Description=Weekly disaster recovery restore test

[Timer]
OnCalendar=Sun *-*-* 04:00:00
RandomizedDelaySec=15m
Persistent=true

[Install]
WantedBy=timers.target

Persistent=true runs a missed test at the next boot if the server was off at the scheduled time. Load the units and enable the timer:

sudo systemctl daemon-reload
sudo systemctl enable --now dr-restore-test.timer

Check the next run:

systemctl list-timers dr-restore-test.timer
NEXT                        LEFT     LAST PASSED UNIT                  ACTIVATES
Sun 2026-09-27 04:07:31 UTC 1 day 16h -    -      dr-restore-test.timer dr-restore-test.service

Step 7 - Verifying the scheduled run

Trigger the service once through systemd to confirm it works in that environment, not only from your shell:

sudo systemctl start dr-restore-test.service

The command returns when the test finishes. Read its output from the journal:

sudo journalctl -u dr-restore-test.service -n 20 --no-pager
Sep 25 11:32:08 drtest dr-restore-test[4121]: Database OK: users has 48213 rows
Sep 25 11:32:08 drtest dr-restore-test[4121]: PASS: restore test finished in 405s, snapshot age 9h
Sep 25 11:32:08 drtest systemd[1]: dr-restore-test.service: Deactivated successfully.

In Healthchecks.io, set the check's period to one week and the grace time to a few hours. You will then get an alert if the test fails and also if it silently stops running.

Review the result history from time to time to spot trends, such as restore times growing with the data:

cat /var/lib/dr-test/results.log

Step 8 - Complementing automation with manual drills

The automated test proves that the data can be recovered. It does not prove that your team can rebuild the service. Schedule a manual drill every quarter that covers what a script cannot:

  • Build a new server from scratch following only your written runbook.
  • Restore the application, point a test host name at it and run the application's own checks.
  • Time the whole exercise and compare it with your recovery time objective.
  • Fix every step of the runbook that was wrong or missing.

Troubleshooting

Fatal: unable to open config file or wrong password. The environment file is not loaded or points to the wrong repository. Run the listing command from Step 2 to test it.

repository is already locked. A backup or another restic command is running against the repository. Schedule the test outside the backup window, or remove a stale lock with restic unlock once you are sure no process is using it.

The dump loads but the row count is zero. The backup job probably dumps a different database or the dump ran before the data was written. Check the dump command on the production server.

The service times out. Raise TimeoutStartSec= or reduce --read-data-subset; a full restic check --read-data can run less often, for example monthly.

Conclusion

You now have a weekly, fully automated restore test that checks repository integrity, backup freshness, file contents and database contents, records recovery times and alerts you both on failure and on silence. Extend it with checks specific to your application, such as starting the application against the restored database and requesting a page. Next, add the same test for every other system you back up, and keep the quarterly manual drill on the calendar.