A backup you have never restored is only a hope. Backups fail silently all the time: a job keeps running but excludes a new directory, a database dump is truncated, a repository password is lost. The only reliable check is to restore the backup regularly and verify the result. In this tutorial you will set up a dedicated test server on Ubuntu 24.04 that, every week, verifies a restic repository, restores the latest snapshot, loads the MySQL dump it contains into a scratch database, checks the data and reports to a monitoring service, alerting you if any step fails.
Prerequisites
To follow this guide you need:
- An existing restic repository that backs up a production server (called
web01below), including the site files and a MySQL dump. This guide assumes the dump is created before each backup with a command such asmysqldump --single-transaction appdb | gzip > /var/backups/mysql/appdb.sql.gz, without--databases, so it can be loaded into a database with a different name. - A separate test server running Ubuntu 24.04 LTS with a non-root sudo user, for example a small CubePath VPS. It needs enough disk for a full restore and must be able to reach the repository. Do not run the test on the production server: a restore test must prove that the backup works without the original machine.
- The repository location and password.
- Optionally, a free account on a dead man's switch service such as Healthchecks.io, which alerts you both when the test fails and when it stops running.
Step 1 - Installing the tools
Install restic, the MySQL server that will host the scratch database, and jq to parse restic's JSON output:
sudo apt update
sudo apt install restic mysql-server jq
Check the restic version:
restic version
restic 0.16.4 compiled with go1.22.2 on linux/amd64
The test script runs as root, and on Ubuntu the MySQL root account authenticates through the Unix socket, so mysql works as root without a password.
Step 2 - Storing the repository credentials
Keep the repository settings in a root-only directory. Create it:
sudo install -d -m 0700 /etc/dr-test
Store the repository password in its own file. Replace your_repository_password with the real one:
echo 'your_repository_password' | sudo tee /etc/dr-test/restic.pass > /dev/null
sudo chmod 600 /etc/dr-test/restic.pass
Create the environment file:
sudo nano /etc/dr-test/restic.env
Add the repository location. This example uses an SFTP repository; for S3-compatible storage, use the s3: URL and add AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY:
RESTIC_REPOSITORY=sftp:backup@your_backup_host:/srv/restic/web01
RESTIC_PASSWORD_FILE=/etc/dr-test/restic.pass
Protect it:
sudo chmod 600 /etc/dr-test/restic.env
For an SFTP repository, root on the test server needs an SSH key that the backup host accepts. Create it with sudo ssh-keygen -t ed25519 and add the public key to the backup account on your_backup_host.
Test access by listing the snapshots:
sudo bash -c 'set -a; . /etc/dr-test/restic.env; restic snapshots --host web01'
ID Time Host Tags Paths
------------------------------------------------------------------------
4f2a9c1e 2026-09-24 02:00:05 web01 /etc
/var/backups/mysql
/var/www
9b7d03aa 2026-09-25 02:00:04 web01 /etc
/var/backups/mysql
/var/www
------------------------------------------------------------------------
2 snapshots
Step 3 - Deciding what "restored correctly" means
Before writing the script, decide which checks prove the backup is usable. Good checks are specific to your application:
| Check | Why it matters |
|---|---|
Repository integrity (restic check) | Detects corrupted or missing data in the repository. |
| Age of the latest snapshot | Detects a backup job that stopped running. This is your real recovery point. |
| A known file exists and is not empty | Detects excluded or missing directories. |
| The dump loads without errors | Detects truncated or corrupt dumps. |
| A key table has a plausible row count | Detects empty or partial dumps that still load. |
| Total restore time | Measures your real recovery time, not the one in a document. |
For this example, the key file is /var/www/your_domain/index.php, and the key table is users, which should always have more than 1,000 rows. Adjust these values to your application.
Step 4 - Writing the restore test script
Create the script:
sudo nano /usr/local/sbin/dr-restore-test
Add the following content and adjust the settings block at the top:
#!/usr/bin/env bash
# Restore the latest restic snapshot and verify that it is usable.
set -euo pipefail
# Settings
BACKUP_HOST="web01"
MAX_AGE_HOURS=26
WEB_ROOT="/var/www/your_domain"
KEY_FILE="index.php"
DUMP_FILE="/var/backups/mysql/appdb.sql.gz"
CHECK_TABLE="users"
MIN_ROWS=1000
TEST_DB="dr_test"
PING_URL="" # for example https://hc-ping.com/your_check_uuid
set -a
. /etc/dr-test/restic.env
set +a
workdir="$(mktemp -d /var/tmp/dr-test.XXXXXX)"
cleanup() {
rm -rf "$workdir"
mysql -e "DROP DATABASE IF EXISTS \`$TEST_DB\`" || true
}
trap cleanup EXIT
fail() {
echo "FAIL: $*" >&2
exit 1
}
started="$(date +%s)"
echo "Checking repository integrity (reading 5% of the data)"
restic check --read-data-subset=5%
snapshot_time="$(restic snapshots --host "$BACKUP_HOST" --json | jq -r 'sort_by(.time) | last | .time')"
[[ -n "$snapshot_time" && "$snapshot_time" != "null" ]] || fail "no snapshots found for $BACKUP_HOST"
age_hours=$(( ($(date +%s) - $(date -d "$snapshot_time" +%s)) / 3600 ))
echo "Latest snapshot: $snapshot_time (${age_hours}h old)"
(( age_hours <= MAX_AGE_HOURS )) || fail "latest snapshot is ${age_hours}h old (limit ${MAX_AGE_HOURS}h)"
echo "Restoring latest snapshot to $workdir"
restic restore latest --host "$BACKUP_HOST" --target "$workdir" \
--include "$WEB_ROOT" --include "$DUMP_FILE"
[[ -s "$workdir$WEB_ROOT/$KEY_FILE" ]] || fail "$WEB_ROOT/$KEY_FILE missing or empty"
echo "Files OK: $(find "$workdir$WEB_ROOT" -type f | wc -l) files restored"
echo "Loading database dump into $TEST_DB"
gzip -t "$workdir$DUMP_FILE" || fail "dump file is corrupt"
mysql -e "DROP DATABASE IF EXISTS \`$TEST_DB\`; CREATE DATABASE \`$TEST_DB\`"
gunzip -c "$workdir$DUMP_FILE" | mysql "$TEST_DB"
rows="$(mysql -N -B -e "SELECT COUNT(*) FROM \`$TEST_DB\`.\`$CHECK_TABLE\`")"
(( rows >= MIN_ROWS )) || fail "table $CHECK_TABLE has $rows rows (expected at least $MIN_ROWS)"
echo "Database OK: $CHECK_TABLE has $rows rows"
duration=$(( $(date +%s) - started ))
echo "PASS: restore test finished in ${duration}s, snapshot age ${age_hours}h"
printf '%s snapshot_age_hours=%s duration_seconds=%s\n' \
"$(date --iso-8601=seconds)" "$age_hours" "$duration" >> /var/lib/dr-test/results.log
if [[ -n "$PING_URL" ]]; then
curl -fsS -m 10 --retry 3 "$PING_URL" > /dev/null
fi
A few design choices worth noting:
set -euo pipefailstops the script at the first failed command, including a failure in the middle of thegunzip | mysqlpipeline, and the non-zero exit marks the systemd unit as failed.- The
trapremoves the restored files and the scratch database even when a check fails, so the next run starts clean. - Each result is appended to
/var/lib/dr-test/results.log, which gives you a history of recovery times.
Make the script executable:
sudo chmod 750 /usr/local/sbin/dr-restore-test
Step 5 - Running the test manually
Create the results directory for this first manual run (systemd will manage it later) and run the script:
sudo install -d -m 0750 /var/lib/dr-test
sudo /usr/local/sbin/dr-restore-test
Checking repository integrity (reading 5% of the data)
...
no errors were found
Latest snapshot: 2026-09-25T02:00:04.512837194Z (9h old)
Restoring latest snapshot to /var/tmp/dr-test.Xk31pQ
...
Files OK: 18342 files restored
Loading database dump into dr_test
Database OK: users has 48213 rows
PASS: restore test finished in 412s, snapshot age 9h
If the dump fails to load with an error about a missing DEFINER user, create that user on the test server or strip definers when dumping. That kind of finding is exactly what this test is for.
Now prove that the test can fail. Temporarily set MIN_ROWS=100000000 in the script and run it again:
FAIL: table users has 48213 rows (expected at least 100000000)
Right after the failing run, check the exit code, which must not be zero. Then restore the real value in the script:
echo $?
Step 6 - Scheduling the test with a systemd timer
Create the service unit:
sudo nano /etc/systemd/system/dr-restore-test.service
[Unit]
Description=Automated disaster recovery restore test
Wants=network-online.target
After=network-online.target mysql.service
OnFailure=dr-restore-test-alert.service
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/dr-restore-test
StateDirectory=dr-test
TimeoutStartSec=3h
Nice=10
IOSchedulingClass=idle
StateDirectory=dr-test makes systemd create /var/lib/dr-test if it does not exist. OnFailure= starts an alert unit when the test fails. Create that unit, which reports the failure to your monitoring service:
sudo nano /etc/systemd/system/dr-restore-test-alert.service
[Unit]
Description=Report a failed disaster recovery restore test
[Service]
Type=oneshot
ExecStart=/usr/bin/curl -fsS -m 10 --retry 3 https://hc-ping.com/your_check_uuid/fail
If you do not use Healthchecks.io, replace the command with whatever sends an alert in your environment (a webhook, an email with a configured MTA). Now create the timer:
sudo nano /etc/systemd/system/dr-restore-test.timer
[Unit]
Description=Weekly disaster recovery restore test
[Timer]
OnCalendar=Sun *-*-* 04:00:00
RandomizedDelaySec=15m
Persistent=true
[Install]
WantedBy=timers.target
Persistent=true runs a missed test at the next boot if the server was off at the scheduled time. Load the units and enable the timer:
sudo systemctl daemon-reload
sudo systemctl enable --now dr-restore-test.timer
Check the next run:
systemctl list-timers dr-restore-test.timer
NEXT LEFT LAST PASSED UNIT ACTIVATES
Sun 2026-09-27 04:07:31 UTC 1 day 16h - - dr-restore-test.timer dr-restore-test.service
Step 7 - Verifying the scheduled run
Trigger the service once through systemd to confirm it works in that environment, not only from your shell:
sudo systemctl start dr-restore-test.service
The command returns when the test finishes. Read its output from the journal:
sudo journalctl -u dr-restore-test.service -n 20 --no-pager
Sep 25 11:32:08 drtest dr-restore-test[4121]: Database OK: users has 48213 rows
Sep 25 11:32:08 drtest dr-restore-test[4121]: PASS: restore test finished in 405s, snapshot age 9h
Sep 25 11:32:08 drtest systemd[1]: dr-restore-test.service: Deactivated successfully.
In Healthchecks.io, set the check's period to one week and the grace time to a few hours. You will then get an alert if the test fails and also if it silently stops running.
Review the result history from time to time to spot trends, such as restore times growing with the data:
cat /var/lib/dr-test/results.log
Step 8 - Complementing automation with manual drills
The automated test proves that the data can be recovered. It does not prove that your team can rebuild the service. Schedule a manual drill every quarter that covers what a script cannot:
- Build a new server from scratch following only your written runbook.
- Restore the application, point a test host name at it and run the application's own checks.
- Time the whole exercise and compare it with your recovery time objective.
- Fix every step of the runbook that was wrong or missing.
Troubleshooting
Fatal: unable to open config file or wrong password. The environment file is not loaded or points to the wrong repository. Run the listing command from Step 2 to test it.
repository is already locked. A backup or another restic command is running against the repository. Schedule the test outside the backup window, or remove a stale lock with restic unlock once you are sure no process is using it.
The dump loads but the row count is zero. The backup job probably dumps a different database or the dump ran before the data was written. Check the dump command on the production server.
The service times out. Raise TimeoutStartSec= or reduce --read-data-subset; a full restic check --read-data can run less often, for example monthly.
Conclusion
You now have a weekly, fully automated restore test that checks repository integrity, backup freshness, file contents and database contents, records recovery times and alerts you both on failure and on silence. Extend it with checks specific to your application, such as starting the application against the restored database and requesting a page. Next, add the same test for every other system you back up, and keep the quarterly manual drill on the calendar.
