Bash is still the quickest way to automate routine work on a Linux server: it is installed everywhere and talks directly to the tools you already use. The downside is that careless scripts fail silently, delete the wrong files or break on a filename with a space. In this tutorial you will write three small but production-worthy scripts on Ubuntu 24.04: a disk usage alert, a backup job with rotation, and a user provisioning tool. You will lint them with ShellCheck and run the first two on a schedule with systemd timers.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS, such as a CubePath VPS. The scripts also work unchanged on Debian 12.
- A non-root user with
sudoprivileges. - Basic command-line knowledge: editing files with
nano, file permissions and running commands withsudo.
Step 1 - Installing ShellCheck
ShellCheck is a static analyzer for shell scripts. It catches unquoted variables, broken conditionals and many other bugs before they run on a real server. Install it from the Ubuntu repositories:
sudo apt update
sudo apt install -y shellcheck
Verify it is available:
shellcheck --version
ShellCheck - shell script analysis tool
version: 0.9.0
...
The scripts in this tutorial run as root, so you will install them in /usr/local/sbin, which is on root's PATH and is never touched by the package manager.
Step 2 - Understanding the safe script header
Every script in this guide starts with the same two lines:
#!/usr/bin/env bash
set -euo pipefail
Here is what each option does:
| Option | Effect |
|---|---|
-e | Exit as soon as a command fails, instead of carrying on with bad data |
-u | Treat unset variables as an error, so a typo like $BACKUP_DRI does not expand to an empty string |
-o pipefail | Make a pipeline fail if any command in it fails, not only the last one |
Two more habits matter just as much:
- Quote every variable (
"$dir","${files[@]}"). Unquoted variables are split on spaces and expanded as globs. - Log to the journal with
loggerinstead of inventing your own log files. Messages end up injournalctl, get rotated automatically and can be filtered by tag.
set -e has one well-known trap: a cmd1 && cmd2 list that fails as the last command of a loop or function can make the whole script exit. Use explicit if statements for conditional actions, as the scripts below do.
Step 3 - Writing a disk usage alert
Full disks are one of the most common causes of outages: databases stop writing, logs stop rotating and package upgrades fail. This script checks every real filesystem and reports any that are above a threshold.
Create the script:
sudo nano /usr/local/sbin/disk-alert
#!/usr/bin/env bash
# disk-alert: warn when a filesystem is above a usage threshold.
# Usage: disk-alert [THRESHOLD_PERCENT]
set -euo pipefail
THRESHOLD="${1:-85}"
if [[ ! "$THRESHOLD" =~ ^[0-9]+$ ]] || (( THRESHOLD > 100 )); then
echo "Usage: $(basename "$0") [THRESHOLD_PERCENT]" >&2
exit 2
fi
status=0
while read -r pcent target; do
usage="${pcent%\%}"
# Skip filesystems that report no percentage
if [[ ! "$usage" =~ ^[0-9]+$ ]]; then
continue
fi
if (( usage >= THRESHOLD )); then
message="${target} is at ${usage}% (threshold ${THRESHOLD}%)"
echo "WARNING: ${message}"
logger -t disk-alert -p user.warning "$message"
status=1
fi
done < <(df --output=pcent,target -x tmpfs -x devtmpfs -x squashfs -x overlay | tail -n +2)
exit "$status"
A few details make this reliable:
df --output=pcent,targetprints only the two columns you need, so there is no fragileawkcolumn counting.-xexcludes pseudo filesystems such astmpfsand snapsquashfsmounts, which are always at 100% and would trigger false alarms.- The script exits with status
1when any filesystem is over the limit. systemd marks that run as failed, which is exactly what you want to notice.
Make it executable and lint it:
sudo chmod 755 /usr/local/sbin/disk-alert
shellcheck /usr/local/sbin/disk-alert
ShellCheck prints nothing when it finds no problems. Now test the script with a threshold of 1% so that it has something to report:
sudo disk-alert 1; echo "exit code: $?"
WARNING: / is at 23% (threshold 1%)
WARNING: /boot is at 14% (threshold 1%)
exit code: 1
Confirm the warnings reached the journal:
journalctl -t disk-alert --since "5 minutes ago"
Sep 25 10:12:04 web01 disk-alert[4121]: / is at 23% (threshold 1%)
Sep 25 10:12:04 web01 disk-alert[4121]: /boot is at 14% (threshold 1%)
Step 4 - Writing a backup script with rotation
This script archives important directories into a timestamped .tar.gz file, checks that the archive is readable and deletes archives older than a retention period.
sudo nano /usr/local/sbin/backup-local
#!/usr/bin/env bash
# backup-local: archive key directories and keep the last KEEP_DAYS days.
set -euo pipefail
BACKUP_DIR="/var/backups/local"
SOURCES=(/etc /home /var/www)
KEEP_DAYS=7
umask 077
mkdir -p "$BACKUP_DIR"
# Only archive sources that exist, as paths relative to /
paths=()
for src in "${SOURCES[@]}"; do
if [[ -e "$src" ]]; then
paths+=("${src#/}")
fi
done
if (( ${#paths[@]} == 0 )); then
logger -t backup-local -p user.err "No source directories found"
exit 1
fi
archive="${BACKUP_DIR}/backup-$(hostname -s)-$(date +%Y%m%d-%H%M%S).tar.gz"
# GNU tar exits with 1 when a file changed while being read (for example
# an active log). That is acceptable; anything higher is a real error.
rc=0
tar -czf "$archive" -C / "${paths[@]}" || rc=$?
if (( rc > 1 )); then
rm -f -- "$archive"
logger -t backup-local -p user.err "tar failed with exit code ${rc}"
exit "$rc"
fi
# Make sure the archive can be read back
tar -tzf "$archive" > /dev/null
# Delete archives older than KEEP_DAYS
find "$BACKUP_DIR" -maxdepth 1 -type f -name 'backup-*.tar.gz' -mtime +"$KEEP_DAYS" -delete
size="$(du -h "$archive" | cut -f1)"
logger -t backup-local "Created ${archive} (${size})"
echo "Created ${archive} (${size})"
Why it is written this way:
umask 077makes the archives readable only by root. They contain/etc/shadowand private keys.tar -C /with relative paths avoids the "Removing leading/" warning and makes restores predictable.- Rotation uses
find -deletewith an exact name pattern, so it can never remove anything outsidebackup-*.tar.gzin that directory.
Make it executable, lint it and run it once:
sudo chmod 755 /usr/local/sbin/backup-local
shellcheck /usr/local/sbin/backup-local
sudo backup-local
Created /var/backups/local/backup-web01-20260925-101530.tar.gz (4.2M)
List the first entries of the archive, using the file name the script printed:
sudo tar -tzf /var/backups/local/backup-web01-20260925-101530.tar.gz | head -n 5
etc/
etc/hostname
etc/fstab
etc/ssh/
etc/ssh/sshd_config
Importanta backup stored on the same disk does not survive a disk failure or a deleted server. Copy these archives to another machine or to object storage, for example with
rsyncorrclone.
Step 5 - Scheduling the scripts with systemd timers
systemd timers are the modern replacement for cron on Ubuntu. Every run is logged in the journal, failed runs are visible in systemctl --failed, and Persistent=true catches up on runs missed while the server was off.
Create a service unit for the backup:
sudo nano /etc/systemd/system/backup-local.service
[Unit]
Description=Local backup of /etc, /home and /var/www
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/backup-local
Nice=10
IOSchedulingClass=idle
Create a timer that runs it every night at 02:30:
sudo nano /etc/systemd/system/backup-local.timer
[Unit]
Description=Nightly local backup
[Timer]
OnCalendar=*-*-* 02:30:00
RandomizedDelaySec=15min
Persistent=true
[Install]
WantedBy=timers.target
Do the same for the disk alert, running every 15 minutes:
sudo nano /etc/systemd/system/disk-alert.service
[Unit]
Description=Check filesystem usage
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/disk-alert 85
sudo nano /etc/systemd/system/disk-alert.timer
[Unit]
Description=Check filesystem usage every 15 minutes
[Timer]
OnCalendar=*:0/15
Persistent=true
[Install]
WantedBy=timers.target
Reload systemd and enable both timers:
sudo systemctl daemon-reload
sudo systemctl enable --now backup-local.timer disk-alert.timer
Check when they will run next:
systemctl list-timers backup-local.timer disk-alert.timer
NEXT LEFT LAST PASSED UNIT ACTIVATES
Thu 2026-09-25 10:30:00 UTC 12min - - disk-alert.timer disk-alert.service
Fri 2026-09-26 02:38:11 UTC 16h - - backup-local.timer backup-local.service
2 timers listed.
You can trigger a run immediately and read its output:
sudo systemctl start backup-local.service
journalctl -u backup-local.service -n 5 --no-pager
Step 6 - Writing a user provisioning script
Creating users by hand tends to leave .ssh directories with the wrong owner or permissions, which makes SSH silently ignore the key. This script creates a user, installs one SSH public key with the correct permissions and optionally adds the user to the sudo group. It validates both inputs before changing anything.
sudo nano /usr/local/sbin/add-ssh-user
#!/usr/bin/env bash
# add-ssh-user: create a user with an SSH public key.
# Usage: add-ssh-user USERNAME "PUBLIC_KEY" [--sudo]
set -euo pipefail
die() {
echo "Error: $*" >&2
exit 1
}
if [[ $EUID -ne 0 ]]; then
die "run this script as root (with sudo)"
fi
if (( $# < 2 )); then
echo "Usage: $(basename "$0") USERNAME \"PUBLIC_KEY\" [--sudo]" >&2
exit 2
fi
username="$1"
pubkey="$2"
grant_sudo="${3:-}"
if [[ ! "$username" =~ ^[a-z_][a-z0-9_-]{0,31}$ ]]; then
die "invalid username: ${username}"
fi
if id "$username" &> /dev/null; then
die "user ${username} already exists"
fi
tmpkey="$(mktemp)"
trap 'rm -f -- "$tmpkey"' EXIT
printf '%s\n' "$pubkey" > "$tmpkey"
if ! ssh-keygen -l -f "$tmpkey" > /dev/null 2>&1; then
die "the public key is not valid"
fi
useradd --create-home --shell /bin/bash "$username"
install -d -m 700 -o "$username" -g "$username" "/home/${username}/.ssh"
install -m 600 -o "$username" -g "$username" "$tmpkey" "/home/${username}/.ssh/authorized_keys"
if [[ "$grant_sudo" == "--sudo" ]]; then
usermod -aG sudo "$username"
fi
logger -t add-ssh-user "Created user ${username} (sudo: ${grant_sudo:-no})"
echo "Created user ${username}"
ssh-keygen -l only succeeds on a well-formed public key, so a truncated copy and paste is rejected before the account exists. install creates the directory and file with the right owner and mode in a single step.
Make it executable, lint it and create a test user. Replace the key with a real public key, such as the content of ~/.ssh/id_ed25519.pub on your workstation:
sudo chmod 755 /usr/local/sbin/add-ssh-user
shellcheck /usr/local/sbin/add-ssh-user
sudo add-ssh-user deploy "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAI... you@workstation"
Created user deploy
Check the permissions:
sudo ls -la /home/deploy/.ssh
drwx------ 2 deploy deploy 4096 Sep 25 10:40 .
drwxr-x--- 3 deploy deploy 4096 Sep 25 10:40 ..
-rw------- 1 deploy deploy 97 Sep 25 10:40 authorized_keys
Finally, test the login from your workstation with ssh deploy@your_server_ip. The new account has no password, so SSH keys are the only way in. If you used --sudo, set a password with sudo passwd deploy so the user can authenticate to sudo.
Troubleshooting
- A script works in your shell but fails under systemd: systemd runs it with a minimal environment and no terminal. Use absolute paths for anything outside
/usr/binand/usr/sbin, and read the error withjournalctl -u NAME.service. unbound variableerrors:set -uis doing its job. Give optional variables a default, as in"${1:-85}".$'\r': command not found: the file was saved with Windows line endings. Convert it withsed -i 's/\r$//' FILE.- The timer never fires: check that you enabled the
.timerunit, not the.service, and thatsystemctl list-timersshows it.
Conclusion
You wrote three Bash scripts that fail loudly instead of silently, log to the journal, pass ShellCheck and run on a schedule with systemd timers. The same header, quoting rules and validation patterns apply to any other script you add to /usr/local/sbin.
From here, you could send the backup archives off the server with rclone, add an OnFailure= unit that notifies you when a timer run fails, or keep your scripts in a Git repository and deploy them to every server with Ansible.
