ZFS is a filesystem and volume manager in one: it checksums every block, repairs corrupted data from redundant copies, and gives you cheap snapshots, transparent compression and efficient replication. In this tutorial you will install OpenZFS on Ubuntu 24.04, build a mirrored pool from two disks, organize it into datasets with quotas, schedule automatic snapshots and send them to another server.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS, with a non-root user that has
sudoprivileges. - Two additional empty disks of the same size for the pool. This guide uses
/dev/vdband/dev/vdc. The root disk stays on its current filesystem. - At least 2 GB of RAM. ZFS uses free memory as a read cache (the ARC) and gives it back under pressure.
- For the replication step only: a second Ubuntu 24.04 server with ZFS installed and a pool named
backup, reachable over SSH with key authentication.
Step 1 - Installing ZFS
Ubuntu builds the ZFS kernel module into its standard kernels, so you only need the userland tools:
sudo apt update
sudo apt install zfsutils-linux
Check that the module is loaded and the tools match it:
zfs version
zfs-2.2.2-0ubuntu9
zfs-kmod-2.2.2-0ubuntu9
If you see The ZFS modules are not loaded instead, your kernel flavor does not include ZFS. Install the generic kernel with sudo apt install linux-generic, reboot into it, and run zfs version again.
Step 2 - Identifying the disks
ZFS takes over whole disks, so be sure which ones are free. List the block devices:
lsblk -o NAME,SIZE,TYPE,MOUNTPOINTS
NAME SIZE TYPE MOUNTPOINTS
vda 40G disk
vda1 39G part /
vda15 106M part /boot/efi
vdb 100G disk
vdc 100G disk
vdb and vdc have no partitions and no mount points. On physical servers, prefer the stable names in /dev/disk/by-id/ (for example ata-SAMSUNG_MZ7L3960HCJR_S662NN0W123456), which make it obvious which physical disk to pull when one fails:
ls -l /dev/disk/by-id/
Virtual disks often have no serial number and therefore no by-id entry. In that case /dev/vdb and /dev/vdc are fine: ZFS writes its own labels to each disk and finds them on import even if the device names change.
Step 3 - Creating a mirrored pool
A mirror keeps an identical copy of the data on both disks. It survives the loss of one disk and lets ZFS repair any block that fails its checksum by reading the other copy. Create a pool named tank:
sudo zpool create -o ashift=12 \
-O compression=lz4 -O atime=off -O xattr=sa -O acltype=posixacl \
tank mirror /dev/vdb /dev/vdc
What the options do:
-o ashift=12aligns writes to 4 KiB sectors. It cannot be changed later, and 12 is correct for virtually every modern disk and SSD.-O compression=lz4compresses data transparently. LZ4 is fast enough that it usually speeds up I/O.-O atime=offstops a write on every file read.-O xattr=sa -O acltype=posixaclstore extended attributes efficiently and enable POSIX ACLs, which Linux software expects.
Options set with -O apply to the root dataset and are inherited by every dataset you create below it. Check the pool:
zpool status tank
pool: tank
state: ONLINE
config:
NAME STATE READ WRITE CKSUM
tank ONLINE 0 0 0
mirror-0 ONLINE 0 0 0
vdb ONLINE 0 0 0
vdc ONLINE 0 0 0
errors: No known data errors
The pool is mounted at /tank and is imported and mounted automatically at boot by the zfs-import-cache and zfs-mount services.
NoteOther layouts use the same command with a different keyword:
raidz1,raidz2orraidz3followed by three or more disks tolerate one, two or three failed disks. For a small number of disks, mirrors are simpler to grow and faster to rebuild.
Step 4 - Creating datasets
Datasets are filesystems inside the pool. Each has its own properties, quotas and snapshots, so create one per kind of data instead of storing everything in /tank:
sudo zfs create tank/data
sudo zfs create tank/db
sudo zfs create -o mountpoint=/var/backups/local tank/backups
Set a quota so backups cannot fill the pool, and a record size suited to database pages on the database dataset:
sudo zfs set quota=40G tank/backups
sudo zfs set recordsize=16K tank/db
List the datasets and their mount points:
zfs list -o name,used,avail,quota,mountpoint
NAME USED AVAIL QUOTA MOUNTPOINT
tank 744K 96.4G none /tank
tank/backups 96K 40.0G 40G /var/backups/local
tank/data 96K 96.4G none /tank/data
tank/db 96K 96.4G none /tank/db
The datasets belong to root. Give your user ownership of the data dataset so you can write to it:
sudo chown your_user: /tank/data
Copy some data in and check how well it compresses:
cp -r /usr/share/doc /tank/data/
zfs get compressratio tank/data
NAME PROPERTY VALUE SOURCE
tank/data compressratio 2.87x -
WarningDo not enable deduplication (
dedup=on) unless you have measured a real need. It keeps a table in RAM that grows with the data and can make the pool extremely slow once it no longer fits. Compression gives most of the benefit at almost no cost.
Step 5 - Working with snapshots
A snapshot freezes the state of a dataset instantly and only uses space as the live data changes. Take one before a risky change:
sudo zfs snapshot tank/data@before-cleanup
zfs list -t snapshot
NAME USED AVAIL REFER MOUNTPOINT
tank/data@before-cleanup 0B - 21.4M -
Now simulate a mistake by deleting data:
rm -rf /tank/data/doc
Each dataset has a hidden, read-only .zfs/snapshot directory. Use it to recover individual files without touching anything else:
ls /tank/data/.zfs/snapshot/before-cleanup/doc | head -n 3
To return the whole dataset to the snapshot, roll it back. This discards every change made after the snapshot:
sudo zfs rollback tank/data@before-cleanup
ls /tank/data
doc
Delete a snapshot when you no longer need it:
sudo zfs destroy tank/data@before-cleanup
Step 6 - Automating snapshots with Sanoid
Manual snapshots are easy to forget. Sanoid, packaged in Ubuntu, takes snapshots on a schedule and prunes old ones according to a retention policy:
sudo apt install sanoid
Create its configuration file:
sudo mkdir -p /etc/sanoid
sudo nano /etc/sanoid/sanoid.conf
[tank/data]
use_template = production
[tank/db]
use_template = production
[template_production]
frequently = 0
hourly = 36
daily = 30
monthly = 3
yearly = 0
autosnap = yes
autoprune = yes
This keeps 36 hourly, 30 daily and 3 monthly snapshots of each dataset. The package ships a systemd timer that runs Sanoid every 15 minutes. Make sure it is enabled, then run Sanoid once by hand to create the first snapshots:
sudo systemctl enable --now sanoid.timer
sudo sanoid --cron --verbose
Verify:
zfs list -t snapshot -o name,creation tank/data
NAME CREATION
tank/data@autosnap_2026-09-25_10:15:02_monthly Fri Sep 25 10:15 2026
tank/data@autosnap_2026-09-25_10:15:02_daily Fri Sep 25 10:15 2026
tank/data@autosnap_2026-09-25_10:15:02_hourly Fri Sep 25 10:15 2026
NoteA snapshot of a live database is crash consistent, like pulling the power cord. MySQL/InnoDB and PostgreSQL recover from that state, but for application-consistent backups also keep regular logical dumps in
tank/backups.
Step 7 - Replicating datasets to another server
Snapshots on the same pool do not protect you from losing the server. zfs send turns a snapshot into a stream, and zfs receive rebuilds it on another pool. Because only changed blocks are sent after the first transfer, this is much faster than copying files.
On the backup server, allow your SSH user to receive into the backup pool without root, and make the received datasets read-only so incremental updates always apply cleanly:
sudo zfs allow your_user create,mount,receive backup
sudo zfs set readonly=on backup
On the source server, take a snapshot and send it in full. your_backup_server is the hostname or IP of the backup server:
sudo zfs snapshot tank/data@repl-1
sudo zfs send tank/data@repl-1 | ssh your_user@your_backup_server zfs receive -u backup/data
The -u flag leaves the dataset unmounted, which non-root users cannot do on Linux anyway. Later, send only the changes since the previous snapshot with -i:
sudo zfs snapshot tank/data@repl-2
sudo zfs send -i tank/data@repl-1 tank/data@repl-2 | ssh your_user@your_backup_server zfs receive -u backup/data
Check the result on the backup server:
zfs list -t snapshot -r backup/data
NAME USED AVAIL REFER MOUNTPOINT
backup/data@repl-1 64K - 21.4M -
backup/data@repl-2 0B - 21.5M -
Keep at least the most recent common snapshot on both sides, since the next incremental send needs it. To automate this, the sanoid package also includes syncoid, which finds the common snapshot and runs the incremental send for you.
Step 8 - Limiting ARC memory
By default the ARC can grow to half of the RAM on Linux. On a server that also runs databases or applications, cap it. This example sets the limit to 4 GiB (4 x 1024^3 bytes):
echo "options zfs zfs_arc_max=4294967296" | sudo tee /etc/modprobe.d/zfs.conf
sudo update-initramfs -u
Apply the same value immediately without rebooting:
echo 4294967296 | sudo tee /sys/module/zfs/parameters/zfs_arc_max
Verify the current size and the maximum:
grep -E "^(size|c_max) " /proc/spl/kstat/zfs/arcstats
c_max 4 4294967296
size 4 138502144
c_max is the limit in bytes and size the memory the ARC uses right now. For a full human-readable report, run arc_summary.
Step 9 - Scrubbing and replacing disks
A scrub reads every block in the pool and repairs anything that fails its checksum. Ubuntu's zfsutils-linux package already schedules one on the second Sunday of every month:
cat /etc/cron.d/zfsutils-linux
You can start one manually and follow its progress in zpool status:
sudo zpool scrub tank
zpool status tank | grep -A 2 scan
scan: scrub repaired 0B in 00:00:04 with 0 errors on Fri Sep 25 10:30:11 2026
When a disk fails, zpool status shows it as FAULTED or UNAVAIL and the pool as DEGRADED. Attach a new disk of at least the same size and replace the failed one. Here /dev/vdc failed and /dev/vdd is the new disk:
sudo zpool replace tank /dev/vdc /dev/vdd
ZFS copies the data from the healthy mirror side (called resilvering). The pool returns to ONLINE when zpool status tank reports resilvered ... with 0 errors.
Troubleshooting
invalid vdev specification ... contains a filesystem of type 'ext4'. The disk has old data. After confirming it is the right disk, wipe its signatures with sudo wipefs -a /dev/vdb and create the pool again. Do not use -f to skip the check.
cannot receive incremental stream: destination backup/data has been modified. Something changed the dataset on the backup side. Keep it readonly=on, or add -F to zfs receive to roll it back to the last common snapshot before applying the stream.
The pool is not mounted after a reboot. Check the import and mount services:
systemctl status zfs-import-cache zfs-mount --no-pager
sudo zpool import tank
If zpool import lists the pool as importable, the cache file was missing. Importing it once updates /etc/zfs/zpool.cache.
Conclusion
You installed OpenZFS on Ubuntu 24.04, created a compressed mirrored pool, split it into datasets with quotas, and set up automatic snapshots with Sanoid and incremental replication with zfs send. As next steps, automate replication with syncoid from a systemd timer, enable email alerts from the ZFS event daemon (zed) by setting ZED_EMAIL_ADDR in /etc/zfs/zed.d/zed.rc, and test a full restore from the backup server.
