ZFS is a filesystem and volume manager in one: it checksums every block, repairs corrupted data from redundant copies, and gives you cheap snapshots, transparent compression and efficient replication. In this tutorial you will install OpenZFS on Ubuntu 24.04, build a mirrored pool from two disks, organize it into datasets with quotas, schedule automatic snapshots and send them to another server.

Prerequisites

To follow this guide you need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS, with a non-root user that has sudo privileges.
  • Two additional empty disks of the same size for the pool. This guide uses /dev/vdb and /dev/vdc. The root disk stays on its current filesystem.
  • At least 2 GB of RAM. ZFS uses free memory as a read cache (the ARC) and gives it back under pressure.
  • For the replication step only: a second Ubuntu 24.04 server with ZFS installed and a pool named backup, reachable over SSH with key authentication.

Step 1 - Installing ZFS

Ubuntu builds the ZFS kernel module into its standard kernels, so you only need the userland tools:

sudo apt update
sudo apt install zfsutils-linux

Check that the module is loaded and the tools match it:

zfs version
zfs-2.2.2-0ubuntu9
zfs-kmod-2.2.2-0ubuntu9

If you see The ZFS modules are not loaded instead, your kernel flavor does not include ZFS. Install the generic kernel with sudo apt install linux-generic, reboot into it, and run zfs version again.

Step 2 - Identifying the disks

ZFS takes over whole disks, so be sure which ones are free. List the block devices:

lsblk -o NAME,SIZE,TYPE,MOUNTPOINTS
NAME    SIZE TYPE MOUNTPOINTS
vda      40G disk
vda1     39G part /
vda15   106M part /boot/efi
vdb     100G disk
vdc     100G disk

vdb and vdc have no partitions and no mount points. On physical servers, prefer the stable names in /dev/disk/by-id/ (for example ata-SAMSUNG_MZ7L3960HCJR_S662NN0W123456), which make it obvious which physical disk to pull when one fails:

ls -l /dev/disk/by-id/

Virtual disks often have no serial number and therefore no by-id entry. In that case /dev/vdb and /dev/vdc are fine: ZFS writes its own labels to each disk and finds them on import even if the device names change.

Step 3 - Creating a mirrored pool

A mirror keeps an identical copy of the data on both disks. It survives the loss of one disk and lets ZFS repair any block that fails its checksum by reading the other copy. Create a pool named tank:

sudo zpool create -o ashift=12 \
  -O compression=lz4 -O atime=off -O xattr=sa -O acltype=posixacl \
  tank mirror /dev/vdb /dev/vdc

What the options do:

  • -o ashift=12 aligns writes to 4 KiB sectors. It cannot be changed later, and 12 is correct for virtually every modern disk and SSD.
  • -O compression=lz4 compresses data transparently. LZ4 is fast enough that it usually speeds up I/O.
  • -O atime=off stops a write on every file read.
  • -O xattr=sa -O acltype=posixacl store extended attributes efficiently and enable POSIX ACLs, which Linux software expects.

Options set with -O apply to the root dataset and are inherited by every dataset you create below it. Check the pool:

zpool status tank
  pool: tank
 state: ONLINE
config:

        NAME        STATE     READ WRITE CKSUM
        tank        ONLINE       0     0     0
          mirror-0  ONLINE       0     0     0
            vdb     ONLINE       0     0     0
            vdc     ONLINE       0     0     0

errors: No known data errors

The pool is mounted at /tank and is imported and mounted automatically at boot by the zfs-import-cache and zfs-mount services.

Step 4 - Creating datasets

Datasets are filesystems inside the pool. Each has its own properties, quotas and snapshots, so create one per kind of data instead of storing everything in /tank:

sudo zfs create tank/data
sudo zfs create tank/db
sudo zfs create -o mountpoint=/var/backups/local tank/backups

Set a quota so backups cannot fill the pool, and a record size suited to database pages on the database dataset:

sudo zfs set quota=40G tank/backups
sudo zfs set recordsize=16K tank/db

List the datasets and their mount points:

zfs list -o name,used,avail,quota,mountpoint
NAME          USED  AVAIL  QUOTA  MOUNTPOINT
tank          744K  96.4G   none  /tank
tank/backups   96K  40.0G    40G  /var/backups/local
tank/data      96K  96.4G   none  /tank/data
tank/db        96K  96.4G   none  /tank/db

The datasets belong to root. Give your user ownership of the data dataset so you can write to it:

sudo chown your_user: /tank/data

Copy some data in and check how well it compresses:

cp -r /usr/share/doc /tank/data/
zfs get compressratio tank/data
NAME       PROPERTY       VALUE  SOURCE
tank/data  compressratio  2.87x  -

Step 5 - Working with snapshots

A snapshot freezes the state of a dataset instantly and only uses space as the live data changes. Take one before a risky change:

sudo zfs snapshot tank/data@before-cleanup
zfs list -t snapshot
NAME                       USED  AVAIL  REFER  MOUNTPOINT
tank/data@before-cleanup     0B      -  21.4M  -

Now simulate a mistake by deleting data:

rm -rf /tank/data/doc

Each dataset has a hidden, read-only .zfs/snapshot directory. Use it to recover individual files without touching anything else:

ls /tank/data/.zfs/snapshot/before-cleanup/doc | head -n 3

To return the whole dataset to the snapshot, roll it back. This discards every change made after the snapshot:

sudo zfs rollback tank/data@before-cleanup
ls /tank/data
doc

Delete a snapshot when you no longer need it:

sudo zfs destroy tank/data@before-cleanup

Step 6 - Automating snapshots with Sanoid

Manual snapshots are easy to forget. Sanoid, packaged in Ubuntu, takes snapshots on a schedule and prunes old ones according to a retention policy:

sudo apt install sanoid

Create its configuration file:

sudo mkdir -p /etc/sanoid
sudo nano /etc/sanoid/sanoid.conf
[tank/data]
        use_template = production

[tank/db]
        use_template = production

[template_production]
        frequently = 0
        hourly = 36
        daily = 30
        monthly = 3
        yearly = 0
        autosnap = yes
        autoprune = yes

This keeps 36 hourly, 30 daily and 3 monthly snapshots of each dataset. The package ships a systemd timer that runs Sanoid every 15 minutes. Make sure it is enabled, then run Sanoid once by hand to create the first snapshots:

sudo systemctl enable --now sanoid.timer
sudo sanoid --cron --verbose

Verify:

zfs list -t snapshot -o name,creation tank/data
NAME                                             CREATION
tank/data@autosnap_2026-09-25_10:15:02_monthly   Fri Sep 25 10:15 2026
tank/data@autosnap_2026-09-25_10:15:02_daily     Fri Sep 25 10:15 2026
tank/data@autosnap_2026-09-25_10:15:02_hourly    Fri Sep 25 10:15 2026

Step 7 - Replicating datasets to another server

Snapshots on the same pool do not protect you from losing the server. zfs send turns a snapshot into a stream, and zfs receive rebuilds it on another pool. Because only changed blocks are sent after the first transfer, this is much faster than copying files.

On the backup server, allow your SSH user to receive into the backup pool without root, and make the received datasets read-only so incremental updates always apply cleanly:

sudo zfs allow your_user create,mount,receive backup
sudo zfs set readonly=on backup

On the source server, take a snapshot and send it in full. your_backup_server is the hostname or IP of the backup server:

sudo zfs snapshot tank/data@repl-1
sudo zfs send tank/data@repl-1 | ssh your_user@your_backup_server zfs receive -u backup/data

The -u flag leaves the dataset unmounted, which non-root users cannot do on Linux anyway. Later, send only the changes since the previous snapshot with -i:

sudo zfs snapshot tank/data@repl-2
sudo zfs send -i tank/data@repl-1 tank/data@repl-2 | ssh your_user@your_backup_server zfs receive -u backup/data

Check the result on the backup server:

zfs list -t snapshot -r backup/data
NAME                 USED  AVAIL  REFER  MOUNTPOINT
backup/data@repl-1    64K      -  21.4M  -
backup/data@repl-2     0B      -  21.5M  -

Keep at least the most recent common snapshot on both sides, since the next incremental send needs it. To automate this, the sanoid package also includes syncoid, which finds the common snapshot and runs the incremental send for you.

Step 8 - Limiting ARC memory

By default the ARC can grow to half of the RAM on Linux. On a server that also runs databases or applications, cap it. This example sets the limit to 4 GiB (4 x 1024^3 bytes):

echo "options zfs zfs_arc_max=4294967296" | sudo tee /etc/modprobe.d/zfs.conf
sudo update-initramfs -u

Apply the same value immediately without rebooting:

echo 4294967296 | sudo tee /sys/module/zfs/parameters/zfs_arc_max

Verify the current size and the maximum:

grep -E "^(size|c_max) " /proc/spl/kstat/zfs/arcstats
c_max                           4    4294967296
size                            4    138502144

c_max is the limit in bytes and size the memory the ARC uses right now. For a full human-readable report, run arc_summary.

Step 9 - Scrubbing and replacing disks

A scrub reads every block in the pool and repairs anything that fails its checksum. Ubuntu's zfsutils-linux package already schedules one on the second Sunday of every month:

cat /etc/cron.d/zfsutils-linux

You can start one manually and follow its progress in zpool status:

sudo zpool scrub tank
zpool status tank | grep -A 2 scan
  scan: scrub repaired 0B in 00:00:04 with 0 errors on Fri Sep 25 10:30:11 2026

When a disk fails, zpool status shows it as FAULTED or UNAVAIL and the pool as DEGRADED. Attach a new disk of at least the same size and replace the failed one. Here /dev/vdc failed and /dev/vdd is the new disk:

sudo zpool replace tank /dev/vdc /dev/vdd

ZFS copies the data from the healthy mirror side (called resilvering). The pool returns to ONLINE when zpool status tank reports resilvered ... with 0 errors.

Troubleshooting

invalid vdev specification ... contains a filesystem of type 'ext4'. The disk has old data. After confirming it is the right disk, wipe its signatures with sudo wipefs -a /dev/vdb and create the pool again. Do not use -f to skip the check.

cannot receive incremental stream: destination backup/data has been modified. Something changed the dataset on the backup side. Keep it readonly=on, or add -F to zfs receive to roll it back to the last common snapshot before applying the stream.

The pool is not mounted after a reboot. Check the import and mount services:

systemctl status zfs-import-cache zfs-mount --no-pager
sudo zpool import tank

If zpool import lists the pool as importable, the cache file was missing. Importing it once updates /etc/zfs/zpool.cache.

Conclusion

You installed OpenZFS on Ubuntu 24.04, created a compressed mirrored pool, split it into datasets with quotas, and set up automatic snapshots with Sanoid and incremental replication with zfs send. As next steps, automate replication with syncoid from a systemd timer, enable email alerts from the ZFS event daemon (zed) by setting ZED_EMAIL_ADDR in /etc/zfs/zed.d/zed.rc, and test a full restore from the backup server.