The I/O scheduler decides in which order the Linux kernel sends read and write requests to a block device. On modern NVMe drives the best choice is usually no scheduler at all, while slower SATA SSDs and spinning disks can benefit from reordering and fairness. In this tutorial you will inspect the scheduler of every disk on Ubuntu 24.04, switch it at runtime, make the choice persistent with a udev rule, tune the most useful parameters and measure the result with fio.
Prerequisites
- A server running Ubuntu 24.04 LTS (the commands also work on Debian 12 and Rocky Linux 9, which use the same multi-queue block layer).
- A non-root user with
sudoprivileges. - A disk or partition with a few GB of free space for benchmark files.
NoteOn a virtual machine, such as a CubePath VPS, the guest sees a virtual disk (
vdaorsda) and the hypervisor schedules the real hardware. Inside the guest,noneormq-deadlineis almost always the right answer. Most of the gains in this guide apply to dedicated servers with physical drives.
Step 1 - Understanding the available schedulers
Since kernel 5.0 Linux only uses the multi-queue block layer (blk-mq). The legacy schedulers noop, deadline and cfq no longer exist, and the old elevator= boot parameter has no effect. Current kernels offer four options:
| Scheduler | How it works | Good fit |
|---|---|---|
none | Passes requests straight to the device, no reordering | NVMe, fast SSD arrays, most virtual disks |
mq-deadline | Sorts requests and gives each one an expiry time, reads are favored | SATA SSDs, virtual disks, database servers on SSD |
bfq | Budget Fair Queueing, divides bandwidth fairly between processes | Spinning disks, desktops, mixed interactive workloads |
kyber | Throttles queues to hit target read/write latencies | Fast devices where you want latency control with little CPU cost |
By default the kernel picks none for devices with multiple hardware queues (NVMe, multi-queue virtio) and mq-deadline for the rest.
Step 2 - Checking the current scheduler
List every disk with its type and active scheduler:
lsblk -d -o NAME,ROTA,TYPE,SIZE,SCHED
NAME ROTA TYPE SIZE SCHED
sda 1 disk 3.6T mq-deadline
sdb 0 disk 960G mq-deadline
nvme0n1 0 disk 1.8T none
ROTA is 1 for rotational disks (HDD) and 0 for SSD and NVMe. To see every scheduler a device supports, read its scheduler file. The active one is shown in brackets:
cat /sys/block/sda/queue/scheduler
[mq-deadline] none
If bfq or kyber does not appear, their kernel modules are not loaded yet. On Ubuntu they ship as modules; load them with:
sudo modprobe bfq
sudo modprobe kyber-iosched
Check again and both should now be listed:
[mq-deadline] kyber bfq none
Step 3 - Changing the scheduler at runtime
Write the scheduler name to the device's scheduler file. Replace sda with your device:
echo bfq | sudo tee /sys/block/sda/queue/scheduler
Verify the change:
cat /sys/block/sda/queue/scheduler
mq-deadline kyber [bfq] none
The change takes effect immediately for new requests, with no reboot or unmount, but it is lost on the next boot. Test first, then make it permanent in the next step.
Step 4 - Making the scheduler persistent with udev
A udev rule applies the scheduler every time the kernel detects the device, including at boot and after hotplug. It can also match on the rotational attribute, so the same rule works across servers with different hardware.
If you plan to use bfq or kyber, have the module loaded at boot first:
echo bfq | sudo tee /etc/modules-load.d/bfq.conf
Create the rule file:
sudo nano /etc/udev/rules.d/60-io-scheduler.rules
Add the following rules, which use BFQ for HDDs, mq-deadline for SATA SSDs and no scheduler for NVMe:
# Spinning disks
ACTION=="add|change", KERNEL=="sd[a-z]*", ATTR{queue/rotational}=="1", ATTR{queue/scheduler}="bfq"
# SATA/SAS SSDs
ACTION=="add|change", KERNEL=="sd[a-z]*", ATTR{queue/rotational}=="0", ATTR{queue/scheduler}="mq-deadline"
# NVMe namespaces
ACTION=="add|change", KERNEL=="nvme[0-9]*n[0-9]*", ATTR{queue/scheduler}="none"
On a VPS with virtio disks, use KERNEL=="vd[a-z]*" instead of sd[a-z]*.
Reload the rules and apply them to existing devices without rebooting:
sudo udevadm control --reload
sudo udevadm trigger --subsystem-match=block --action=change
Confirm the result:
lsblk -d -o NAME,ROTA,SCHED
NAME ROTA SCHED
sda 1 bfq
sdb 0 mq-deadline
nvme0n1 0 none
Reboot once during a maintenance window to confirm the rule is also applied at boot.
Step 5 - Tuning scheduler parameters
Each scheduler exposes its tunables under /sys/block/<device>/queue/iosched/. The defaults are sensible, so only change a value when a benchmark shows a benefit. Runtime changes, like the scheduler itself, can be made persistent by adding more ATTR{...} assignments to the udev rule.
mq-deadline
List the parameters and their values:
grep . /sys/block/sdb/queue/iosched/*
/sys/block/sdb/queue/iosched/async_depth:48
/sys/block/sdb/queue/iosched/fifo_batch:16
/sys/block/sdb/queue/iosched/front_merges:1
/sys/block/sdb/queue/iosched/prio_aging_expire:10000
/sys/block/sdb/queue/iosched/read_expire:500
/sys/block/sdb/queue/iosched/write_expire:5000
/sys/block/sdb/queue/iosched/writes_starved:2
The most relevant ones are:
read_expireandwrite_expire: time in milliseconds after which a request must be served. Loweringread_expire(for example to250) tightens read latency for databases.fifo_batch: how many requests are dispatched in one batch. Lower values reduce latency, higher values increase throughput.writes_starved: how many read batches can go before a pending write batch is served.
For example, to favor read latency on a database SSD:
echo 250 | sudo tee /sys/block/sdb/queue/iosched/read_expire
echo 8 | sudo tee /sys/block/sdb/queue/iosched/fifo_batch
BFQ
BFQ has more options, but two matter most:
low_latency(default1): boosts interactive and soft real-time applications. Set it to0on a pure throughput server such as a backup target.slice_idle(default8ms): how long BFQ waits for the next request from the same process. It helps sequential reads on HDDs;0improves throughput on SSDs and RAID controllers with their own cache.
echo 0 | sudo tee /sys/block/sda/queue/iosched/low_latency
BFQ also honors I/O priorities set with ionice, which is useful to push a backup job to the background:
sudo ionice -c 3 tar -czf /backup/home.tar.gz /home
Kyber
Kyber has only two targets, in nanoseconds: read_lat_nsec (default 2 ms) and write_lat_nsec (default 10 ms). Lower them to make Kyber throttle more aggressively.
Queue settings common to all schedulers
These live in /sys/block/<device>/queue/ and apply regardless of the scheduler:
read_ahead_kb: how much data is read ahead on sequential access. Increase it (for example to1024) for streaming large files from HDDs, keep it low for random database I/O.nr_requests: queue depth managed by the scheduler.rotational: should match the hardware. Some RAID controllers report SSD arrays as rotational; set it to0so the kernel and your udev rules treat them correctly.
Step 6 - Benchmarking with fio
Never change schedulers based on theory alone. Install fio:
sudo apt update
sudo apt install fio
Run a random read test at 4 KiB, which approximates a database workload. It creates a 4 GB file in the current directory, so run it on the disk you are testing:
fio --name=randread --filename=fio-test --size=4G --rw=randread --bs=4k \
--ioengine=libaio --direct=1 --iodepth=32 --numjobs=4 --group_reporting \
--runtime=60 --time_based
Look at the IOPS figure and the latency percentiles in the output:
read: IOPS=94.2k, BW=368MiB/s (386MB/s)(21.6GiB/60001msec)
...
| 99.00th=[ 2769], 99.50th=[ 3392], 99.90th=[ 5145], 99.95th=[ 6128],
Switch the scheduler as shown in Step 3, run the same command again and compare. Repeat with a mixed workload (--rw=randrw --rwmixread=70) or a sequential one (--rw=read --bs=1M) that matches what the server really does. Remove the test file when you finish:
rm fio-test
Step 7 - Monitoring I/O in production
The iostat tool from the sysstat package shows per-device latency and utilization:
sudo apt install sysstat
iostat -xz 2
Focus on these columns:
r_awaitandw_await: average time in milliseconds a read or write takes, including queueing. This is the latency your applications feel.aqu-sz: average queue size. A persistently high value means the device cannot keep up.%util: time the device was busy. On NVMe and RAID it can reach 100% while still having spare capacity, so trustawaitmore.
To find which process generates the I/O, use iotop:
sudo apt install iotop-c
sudo iotop-c -o
Troubleshooting
tee: /sys/block/sda/queue/scheduler: Invalid argument: the scheduler is not available. Load its module withmodprobe(Step 2) and check the spelling.- The udev rule does not apply after reboot: make sure the module is listed in
/etc/modules-load.d/and that theKERNEL==pattern matches your device name (sd*,vd*ornvme*n*). Test withsudo udevadm test /sys/block/sda. - No difference in benchmarks: expected on NVMe and virtual disks, where the device or the hypervisor does its own scheduling. Keep
nonein that case, since it has the lowest CPU overhead.
Conclusion
You now know which scheduler each of your disks uses, how to switch it safely, how to make the choice persistent with a udev rule and how to prove the benefit with fio. As a rule of thumb, use none on NVMe and virtual disks, mq-deadline on SATA SSDs and bfq on spinning disks. As next steps, review memory and swap tuning to reduce unnecessary disk I/O, and set up monitoring of await latency so you notice when a disk starts to saturate.
