The I/O scheduler decides in which order the Linux kernel sends read and write requests to a block device. On modern NVMe drives the best choice is usually no scheduler at all, while slower SATA SSDs and spinning disks can benefit from reordering and fairness. In this tutorial you will inspect the scheduler of every disk on Ubuntu 24.04, switch it at runtime, make the choice persistent with a udev rule, tune the most useful parameters and measure the result with fio.

Prerequisites

  • A server running Ubuntu 24.04 LTS (the commands also work on Debian 12 and Rocky Linux 9, which use the same multi-queue block layer).
  • A non-root user with sudo privileges.
  • A disk or partition with a few GB of free space for benchmark files.

Step 1 - Understanding the available schedulers

Since kernel 5.0 Linux only uses the multi-queue block layer (blk-mq). The legacy schedulers noop, deadline and cfq no longer exist, and the old elevator= boot parameter has no effect. Current kernels offer four options:

SchedulerHow it worksGood fit
nonePasses requests straight to the device, no reorderingNVMe, fast SSD arrays, most virtual disks
mq-deadlineSorts requests and gives each one an expiry time, reads are favoredSATA SSDs, virtual disks, database servers on SSD
bfqBudget Fair Queueing, divides bandwidth fairly between processesSpinning disks, desktops, mixed interactive workloads
kyberThrottles queues to hit target read/write latenciesFast devices where you want latency control with little CPU cost

By default the kernel picks none for devices with multiple hardware queues (NVMe, multi-queue virtio) and mq-deadline for the rest.

Step 2 - Checking the current scheduler

List every disk with its type and active scheduler:

lsblk -d -o NAME,ROTA,TYPE,SIZE,SCHED
NAME    ROTA TYPE  SIZE SCHED
sda        1 disk  3.6T mq-deadline
sdb        0 disk  960G mq-deadline
nvme0n1    0 disk  1.8T none

ROTA is 1 for rotational disks (HDD) and 0 for SSD and NVMe. To see every scheduler a device supports, read its scheduler file. The active one is shown in brackets:

cat /sys/block/sda/queue/scheduler
[mq-deadline] none

If bfq or kyber does not appear, their kernel modules are not loaded yet. On Ubuntu they ship as modules; load them with:

sudo modprobe bfq
sudo modprobe kyber-iosched

Check again and both should now be listed:

[mq-deadline] kyber bfq none

Step 3 - Changing the scheduler at runtime

Write the scheduler name to the device's scheduler file. Replace sda with your device:

echo bfq | sudo tee /sys/block/sda/queue/scheduler

Verify the change:

cat /sys/block/sda/queue/scheduler
mq-deadline kyber [bfq] none

The change takes effect immediately for new requests, with no reboot or unmount, but it is lost on the next boot. Test first, then make it permanent in the next step.

Step 4 - Making the scheduler persistent with udev

A udev rule applies the scheduler every time the kernel detects the device, including at boot and after hotplug. It can also match on the rotational attribute, so the same rule works across servers with different hardware.

If you plan to use bfq or kyber, have the module loaded at boot first:

echo bfq | sudo tee /etc/modules-load.d/bfq.conf

Create the rule file:

sudo nano /etc/udev/rules.d/60-io-scheduler.rules

Add the following rules, which use BFQ for HDDs, mq-deadline for SATA SSDs and no scheduler for NVMe:

# Spinning disks
ACTION=="add|change", KERNEL=="sd[a-z]*", ATTR{queue/rotational}=="1", ATTR{queue/scheduler}="bfq"
# SATA/SAS SSDs
ACTION=="add|change", KERNEL=="sd[a-z]*", ATTR{queue/rotational}=="0", ATTR{queue/scheduler}="mq-deadline"
# NVMe namespaces
ACTION=="add|change", KERNEL=="nvme[0-9]*n[0-9]*", ATTR{queue/scheduler}="none"

On a VPS with virtio disks, use KERNEL=="vd[a-z]*" instead of sd[a-z]*.

Reload the rules and apply them to existing devices without rebooting:

sudo udevadm control --reload
sudo udevadm trigger --subsystem-match=block --action=change

Confirm the result:

lsblk -d -o NAME,ROTA,SCHED
NAME    ROTA SCHED
sda        1 bfq
sdb        0 mq-deadline
nvme0n1    0 none

Reboot once during a maintenance window to confirm the rule is also applied at boot.

Step 5 - Tuning scheduler parameters

Each scheduler exposes its tunables under /sys/block/<device>/queue/iosched/. The defaults are sensible, so only change a value when a benchmark shows a benefit. Runtime changes, like the scheduler itself, can be made persistent by adding more ATTR{...} assignments to the udev rule.

mq-deadline

List the parameters and their values:

grep . /sys/block/sdb/queue/iosched/*
/sys/block/sdb/queue/iosched/async_depth:48
/sys/block/sdb/queue/iosched/fifo_batch:16
/sys/block/sdb/queue/iosched/front_merges:1
/sys/block/sdb/queue/iosched/prio_aging_expire:10000
/sys/block/sdb/queue/iosched/read_expire:500
/sys/block/sdb/queue/iosched/write_expire:5000
/sys/block/sdb/queue/iosched/writes_starved:2

The most relevant ones are:

  • read_expire and write_expire: time in milliseconds after which a request must be served. Lowering read_expire (for example to 250) tightens read latency for databases.
  • fifo_batch: how many requests are dispatched in one batch. Lower values reduce latency, higher values increase throughput.
  • writes_starved: how many read batches can go before a pending write batch is served.

For example, to favor read latency on a database SSD:

echo 250 | sudo tee /sys/block/sdb/queue/iosched/read_expire
echo 8 | sudo tee /sys/block/sdb/queue/iosched/fifo_batch

BFQ

BFQ has more options, but two matter most:

  • low_latency (default 1): boosts interactive and soft real-time applications. Set it to 0 on a pure throughput server such as a backup target.
  • slice_idle (default 8 ms): how long BFQ waits for the next request from the same process. It helps sequential reads on HDDs; 0 improves throughput on SSDs and RAID controllers with their own cache.
echo 0 | sudo tee /sys/block/sda/queue/iosched/low_latency

BFQ also honors I/O priorities set with ionice, which is useful to push a backup job to the background:

sudo ionice -c 3 tar -czf /backup/home.tar.gz /home

Kyber

Kyber has only two targets, in nanoseconds: read_lat_nsec (default 2 ms) and write_lat_nsec (default 10 ms). Lower them to make Kyber throttle more aggressively.

Queue settings common to all schedulers

These live in /sys/block/<device>/queue/ and apply regardless of the scheduler:

  • read_ahead_kb: how much data is read ahead on sequential access. Increase it (for example to 1024) for streaming large files from HDDs, keep it low for random database I/O.
  • nr_requests: queue depth managed by the scheduler.
  • rotational: should match the hardware. Some RAID controllers report SSD arrays as rotational; set it to 0 so the kernel and your udev rules treat them correctly.

Step 6 - Benchmarking with fio

Never change schedulers based on theory alone. Install fio:

sudo apt update
sudo apt install fio

Run a random read test at 4 KiB, which approximates a database workload. It creates a 4 GB file in the current directory, so run it on the disk you are testing:

fio --name=randread --filename=fio-test --size=4G --rw=randread --bs=4k \
    --ioengine=libaio --direct=1 --iodepth=32 --numjobs=4 --group_reporting \
    --runtime=60 --time_based

Look at the IOPS figure and the latency percentiles in the output:

  read: IOPS=94.2k, BW=368MiB/s (386MB/s)(21.6GiB/60001msec)
...
     | 99.00th=[ 2769], 99.50th=[ 3392], 99.90th=[ 5145], 99.95th=[ 6128],

Switch the scheduler as shown in Step 3, run the same command again and compare. Repeat with a mixed workload (--rw=randrw --rwmixread=70) or a sequential one (--rw=read --bs=1M) that matches what the server really does. Remove the test file when you finish:

rm fio-test

Step 7 - Monitoring I/O in production

The iostat tool from the sysstat package shows per-device latency and utilization:

sudo apt install sysstat
iostat -xz 2

Focus on these columns:

  • r_await and w_await: average time in milliseconds a read or write takes, including queueing. This is the latency your applications feel.
  • aqu-sz: average queue size. A persistently high value means the device cannot keep up.
  • %util: time the device was busy. On NVMe and RAID it can reach 100% while still having spare capacity, so trust await more.

To find which process generates the I/O, use iotop:

sudo apt install iotop-c
sudo iotop-c -o

Troubleshooting

  • tee: /sys/block/sda/queue/scheduler: Invalid argument: the scheduler is not available. Load its module with modprobe (Step 2) and check the spelling.
  • The udev rule does not apply after reboot: make sure the module is listed in /etc/modules-load.d/ and that the KERNEL== pattern matches your device name (sd*, vd* or nvme*n*). Test with sudo udevadm test /sys/block/sda.
  • No difference in benchmarks: expected on NVMe and virtual disks, where the device or the hypervisor does its own scheduling. Keep none in that case, since it has the lowest CPU overhead.

Conclusion

You now know which scheduler each of your disks uses, how to switch it safely, how to make the choice persistent with a udev rule and how to prove the benefit with fio. As a rule of thumb, use none on NVMe and virtual disks, mq-deadline on SATA SSDs and bfq on spinning disks. As next steps, review memory and swap tuning to reduce unnecessary disk I/O, and set up monitoring of await latency so you notice when a disk starts to saturate.