Before a server can run CUDA, PyTorch, Ollama or any other GPU workload, it needs the NVIDIA kernel driver and its user space tools. On Ubuntu the most reliable way to get them is through Canonical's own packages, which ship prebuilt and signed kernel modules that follow kernel updates automatically. In this tutorial you will install the NVIDIA server driver on a headless Ubuntu 24.04 machine, verify it with nvidia-smi, and learn how to check it after kernel upgrades.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS with an NVIDIA GPU attached (bare metal, or a virtual machine with the GPU passed through).
- A non-root user with
sudoprivileges. - Console access (IPMI, KVM or the provider's web console) is recommended, in case a reboot does not come back as expected.
This guide covers compute servers without a desktop. It installs the headless driver, which has no X11 or Wayland components.
Step 1 - Checking the GPU and the current state
First confirm that the operating system sees the GPU on the PCI bus:
lspci | grep -i nvidia
01:00.0 3D controller: NVIDIA Corporation AD102GL [L40S] (rev a1)
If the command returns nothing, the GPU is not visible to the OS. On a virtual machine this usually means passthrough is not configured; no driver will fix that.
Next, check whether an NVIDIA driver is already installed, so you do not end up with two conflicting versions:
dpkg -l | grep -E 'nvidia-(driver|headless|utils)'
No output means a clean system. If you see packages from an older installation, or you previously used NVIDIA's .run installer, remove that driver first (sudo apt purge '^nvidia-.*' for packages, or sudo nvidia-uninstall for the runfile) and reboot.
Finally, check whether Secure Boot is enabled. It matters for how the kernel module gets loaded:
mokutil --sb-state
SecureBoot enabled
The Ubuntu packages used below ship modules signed by Canonical, so they load with Secure Boot on. Only if you end up using a DKMS build (the NVIDIA repository method) will you need to enroll a key, covered in Troubleshooting.
Step 2 - Choosing the driver branch
Install the helper tool that detects the GPU and lists matching driver packages:
sudo apt update
sudo apt install ubuntu-drivers-common
List the drivers intended for general purpose GPU computing:
sudo ubuntu-drivers list --gpgpu
nvidia-driver-570-server, (kernel modules provided by linux-modules-nvidia-570-server-generic)
nvidia-driver-580-server, (kernel modules provided by linux-modules-nvidia-580-server-generic)
nvidia-driver-580-server-open, (kernel modules provided by linux-modules-nvidia-580-server-open-generic)
The numbers you see depend on your GPU and on when you run the command. Some guidance for choosing:
| Package suffix | Use it when |
|---|---|
-server | Production compute servers. These are NVIDIA's long-lived data center branches, validated for longer. |
-server-open | Turing (T4, RTX 20) or newer GPUs. NVIDIA recommends the open kernel modules for these, and Blackwell GPUs require them. |
| no suffix | Desktop or workstation drivers. Avoid them on servers. |
As a rule, pick the highest -server-open branch for recent data center GPUs (A100, H100, L40S and newer), and the -server branch for older architectures such as Pascal or Volta. Also check the CUDA version your frameworks need: CUDA 13 requires driver 580 or newer.
Step 3 - Installing the driver
Install the branch you chose. The --gpgpu flag selects the headless compute driver and the prebuilt, signed kernel modules for your kernel flavor. Replace 580-server-open with the branch from the previous step:
sudo ubuntu-drivers install --gpgpu nvidia:580-server-open
The --gpgpu install does not include the command line utilities. Install them from the same branch so nvidia-smi is available:
sudo apt install nvidia-utils-580-server
The utilities package is shared between the open and proprietary variants, so there is no -open version of it.
Ubuntu's NVIDIA packages blacklist the open source nouveau driver, which would otherwise claim the GPU at boot. Reboot to load the new module:
sudo reboot
Step 4 - Verifying the installation
After the server comes back, log in and confirm that the nvidia module is loaded and nouveau is not:
lsmod | grep -E '^(nvidia|nouveau)'
nvidia_uvm 2076672 0
nvidia_drm 135168 0
nvidia_modeset 1638400 1 nvidia_drm
nvidia 104071168 2 nvidia_uvm,nvidia_modeset
Now query the GPU with nvidia-smi:
nvidia-smi
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 |
|-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
|=========================================+========================+======================|
| 0 NVIDIA L40S Off | 00000000:01:00.0 Off | 0 |
| N/A 31C P8 33W / 350W | 0MiB / 46068MiB | 0% Default |
+-----------------------------------------+------------------------+----------------------+
The CUDA Version field is the newest CUDA runtime this driver supports, not an installed toolkit. You install the CUDA Toolkit separately if you need nvcc.
For a compact, script-friendly check, query specific fields:
nvidia-smi --query-gpu=name,driver_version,memory.total,compute_cap --format=csv
name, driver_version, memory.total [MiB], compute_cap
NVIDIA L40S, 580.95.05, 46068 MiB, 8.9
Step 5 - Enabling persistence mode
Without a client using it, the driver tears down the GPU state and has to initialize it again on the next job, which adds latency to the first call. Persistence mode keeps the GPU initialized:
sudo nvidia-smi -pm 1
Enabled persistence mode for GPU 00000000:01:00.0.
All done.
This setting does not survive a reboot. To keep it permanently, NVIDIA provides the nvidia-persistenced daemon. Check whether your driver packages ship its systemd unit and enable it:
systemctl list-unit-files | grep nvidia-persistenced
sudo systemctl enable --now nvidia-persistenced
Confirm the state after the next reboot with nvidia-smi --query-gpu=persistence_mode --format=csv.
Step 6 - Keeping the driver working after kernel upgrades
Because ubuntu-drivers installed a meta package (linux-modules-nvidia-580-server-open-generic in this example), every new generic kernel pulls in the matching NVIDIA module automatically. There is nothing to rebuild.
You can confirm that the module package for the running kernel is present:
dpkg -l | grep "linux-modules-nvidia.*$(uname -r)"
ii linux-modules-nvidia-580-server-open-6.8.0-85-generic 6.8.0-85.85 amd64 Linux kernel nvidia modules for version 6.8.0-85
After each apt upgrade that installs a new kernel, reboot and run nvidia-smi again. If it fails, see the next section.
To stop routine upgrades from moving you to a different driver branch, you do not need to hold anything: the branch number is part of the package name, so apt only installs updates within the branch you chose.
Alternative - Using NVIDIA's CUDA repository
If you need a driver version that Ubuntu has not packaged yet, NVIDIA publishes drivers in its CUDA repository. These builds use DKMS, so the module is compiled locally on each kernel update and, with Secure Boot on, needs a Machine Owner Key. Only use this method if the Ubuntu packages do not fit.
Install the kernel headers and the repository keyring package:
sudo apt install linux-headers-$(uname -r)
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
Then install either the open kernel modules or the proprietary ones:
sudo apt install nvidia-open
sudo apt install cuda-drivers
Reboot and verify with dkms status and nvidia-smi. Do not mix this method with the Ubuntu packages on the same machine.
Troubleshooting
nvidia-smi reports "couldn't communicate with the NVIDIA driver". The module is not loaded. Check the kernel log for the reason:
sudo journalctl -k -b | grep -iE 'nvidia|nouveau'
A message like Lockdown: modprobe: unsigned module loading is restricted points to Secure Boot rejecting a DKMS-built module (see below). nouveau messages mean it grabbed the GPU first: confirm the blacklist exists with grep -r nouveau /etc/modprobe.d /usr/lib/modprobe.d, then run sudo update-initramfs -u and reboot.
Secure Boot blocks a DKMS module. With the NVIDIA repository method, DKMS signs modules with a local key that the firmware does not trust yet. Enroll it, set a one-time password, and confirm the enrollment on the console during the next boot (the blue MOK Manager screen):
sudo mokutil --import /var/lib/shim-signed/mok/MOK.der
sudo reboot
This step requires console access to the server.
nvidia-smi fails after a kernel upgrade. If you used ubuntu-drivers, check that the module package for the new kernel is installed (Step 6) and install it if not, for example sudo apt install linux-modules-nvidia-580-server-open-$(uname -r). With DKMS, rebuild the module with sudo dkms autoinstall and look at dkms status for build errors, which usually mean missing kernel headers.
"Driver/library version mismatch". The user space libraries were upgraded but the old kernel module is still loaded. Reboot the server to load the matching module.
Conclusion
Your Ubuntu 24.04 server now has the NVIDIA server driver installed from Canonical's packages, with signed modules that load under Secure Boot and follow kernel updates on their own. nvidia-smi confirms the GPU, driver version and the highest CUDA version supported. From here you can install the CUDA Toolkit to compile GPU code, install Docker with the NVIDIA Container Toolkit to run GPU containers, or deploy an inference server such as Ollama.
