The NVIDIA CUDA Toolkit contains the nvcc compiler, the CUDA runtime and math libraries, and the headers needed to build GPU-accelerated software. In this tutorial you will install the toolkit from NVIDIA's official APT repository on Ubuntu 24.04, add cuDNN, configure your shell, switch between toolkit versions, and compile a small CUDA program to prove everything works.

Prerequisites

To follow this guide you need:

  • A server running Ubuntu 24.04 LTS with an NVIDIA GPU.
  • A working NVIDIA driver: nvidia-smi must list your GPU. If it does not, follow a driver installation guide first.
  • A non-root user with sudo privileges.
  • About 8 GB of free disk space for the toolkit and cuDNN.

Step 1 - Choosing a CUDA version

Every CUDA release needs a minimum driver version. Check the driver you have and the newest CUDA version it supports:

nvidia-smi --query-gpu=name,driver_version,compute_cap --format=csv
nvidia-smi | grep "CUDA Version"
name, driver_version, compute_cap
NVIDIA L40S, 580.95.05, 8.9
| NVIDIA-SMI 580.95.05              Driver Version: 580.95.05      CUDA Version: 13.0     |

The toolkit you install must not be newer than the CUDA Version shown. As a reference, CUDA 12.x works with drivers 525 and newer, and CUDA 13.x needs driver 580 or newer. Also check which version your frameworks target: the PyTorch "Get Started" page and the TensorFlow install page list the CUDA versions they support. Note the compute_cap value; you will use it when compiling.

Step 2 - Adding the NVIDIA CUDA repository

NVIDIA publishes a cuda-keyring package that installs its signing key under /usr/share/keyrings and the repository definition under /etc/apt/sources.list.d. Download and install it:

wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update

On an arm64 server (for example Grace Hopper), replace x86_64 with sbsa in the URL.

List the toolkit versions available to you:

apt-cache search --names-only '^cuda-toolkit-[0-9]+-[0-9]+$'
cuda-toolkit-12-8 - CUDA Toolkit 12.8 meta-package
cuda-toolkit-12-9 - CUDA Toolkit 12.9 meta-package
cuda-toolkit-13-0 - CUDA Toolkit 13.0 meta-package

Step 3 - Installing the toolkit

Install the compiler toolchain that nvcc relies on for host code, then the toolkit version you chose. Use the versioned cuda-toolkit-X-Y package: it installs only the toolkit and never touches your driver. Avoid the cuda and cuda-X-Y meta packages, which also pull in a driver.

sudo apt install build-essential
sudo apt install cuda-toolkit-13-0

The toolkit is installed under /usr/local/cuda-13.0, and a /usr/local/cuda symlink points to the default version:

ls -l /usr/local | grep cuda
lrwxrwxrwx  1 root root   22 Sep 25 10:12 cuda -> /etc/alternatives/cuda
lrwxrwxrwx  1 root root   25 Sep 25 10:12 cuda-13 -> /etc/alternatives/cuda-13
drwxr-xr-x 15 root root 4096 Sep 25 10:12 cuda-13.0

Step 4 - Configuring your environment

The toolkit's binaries are not on the default PATH. Add them system-wide with a profile script so every login shell finds nvcc:

sudo nano /etc/profile.d/cuda.sh
export CUDA_HOME=/usr/local/cuda
export PATH="$CUDA_HOME/bin:$PATH"

Pointing at /usr/local/cuda rather than a specific version means you can switch versions later without editing this file. Load it in the current session:

source /etc/profile.d/cuda.sh

Verify the compiler:

nvcc --version
nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2025 NVIDIA Corporation
Cuda compilation tools, release 13.0, V13.0.88

The Debian packages register the CUDA library directory with the dynamic linker, so you normally do not need LD_LIBRARY_PATH. Check that the runtime library is found:

ldconfig -p | grep libcudart
	libcudart.so.13 (libc6,x86-64) => /usr/local/cuda-13.0/targets/x86_64-linux/lib/libcudart.so.13

If this prints nothing, add export LD_LIBRARY_PATH="$CUDA_HOME/lib64${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}" to /etc/profile.d/cuda.sh.

Step 5 - Compiling and running a test program

A quick program that queries the GPU and runs a kernel proves that the compiler, runtime and driver all agree. Create a working directory and the source file:

mkdir -p ~/cuda-test && cd ~/cuda-test
nano vector_add.cu
#include <cstdio>
#include <cuda_runtime.h>

#define CHECK(call) do { cudaError_t e = (call); if (e != cudaSuccess) { \
    fprintf(stderr, "CUDA error: %s (%s:%d)\n", cudaGetErrorString(e), __FILE__, __LINE__); return 1; } } while (0)

__global__ void add(const float *a, const float *b, float *c, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < n) c[i] = a[i] + b[i];
}

int main() {
    cudaDeviceProp prop;
    CHECK(cudaGetDeviceProperties(&prop, 0));
    printf("GPU: %s, compute capability %d.%d\n", prop.name, prop.major, prop.minor);

    const int n = 1 << 20;
    float *a, *b, *c;
    CHECK(cudaMallocManaged(&a, n * sizeof(float)));
    CHECK(cudaMallocManaged(&b, n * sizeof(float)));
    CHECK(cudaMallocManaged(&c, n * sizeof(float)));
    for (int i = 0; i < n; i++) { a[i] = 1.0f; b[i] = 2.0f; }

    add<<<(n + 255) / 256, 256>>>(a, b, c, n);
    CHECK(cudaGetLastError());
    CHECK(cudaDeviceSynchronize());

    for (int i = 0; i < n; i++) {
        if (c[i] != 3.0f) { printf("FAILED at %d\n", i); return 1; }
    }
    printf("Result = PASS (%d elements)\n", n);
    cudaFree(a); cudaFree(b); cudaFree(c);
    return 0;
}

Compile it for the GPU in this machine. -arch=native tells nvcc to detect the installed GPU's architecture:

nvcc -arch=native -o vector_add vector_add.cu
./vector_add
GPU: NVIDIA L40S, compute capability 8.9
Result = PASS (1048576 elements)

If you build binaries to run on other machines, target the architectures explicitly instead, for example -arch=sm_89 for compute capability 8.9 or -arch=sm_90 for Hopper.

Step 6 - Installing cuDNN

cuDNN is NVIDIA's library of deep learning primitives (convolutions, attention, normalization). It is available from the same repository. Install the build that matches your CUDA major version:

sudo apt install cudnn9-cuda-13

For a CUDA 12 toolkit, install cudnn9-cuda-12 instead. Confirm which cuDNN packages were installed and their version:

dpkg -l | grep -i cudnn
ii  cudnn9-cuda-13                  9.13.0.50-1  amd64  NVIDIA cuDNN for CUDA 13
ii  libcudnn9-cuda-13               9.13.0.50-1  amd64  cuDNN runtime libraries for CUDA 13
ii  libcudnn9-dev-cuda-13           9.13.0.50-1  amd64  cuDNN development libraries for CUDA 13
ii  libcudnn9-headers-cuda-13       9.13.0.50-1  amd64  cuDNN header files for CUDA 13

Also check that the dynamic linker can find the library:

ldconfig -p | grep libcudnn.so

Step 7 - Managing multiple CUDA versions

Different projects often need different CUDA versions. Toolkits install side by side, each in its own /usr/local/cuda-X.Y directory:

sudo apt install cuda-toolkit-12-9

The /usr/local/cuda symlink is managed by update-alternatives. List the installed versions and choose the system default:

sudo update-alternatives --config cuda
There are 2 choices for the alternative cuda (providing /usr/local/cuda).

  Selection    Path                  Priority   Status
------------------------------------------------------------
* 0            /usr/local/cuda-13.0   130       auto mode
  1            /usr/local/cuda-12.9   129       manual mode
  2            /usr/local/cuda-13.0   130       manual mode

Press <enter> to keep the current choice[*], or type selection number:

Because /etc/profile.d/cuda.sh uses /usr/local/cuda, open a new shell and nvcc --version reports the new default.

To use a different version for a single session or build without changing the system default, override the variables in that shell only:

export CUDA_HOME=/usr/local/cuda-12.9
export PATH="$CUDA_HOME/bin:$PATH"
nvcc --version | grep release
Cuda compilation tools, release 12.9, V12.9.86

Troubleshooting

nvcc: command not found. The toolkit is installed but not on your PATH. Confirm /usr/local/cuda/bin/nvcc exists and that you opened a new login shell or ran source /etc/profile.d/cuda.sh. Note that sudo resets PATH, so use the full path /usr/local/cuda/bin/nvcc in commands run with sudo.

CUDA error: CUDA driver version is insufficient for CUDA runtime version. The toolkit is newer than the driver supports. Compare with the CUDA Version shown by nvidia-smi, then either upgrade the driver or switch to an older toolkit with update-alternatives --config cuda.

no kernel image is available for execution on the device. The binary was compiled for a different GPU architecture. Rebuild with -arch=native, or with the sm_XX value that matches your compute_cap.

apt wants to install a driver together with the toolkit. You installed cuda or cuda-X-Y instead of cuda-toolkit-X-Y. Cancel the transaction and install the cuda-toolkit- package so your existing driver is left alone.

PyTorch reports a different CUDA version than nvcc. This is expected. python3 -c "import torch; print(torch.version.cuda)" shows the runtime bundled in the wheel, which is independent of the system toolkit. Only the driver has to be new enough for both.

Conclusion

You installed the CUDA Toolkit from NVIDIA's repository on Ubuntu 24.04, configured the shell through /etc/profile.d, added cuDNN, and confirmed the setup with a compiled vector addition kernel. Versions live side by side and update-alternatives controls the default. Next, you can build the official CUDA samples from NVIDIA's cuda-samples repository, set up a Python virtual environment with PyTorch, or install the NVIDIA Container Toolkit to run CUDA workloads in Docker.