vLLM is an open source inference engine for large language models. Its PagedAttention memory manager and continuous batching let a single GPU serve many concurrent requests, and it exposes an HTTP API compatible with the OpenAI client libraries. In this tutorial you will install vLLM on Ubuntu 24.04, serve an open model, query it with curl and the OpenAI Python SDK, protect it with an API key and run it as a systemd service.

Prerequisites

To follow this guide you need:

  • A server running Ubuntu 24.04 LTS with an NVIDIA GPU (compute capability 7.0 or newer, such as T4, L4, A10, RTX 4090, A100 or H100).
  • A non-root user with sudo privileges.
  • At least 32 GB of RAM and 60 GB of free disk space for the Python environment and model weights.
  • Optional: a Hugging Face account and access token if you want to serve gated models such as Llama or Gemma.

As a rule of thumb, a model needs about 2 GB of VRAM per billion parameters in 16-bit precision, plus room for the KV cache. The examples use Qwen2.5 7B Instruct, which fits on a 24 GB GPU, and its 4-bit AWQ version for 16 GB GPUs.

Step 1 - Installing the NVIDIA driver

The vLLM wheels ship their own CUDA runtime through PyTorch, so the host only needs a recent NVIDIA driver. Install the recommended one and reboot:

sudo apt update
sudo ubuntu-drivers install
sudo reboot

Verify that the GPU is visible:

nvidia-smi --query-gpu=name,driver_version,memory.total --format=csv
name, driver_version, memory.total [MiB]
NVIDIA L4, 570.xx.xx, 23034 MiB

Step 2 - Creating a service user and a virtual environment

Running the server under a dedicated account keeps downloaded weights and the process separate from your login user. Create a system user vllm with its home in /opt/vllm:

sudo useradd --system --create-home --home-dir /opt/vllm --shell /usr/sbin/nologin vllm

Install the tools to build a virtual environment:

sudo apt install -y python3-venv python3-dev build-essential

vLLM compiles some kernels at startup with Triton, which needs a C compiler and the Python headers, hence build-essential and python3-dev. Create the environment as the vllm user and install vLLM into it:

sudo -u vllm python3 -m venv /opt/vllm/venv
sudo -u vllm /opt/vllm/venv/bin/pip install --upgrade pip
sudo -u vllm /opt/vllm/venv/bin/pip install vllm

The install pulls PyTorch with CUDA support and takes a few minutes. Verify it:

sudo -u vllm /opt/vllm/venv/bin/python -c "import vllm, torch; print(vllm.__version__, torch.cuda.is_available())"
0.x.x True

If the second value is False, PyTorch cannot see the GPU; recheck the driver with nvidia-smi.

Step 3 - Downloading a model

vLLM downloads models from Hugging Face on first start, but pre-downloading makes startup predictable and lets you check disk usage. Weights are stored in the Hugging Face cache, which you will point to /opt/vllm/hf-cache. The hf command line tool is installed with vLLM:

sudo -u vllm HF_HOME=/opt/vllm/hf-cache /opt/vllm/venv/bin/hf download Qwen/Qwen2.5-7B-Instruct

On a 16 GB GPU, download the quantized variant instead: Qwen/Qwen2.5-7B-Instruct-AWQ.

For gated models, accept the license on the model page first and pass your token by adding HF_TOKEN=hf_your_token next to HF_HOME in the command.

Check the space used:

sudo du -sh /opt/vllm/hf-cache
15G	/opt/vllm/hf-cache

Step 4 - Starting the server manually

Start the OpenAI-compatible server in the foreground once to confirm that the model loads:

sudo -u vllm HF_HOME=/opt/vllm/hf-cache /opt/vllm/venv/bin/vllm serve Qwen/Qwen2.5-7B-Instruct \
  --host 127.0.0.1 \
  --port 8000 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.90

The options control memory and exposure:

  • --host 127.0.0.1 keeps the API private to the server. You will add authentication before exposing it.
  • --max-model-len 8192 caps the context window. The model supports longer contexts, but every extra token reserves KV cache memory; lowering it is the easiest fix for out of memory errors at startup.
  • --gpu-memory-utilization 0.90 lets vLLM use 90% of VRAM for weights and KV cache.

Loading takes one to three minutes. The server is ready when it prints:

INFO:     Application startup complete.
INFO:     Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)

In a second SSH session, check the health endpoint and the model list:

curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8000/health
curl -s http://127.0.0.1:8000/v1/models
200
{"object":"list","data":[{"id":"Qwen/Qwen2.5-7B-Instruct","object":"model",...}]}

Step 5 - Sending chat completion requests

The /v1/chat/completions endpoint takes the same JSON body as OpenAI's API. The model field must match the model ID reported by /v1/models:

curl -s http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-7B-Instruct",
    "messages": [
      {"role": "system", "content": "You are a concise Linux administrator."},
      {"role": "user", "content": "How do I list the largest directories under /var?"}
    ],
    "max_tokens": 200,
    "temperature": 0.2
  }'

The response contains the generated text in choices[0].message.content and token counts in usage:

{"id":"chatcmpl-...","object":"chat.completion",...,"choices":[{"index":0,"message":{"role":"assistant","content":"Use du:\n\nsudo du -h --max-depth=1 /var | sort -hr | head"...}}],"usage":{"prompt_tokens":38,"total_tokens":81,"completion_tokens":43}}

Any application that uses the OpenAI SDK can point to vLLM by changing the base URL. On your workstation or application server, install the SDK with pip install openai and run:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[{"role": "user", "content": "Explain what a reverse proxy does in two sentences."}],
    max_tokens=120,
)
print(response.choices[0].message.content)

Add stream=True to the call, or "stream": true to the JSON body, to receive tokens as they are generated.

Stop the foreground server with Ctrl+C before continuing.

Step 6 - Running vLLM as a systemd service with an API key

vLLM accepts any request unless you set an API key. Generate a random key and store it in an environment file readable only by root:

sudo mkdir -p /etc/vllm
echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/vllm/vllm.env > /dev/null
sudo chmod 600 /etc/vllm/vllm.env

If you use a gated model, add a line HF_TOKEN=hf_your_token to the same file. Now create the unit:

sudo nano /etc/systemd/system/vllm.service
[Unit]
Description=vLLM OpenAI-compatible server
After=network-online.target
Wants=network-online.target

[Service]
Type=simple
User=vllm
Group=vllm
WorkingDirectory=/opt/vllm
Environment=HF_HOME=/opt/vllm/hf-cache
EnvironmentFile=/etc/vllm/vllm.env
ExecStart=/opt/vllm/venv/bin/vllm serve Qwen/Qwen2.5-7B-Instruct \
    --host 127.0.0.1 \
    --port 8000 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.90 \
    --api-key ${VLLM_API_KEY}
Restart=on-failure
RestartSec=30
TimeoutStartSec=900

[Install]
WantedBy=multi-user.target

systemd reads the environment file as root before dropping privileges, so the vllm user never needs read access to it. TimeoutStartSec gives large models time to load. Start the service:

sudo systemctl daemon-reload
sudo systemctl enable --now vllm
sudo journalctl -u vllm -f

Wait for Application startup complete, then press Ctrl+C. Verify that requests without the key are rejected and requests with it succeed:

curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8000/v1/models
KEY=$(sudo grep -Po '(?<=VLLM_API_KEY=).*' /etc/vllm/vllm.env)
curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer $KEY" http://127.0.0.1:8000/v1/models
401
200

To reach the API from other machines, put Nginx with TLS in front of 127.0.0.1:8000 or connect over a private network, rather than binding vLLM to 0.0.0.0.

Step 7 - Fitting larger models: quantization and multiple GPUs

When a model does not fit in VRAM, use a pre-quantized checkpoint. vLLM reads the quantization method from the model's config, so you only change the model name:

sudo systemctl edit --full vllm

Replace the model in the ExecStart line, for example with Qwen/Qwen2.5-7B-Instruct-AWQ, save, and restart with sudo systemctl restart vllm. The model ID returned by /v1/models changes accordingly, so update the model field in your clients or add --served-model-name to keep a stable name.

Approximate VRAM for the weights alone:

Parameters16-bit4-bit (AWQ/GPTQ)
7-8B~16 GB~5 GB
14B~28 GB~9 GB
32B~64 GB~18 GB
70-72B~140 GB~40 GB

On a server with several GPUs, split one model across them with tensor parallelism by adding --tensor-parallel-size 2 (or 4, 8) to ExecStart. The value must divide the model's number of attention heads, and all GPUs should be the same model. nvidia-smi then shows a vllm process using memory on each card.

Troubleshooting

ValueError: ... max seq len ... is larger than the maximum number of tokens that can be stored in KV cache. The context length does not fit in the memory left after loading the weights. Lower --max-model-len, raise --gpu-memory-utilization slightly (up to 0.95), or switch to a quantized model.

torch.OutOfMemoryError at startup. Another process is using the GPU. Check nvidia-smi, stop it, and restart vLLM.

401 Client Error or GatedRepoError when downloading. The model is gated. Accept its license on Hugging Face and provide HF_TOKEN in /etc/vllm/vllm.env.

The service is killed after a few minutes with no error. Check journalctl -u vllm for a start timeout and raise TimeoutStartSec, or look for the kernel OOM killer in sudo dmesg, which means the server needs more system RAM.

Conclusion

vLLM is now serving an open model on your Ubuntu 24.04 GPU server through an OpenAI-compatible API protected by a key, and systemd keeps it running across reboots. Good next steps are adding a chat interface such as Open WebUI on top of the API, putting Nginx with HTTPS in front of it for remote clients, and load testing with realistic prompts to tune --max-model-len and concurrency.