Ollama is an open source runtime that downloads, manages and serves large language models such as Llama, Gemma, Qwen and Mistral on your own hardware. It exposes a simple CLI and an HTTP API, including an OpenAI-compatible endpoint, so existing tools can talk to it with minimal changes. In this tutorial you will install Ollama on Ubuntu 24.04, run your first model, use the API, configure the service, and add Open WebUI as a private chat interface.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS or bare metal server.
- A non-root user with
sudoprivileges. - Enough memory for the models you plan to run. As a rough guide for 4-bit models: 8 GB of RAM for 3B to 8B models, 16 GB for 12B to 14B models, and 48 GB or more for 70B models.
- Disk space for the models: a 3B model takes about 2 GB, an 8B model about 5 GB, a 70B model about 40 GB.
- Optional: an NVIDIA GPU with the driver installed (
nvidia-smiworks). Ollama runs on CPU without one, only more slowly. - Optional, for Step 6: Docker Engine installed.
Step 1 - Installing Ollama
Ollama's supported Linux installation is its install script. It downloads the binaries, creates an ollama system user and a systemd service, and detects NVIDIA or AMD GPUs. Download the script first so you can read what it does before running it:
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
When you are satisfied, run it:
sh ollama-install.sh
The script uses sudo for the steps that need root. At the end it should report that the API is available at 127.0.0.1:11434, and whether it found a GPU.
Verify the binary and the service:
ollama -v
systemctl status ollama --no-pager
ollama version is 0.12.3
● ollama.service - Ollama Service
Loaded: loaded (/etc/systemd/system/ollama.service; enabled; preset: enabled)
Active: active (running) since Thu 2026-09-25 11:02:40 UTC; 20s ago
The API should answer on localhost:
curl http://127.0.0.1:11434
Ollama is running
Step 2 - Running your first model
Models are pulled from the Ollama library at ollama.com/library, where each model has tags for sizes and quantizations. Start with a small model that runs well even on CPU:
ollama run llama3.2:3b
The first run downloads the model, then opens an interactive prompt:
pulling manifest
pulling dde5aa3fc5ff: 100% ▕████████████████▏ 2.0 GB
...
success
>>> Explain what a reverse proxy is in two sentences.
Type /bye or press Ctrl+D to leave. You can also pass a prompt directly, which is useful in scripts:
ollama run llama3.2:3b "Write a one-line bash command that shows the 5 largest files in /var/log"
Step 3 - Managing models
Download a model without starting a chat:
ollama pull gemma3:4b
List the models on disk and their sizes:
ollama list
NAME ID SIZE MODIFIED
gemma3:4b a2af6cc3eb7f 3.3 GB 10 seconds ago
llama3.2:3b a80c4f17acd5 2.0 GB 3 minutes ago
See which models are loaded in memory right now, and whether they run on the CPU or the GPU:
ollama ps
NAME ID SIZE PROCESSOR CONTEXT UNTIL
llama3.2:3b a80c4f17acd5 3.1 GB 100% GPU 4096 4 minutes from now
100% GPU means the whole model fits in VRAM. A split such as 45%/55% CPU/GPU means part of it runs on the CPU, which is much slower; use a smaller model or quantization if speed matters. Inspect a model's parameters, template and license with ollama show llama3.2:3b, and delete one you no longer need with ollama rm gemma3:4b.
Creating a custom model with a Modelfile
A Modelfile builds a new model on top of an existing one with your own system prompt and parameters. Create one:
nano Modelfile
FROM llama3.2:3b
SYSTEM """
You are a concise Linux system administration assistant. Answer with the exact commands
needed and a one-line explanation of each.
"""
PARAMETER temperature 0.3
PARAMETER num_ctx 8192
Build and run it:
ollama create sysadmin -f Modelfile
ollama run sysadmin "How do I find which process is listening on port 443?"
Step 4 - Using the REST API
Everything the CLI does goes through the HTTP API on port 11434, so any application on the server can use the models. Generate a single completion:
curl http://127.0.0.1:11434/api/generate -d '{
"model": "llama3.2:3b",
"prompt": "What is a VPS? Answer in one sentence.",
"stream": false
}'
The response is a JSON object; the generated text is in the response field, along with timing data such as eval_count and eval_duration.
For multi-turn conversations, use the chat endpoint and send the message history:
curl http://127.0.0.1:11434/api/chat -d '{
"model": "llama3.2:3b",
"messages": [
{"role": "system", "content": "You answer in one short paragraph."},
{"role": "user", "content": "How does a container differ from a virtual machine?"}
],
"stream": false
}'
Ollama also implements the OpenAI Chat Completions API under /v1, so SDKs and tools built for OpenAI work by changing the base URL:
curl http://127.0.0.1:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.2:3b",
"messages": [{"role": "user", "content": "Say hello"}]
}'
With the official openai Python package, point the client at Ollama. The API key is required by the library but ignored by Ollama:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:11434/v1", api_key="ollama")
reply = client.chat.completions.create(
model="llama3.2:3b",
messages=[{"role": "user", "content": "List three uses of the ss command."}],
)
print(reply.choices[0].message.content)
Step 5 - Configuring the Ollama service
The service is configured through environment variables. Set them in a systemd override so they survive Ollama upgrades:
sudo systemctl edit ollama
Add the settings between the comment lines at the top of the file:
[Service]
Environment="OLLAMA_KEEP_ALIVE=30m"
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_NUM_PARALLEL=2"
Environment="OLLAMA_FLASH_ATTENTION=1"
What each variable does:
OLLAMA_KEEP_ALIVE: how long a model stays in memory after the last request (default 5 minutes). Longer values avoid reload delays.OLLAMA_CONTEXT_LENGTH: the default context window in tokens. Larger windows use more memory.OLLAMA_NUM_PARALLEL: how many requests each loaded model handles at once. Each parallel slot needs its own context memory.OLLAMA_FLASH_ATTENTION: enables flash attention on supported GPUs, which reduces memory use for long contexts.
To store models on a different disk, add Environment="OLLAMA_MODELS=/data/ollama/models" and give the ollama user ownership of that directory with sudo chown -R ollama:ollama /data/ollama. The default location for the service is /usr/share/ollama/.ollama/models.
Apply the changes and check the logs:
sudo systemctl restart ollama
sudo journalctl -u ollama -n 30 --no-pager
The startup log lists the configuration in effect and the GPUs Ollama detected, which is the quickest way to confirm that GPU acceleration is active.
Access from other machines
By default Ollama listens only on 127.0.0.1. The API has no authentication, so anyone who can reach port 11434 can use your models and download or delete them. The safest way to use it from your workstation is an SSH tunnel:
ssh -N -L 11434:127.0.0.1:11434 your_user@your_server_ip
While the tunnel is open, http://127.0.0.1:11434 on your workstation reaches the server. If other servers must call the API directly, add Environment="OLLAMA_HOST=0.0.0.0:11434" to the override, restart the service, and allow only those hosts with UFW:
sudo ufw allow from 203.0.113.10 to any port 11434 proto tcp
Replace 203.0.113.10 with the IP of the client server, and never open the port to everyone.
Step 6 - Adding a web interface with Open WebUI
Open WebUI is a self-hosted chat interface for Ollama with user accounts, conversation history and document upload. Run it in Docker using host networking, so the container can reach Ollama on 127.0.0.1 without exposing Ollama:
docker run -d --network=host \
-v open-webui:/app/backend/data \
-e OLLAMA_BASE_URL=http://127.0.0.1:11434 \
--name open-webui --restart always \
ghcr.io/open-webui/open-webui:main
With host networking, Open WebUI listens on port 8080 of the server. Follow its startup until it reports that it is ready:
docker logs -f open-webui
Do not open port 8080 in the firewall. Reach it through an SSH tunnel from your workstation instead:
ssh -N -L 8080:127.0.0.1:8080 your_user@your_server_ip
Open http://localhost:8080 in your browser and create an account right away: the first account created becomes the administrator. Your Ollama models appear in the model selector. For access by a team, publish Open WebUI behind Nginx with HTTPS instead of the tunnel.
Troubleshooting
Error: model requires more system memory. The model does not fit in the available RAM. Pick a smaller size or quantization tag from the model's page, or reduce OLLAMA_CONTEXT_LENGTH and OLLAMA_NUM_PARALLEL, which both increase memory use.
The GPU is not used (ollama ps shows 100% CPU). Confirm nvidia-smi works, then look at the startup log with sudo journalctl -u ollama -b | grep -iE 'gpu|cuda'. If the driver was installed after Ollama, restart the service with sudo systemctl restart ollama. For more detail, add Environment="OLLAMA_DEBUG=1" to the override temporarily.
could not connect to ollama app, is it running?. The CLI cannot reach the server. Check systemctl status ollama. If you changed OLLAMA_HOST in the service, the CLI still defaults to 127.0.0.1:11434, which is fine for 0.0.0.0; for any other address, export the same OLLAMA_HOST in your shell.
Open WebUI shows no models. Check that Ollama answers with curl http://127.0.0.1:11434/api/tags on the server and that the container was started with --network=host and the OLLAMA_BASE_URL above.
Conclusion
Ollama now runs as a systemd service on Ubuntu 24.04, serving local models through its native and OpenAI-compatible APIs, tuned through a service override and kept private on localhost, with Open WebUI as a chat front end. Next, you can try larger models on a GPU server, connect coding assistants or automation tools to the /v1 endpoint, or put Open WebUI behind Nginx with a Let's Encrypt certificate for your team.
