Paperless-ngx is a self-hosted document management system: it takes scanned or digital documents, runs OCR on them, stores a searchable PDF/A archive and lets you organize everything with tags, correspondents and document types. In this tutorial you will deploy Paperless-ngx on Ubuntu 24.04 with Docker Compose, PostgreSQL and Redis, publish it over HTTPS with Nginx, configure OCR for your languages, feed it documents through a watched folder and schedule full exports for backup.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS, with at least 2 GB of RAM (4 GB recommended if you process many documents or several OCR languages).
- A non-root user with
sudoprivileges. - Docker Engine and the Docker Compose plugin installed. See How to install Docker on Linux.
- A domain name, referred to as
your_domain(for exampledocs.example.com), with a DNS A record pointing to your server.
Step 1 - Creating the project directory
Create a directory for the Compose project. The consume folder is where you drop documents to import, and export receives the backups created in Step 7:
sudo mkdir -p /opt/paperless/{consume,export}
cd /opt/paperless
Paperless-ngx runs as an unprivileged user inside the container. To let it read files that you or a scanner drop into consume, map that user to your own user ID. Check your UID and GID:
id -u; id -g
1000
1000
Give your user ownership of the two bind-mounted folders:
sudo chown -R 1000:1000 /opt/paperless/consume /opt/paperless/export
Step 2 - Writing the configuration
Generate a secret key for Django sessions and a password for PostgreSQL:
openssl rand -hex 32
openssl rand -hex 24
Create an environment file with the Paperless-ngx settings:
sudo nano /opt/paperless/docker-compose.env
Paste the following, replacing your_domain, the time zone and the two generated values:
# Map the container user to your user (from `id -u` / `id -g`)
USERMAP_UID=1000
USERMAP_GID=1000
# Public URL, needed for CSRF protection behind the reverse proxy
PAPERLESS_URL=https://your_domain
PAPERLESS_SECRET_KEY=your_secret_key
PAPERLESS_TIME_ZONE=Europe/Madrid
# Database
PAPERLESS_DBPASS=your_db_password
# OCR: languages used on every document (Tesseract codes joined with +)
PAPERLESS_OCR_LANGUAGE=eng+spa
# Extra Tesseract language packs installed at container start (space separated)
PAPERLESS_OCR_LANGUAGES=spa
The container image ships with English and a few other languages. PAPERLESS_OCR_LANGUAGES installs additional packs (here Spanish) when the container starts, and PAPERLESS_OCR_LANGUAGE tells Tesseract which ones to use. Tesseract uses three-letter codes such as deu, fra, ita or por.
Now create the Compose file:
sudo nano /opt/paperless/compose.yaml
services:
broker:
image: docker.io/library/redis:7
restart: unless-stopped
volumes:
- redisdata:/data
db:
image: docker.io/library/postgres:17
restart: unless-stopped
volumes:
- pgdata:/var/lib/postgresql/data
environment:
POSTGRES_DB: paperless
POSTGRES_USER: paperless
POSTGRES_PASSWORD: ${PAPERLESS_DBPASS}
webserver:
image: ghcr.io/paperless-ngx/paperless-ngx:latest
restart: unless-stopped
depends_on:
- db
- broker
ports:
- "127.0.0.1:8000:8000"
volumes:
- data:/usr/src/paperless/data
- media:/usr/src/paperless/media
- ./export:/usr/src/paperless/export
- ./consume:/usr/src/paperless/consume
env_file: docker-compose.env
environment:
PAPERLESS_REDIS: redis://broker:6379
PAPERLESS_DBHOST: db
volumes:
data:
media:
pgdata:
redisdata:
The PostgreSQL service reads its password from the same variable. Create a .env file so Compose can substitute ${PAPERLESS_DBPASS}, then restrict both files:
grep '^PAPERLESS_DBPASS=' /opt/paperless/docker-compose.env | sudo tee /opt/paperless/.env > /dev/null
sudo chmod 600 /opt/paperless/.env /opt/paperless/docker-compose.env
The web port is bound to 127.0.0.1 because ports published by Docker bypass UFW; Nginx will expose it publicly.
Step 3 - Starting Paperless-ngx and creating the admin user
Validate the Compose file and start the stack:
cd /opt/paperless
sudo docker compose config --quiet
sudo docker compose up -d
The first start runs database migrations and installs the extra OCR language packs, which takes a minute or two. Follow the logs until the web server is up, then press Ctrl+C:
sudo docker compose logs -f webserver
Create the first superuser. The command asks for a username, email and password:
sudo docker compose run --rm webserver createsuperuser
Confirm that the application answers locally:
curl -sI http://127.0.0.1:8000/ | head -n 1
HTTP/1.1 302 Found
The redirect points to the login page, which means Paperless-ngx is running.
Step 4 - Configuring Nginx and HTTPS
Install Nginx and Certbot:
sudo apt update
sudo apt install nginx certbot python3-certbot-nginx
Create the server block:
sudo nano /etc/nginx/sites-available/paperless
server {
listen 80;
listen [::]:80;
server_name your_domain;
client_max_body_size 100M;
location / {
proxy_pass http://127.0.0.1:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
}
The WebSocket headers let the web interface show document processing status in real time, and client_max_body_size allows uploading large scanned PDFs.
Enable the site, open the firewall and request a certificate:
sudo ln -s /etc/nginx/sites-available/paperless /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
sudo ufw allow OpenSSH
sudo ufw allow 'Nginx Full'
sudo ufw enable
sudo certbot --nginx -d your_domain
Open https://your_domain and log in with the superuser you created. If you get a "CSRF verification failed" error, PAPERLESS_URL does not match the address in your browser; fix it and run sudo docker compose up -d again.
Step 5 - Adding documents
There are three ways to get documents into Paperless-ngx:
- Web upload: drag files onto the dashboard or use the Upload button.
- Consume folder: any file placed in
/opt/paperless/consumeis picked up, processed and removed from the folder. Point a network scanner's SMB/FTP target or a sync tool at it. - Email: Paperless-ngx can poll IMAP mailboxes and import attachments (Step 6).
Test the consume folder by copying a PDF into it:
cp ~/invoice.pdf /opt/paperless/consume/
Watch the processing in the logs:
cd /opt/paperless
sudo docker compose logs -f --tail 20 webserver
[INFO] [paperless.consumer] Consuming invoice.pdf
[INFO] [paperless.consumer] Document ... consumption finished
The file disappears from consume and the document appears in Documents, with the OCR text searchable. Paperless stores the original file and an archived PDF/A version with an embedded text layer.
You can tune how OCR behaves with these optional settings in docker-compose.env:
# skip: only OCR pages without text (default). redo / force: always OCR
PAPERLESS_OCR_MODE=skip
# Straighten and rotate scanned pages before OCR
PAPERLESS_OCR_DESKEW=true
PAPERLESS_OCR_ROTATE_PAGES=true
# Use subfolders of consume as tags (consume/bills/x.pdf gets the tag "bills")
PAPERLESS_CONSUMER_RECURSIVE=true
PAPERLESS_CONSUMER_SUBDIRS_AS_TAGS=true
Apply changes to the environment file by recreating the container:
sudo docker compose up -d
Noteif the consume folder is a network share (NFS or SMB), inotify events do not arrive. Set
PAPERLESS_CONSUMER_POLLING=60to scan it every 60 seconds instead.
Step 6 - Organizing documents automatically
Paperless-ngx assigns metadata with three kinds of objects: tags, correspondents (who sent the document) and document types (invoice, contract, payslip). Each one has a matching algorithm:
| Algorithm | Behavior |
|---|---|
| Auto | Learns from documents you tagged manually (default) |
| Any / All | Matches if any or all of the listed words appear |
| Exact | Matches the exact phrase |
| Regular expression | Matches a regex against the OCR text |
| Fuzzy | Tolerates small OCR errors in the phrase |
A practical approach is to create a handful of correspondents and document types, assign them manually on the first documents, and let the Auto algorithm take over as the classifier learns. The classifier is retrained automatically every hour.
For rules that go beyond matching, open Workflows. A workflow has a trigger (for example "Document added" with a filename filter, or "Consumption started" from a given mail rule) and actions such as assigning tags, a document type, an owner or a storage path.
To import email attachments, add the mailbox under Mail > Mail accounts (IMAP server, port 993, SSL, and an app password if your provider requires one), then create a Mail rule that selects the folder, filters by sender or subject and chooses what to do with processed messages.
Search uses full-text queries with field filters, for example:
tag:bills correspondent:electric created:[2025 to now]
Step 7 - Backing up with the document exporter
The document exporter writes all documents, thumbnails and metadata into export in a format that Paperless-ngx can import again on any version. Run it once to test:
cd /opt/paperless
sudo docker compose exec -T webserver document_exporter ../export
Check that the export contains your documents and a manifest.json, which holds all the metadata:
ls -lh /opt/paperless/export
Schedule it every night with a cron file:
sudo nano /etc/cron.d/paperless-export
# Export all Paperless-ngx documents and metadata every night at 02:30
30 2 * * * root cd /opt/paperless && docker compose exec -T webserver document_exporter ../export > /dev/null 2>&1
The exporter updates the existing export in place, so it only rewrites changed documents. Copy /opt/paperless/export off the server with your usual backup tool. To restore on a new server, install Paperless-ngx the same way and run document_importer ../export.
Updating Paperless-ngx
Run the exporter, then pull the new image and recreate the container. Database migrations run automatically on start:
cd /opt/paperless
sudo docker compose exec -T webserver document_exporter ../export
sudo docker compose pull
sudo docker compose up -d
Troubleshooting
Files stay in the consume folder. Check sudo docker compose logs --tail 50 webserver for errors. Permission errors mean the files are not readable by the UID set in USERMAP_UID; fix ownership with sudo chown 1000:1000 /opt/paperless/consume/*. On network shares, enable PAPERLESS_CONSUMER_POLLING.
OCR text is garbled. The document language is probably missing from PAPERLESS_OCR_LANGUAGE. Add it (and to PAPERLESS_OCR_LANGUAGES if it is not built in), recreate the container, then reprocess the document from its detail page with Reprocess.
The server runs out of memory during OCR. Limit parallel processing in docker-compose.env with PAPERLESS_TASK_WORKERS=1 and PAPERLESS_THREADS_PER_WORKER=1, then recreate the container.
Search does not find recent documents. Rebuild the full-text index with sudo docker compose exec webserver document_index reindex.
Conclusion
Paperless-ngx now runs on Ubuntu 24.04 with PostgreSQL and Redis behind Nginx with HTTPS, processes documents from the web, the consume folder or email, and exports everything nightly. As next steps, point your scanner at the consume folder, install a mobile client that supports the Paperless-ngx REST API, and ship the export folder off the server with restic.
