Apache Solr is a mature open source search platform built on Apache Lucene. It provides full-text search with configurable analysis, faceting, highlighting and horizontal scaling through SolrCloud. In this tutorial you will install Solr 10 on Ubuntu 24.04 with the official service installer, create a collection, define its schema with the Schema API, index JSON documents and run relevance-boosted, faceted searches.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS with at least 2 GB of RAM, for example a CubePath VPS. Use 4 GB or more for real indexes; Solr relies on both the Java heap and the operating system cache.
- A non-root user with
sudoprivileges. - SSH access from your local machine, used to open the Solr Admin UI through a tunnel.
Step 1 - Installing Java
Solr 10 requires Java 21 or later. Ubuntu 24.04 ships OpenJDK 21 in its main repository. Install the headless runtime together with lsof, which the Solr start script uses to check ports, and jq to read JSON responses:
sudo apt update
sudo apt install -y openjdk-21-jre-headless lsof curl jq
Verify the Java version:
java -version
openjdk version "21.0.x" ...
OpenJDK Runtime Environment (build 21.0.x+...-Ubuntu-...)
OpenJDK 64-Bit Server VM (build 21.0.x+...-Ubuntu-..., mixed mode, sharing)
Step 2 - Downloading Solr
Download the Solr 10.0.0 binary release and its checksum from the Apache CDN. Check the Solr downloads page for the current version; older releases move to archive.apache.org:
cd /tmp
curl -fLO https://dlcdn.apache.org/solr/solr/10.0.0/solr-10.0.0.tgz
curl -fLO https://dlcdn.apache.org/solr/solr/10.0.0/solr-10.0.0.tgz.sha512
Verify the archive:
sha512sum -c solr-10.0.0.tgz.sha512
solr-10.0.0.tgz: OK
Step 3 - Installing Solr as a service
The release includes install_solr_service.sh, which creates a solr system user, installs Solr under /opt, puts data and logs in /var/solr, and registers a systemd unit. Extract only that script:
tar xzf solr-10.0.0.tgz solr-10.0.0/bin/install_solr_service.sh --strip-components=2
Run it against the archive:
sudo bash ./install_solr_service.sh solr-10.0.0.tgz
When it finishes, the installer starts Solr. The layout is:
| Path | Contents |
|---|---|
/opt/solr | Symlink to /opt/solr-10.0.0, the program files |
/var/solr/data | Solr home: collections and indexes |
/var/solr/logs | Log files |
/etc/default/solr.in.sh | Service settings (heap, port, host) |
Check the service:
sudo systemctl status solr
● solr.service - Apache Solr
Loaded: loaded (/etc/systemd/system/solr.service; enabled; preset: enabled)
Active: active (running) since ...
Solr 10 starts in SolrCloud mode by default, with an embedded ZooKeeper on port 9983 when no external ZooKeeper is configured. This is fine for a single server and lets you move to a real cluster later without changing how you create collections. Confirm the version and mode:
curl -s "http://localhost:8983/solr/admin/info/system" | jq '.lucene."solr-spec-version", .mode'
"10.0.0"
"solrcloud"
Step 4 - Adjusting memory and network settings
The default Java heap is 512 MB, which is too small for most indexes. Open the service settings:
sudo nano /etc/default/solr.in.sh
Uncomment and set SOLR_HEAP. A good starting point is about a quarter to half of the RAM, leaving the rest to the operating system cache that Lucene relies on:
SOLR_HEAP="1g"
By default Solr 10 only listens on 127.0.0.1, so it is not reachable from the Internet. Leave it that way; Solr should sit behind your application, not face users directly. Restart Solr to apply the heap change:
sudo systemctl restart solr
Confirm that port 8983 is bound to localhost only:
sudo ss -tlnp | grep 8983
The local address column must show 127.0.0.1:8983 (or [::ffff:127.0.0.1]:8983), not *:8983 or 0.0.0.0:8983.
To open the Solr Admin UI, create an SSH tunnel from your local machine:
ssh -L 8983:127.0.0.1:8983 your_user@your_server_ip
Then browse to http://localhost:8983/solr/. The dashboard shows the heap you just configured.
Step 5 - Creating a collection
A collection is a logical index. Create one named products. The command copies the built-in _default configset into a new configset with the same name, so changes you make later only affect this collection:
sudo -u solr /opt/solr/bin/solr create -c products
List the collections to confirm it exists:
curl -s "http://localhost:8983/solr/admin/collections?action=LIST" | jq '.collections'
[
"products"
]
Step 6 - Defining the schema
The _default configset is "schemaless": when it sees an unknown field, it guesses a type and adds it. Guesses are often wrong (a price of 10 becomes an integer field and later 10.5 fails), so turn automatic field creation off for this collection:
curl -s -X POST "http://localhost:8983/solr/products/config" \
-H 'Content-Type: application/json' \
-d '{"set-user-property": {"update.autoCreateFields": "false"}}' | jq '.responseHeader.status'
0
Now declare the fields with the Schema API. text_general fields are tokenised for full-text search, string and strings fields match exactly and are ideal for facets, pfloat and pdate are the point-based numeric and date types:
curl -s -X POST "http://localhost:8983/solr/products/schema" \
-H 'Content-Type: application/json' \
-d '{
"add-field": [
{"name": "name", "type": "text_general", "stored": true},
{"name": "description", "type": "text_general", "stored": true},
{"name": "category", "type": "string", "stored": true},
{"name": "price", "type": "pfloat", "stored": true},
{"name": "in_stock", "type": "boolean", "stored": true},
{"name": "tags", "type": "strings", "stored": true},
{"name": "created_at", "type": "pdate", "stored": true}
]
}' | jq '.responseHeader.status'
0
The _default schema already contains id as the unique key and a catch-all _text_ field that receives a copy of every field, which is what Solr searches when a query does not name a field. Check one of the new fields:
curl -s "http://localhost:8983/solr/products/schema/fields/price" | jq '.field'
{
"name": "price",
"type": "pfloat",
"stored": true
}
Step 7 - Indexing documents
Create a file with sample products:
nano products.json
[
{"id": "prod-1", "name": "Mechanical Keyboard TKL", "description": "Tenkeyless keyboard with tactile brown switches", "category": "Electronics", "price": 89.99, "in_stock": true, "tags": ["keyboard", "mechanical"], "created_at": "2026-01-15T00:00:00Z"},
{"id": "prod-2", "name": "Ergonomic Vertical Mouse", "description": "Wireless mouse designed for wrist comfort", "category": "Electronics", "price": 45.00, "in_stock": false, "tags": ["mouse", "ergonomic", "wireless"], "created_at": "2026-02-20T00:00:00Z"},
{"id": "prod-3", "name": "Wireless Keyboard and Mouse Combo", "description": "Quiet keys and a matching wireless mouse", "category": "Electronics", "price": 59.90, "in_stock": true, "tags": ["keyboard", "mouse", "wireless"], "created_at": "2026-03-05T00:00:00Z"},
{"id": "prod-4", "name": "Oak Monitor Stand", "description": "Solid wood riser with keyboard storage space", "category": "Furniture", "price": 120.00, "in_stock": true, "tags": ["desk", "stand"], "created_at": "2026-04-10T00:00:00Z"}
]
Send it to the update handler. commit=true makes the documents visible to searches immediately:
curl -s -X POST "http://localhost:8983/solr/products/update?commit=true" \
-H 'Content-Type: application/json' \
--data-binary @products.json | jq '.responseHeader.status'
0
Committing on every request is expensive under heavy indexing. The _default configset already performs a hard commit every 15 seconds for durability; add a soft commit so new documents become searchable within 5 seconds without explicit commits:
curl -s -X POST "http://localhost:8983/solr/products/config" \
-H 'Content-Type: application/json' \
-d '{"set-property": {"updateHandler.autoSoftCommit.maxTime": 5000}}' | jq '.responseHeader.status'
To delete a document, send a delete command by id (or by query, for example {"delete": {"query": "in_stock:false"}}):
curl -s -X POST "http://localhost:8983/solr/products/update?commit=true" \
-H 'Content-Type: application/json' \
-d '{"delete": {"id": "prod-4"}}' | jq '.responseHeader.status'
Step 8 - Searching, boosting and faceting
A basic query searches the catch-all field. fl limits the returned fields:
curl -s "http://localhost:8983/solr/products/select?q=keyboard&fl=id,name" | jq '.response'
The response looks similar to this (the order depends on the relevance score):
{
"numFound": 2,
"start": 0,
"numFoundExact": true,
"docs": [
{
"id": "prod-1",
"name": "Mechanical Keyboard TKL"
},
{
"id": "prod-3",
"name": "Wireless Keyboard and Mouse Combo"
}
]
}
For user-facing search, use the eDisMax query parser. It accepts plain text, searches several fields at once and lets you weight them with qf. Here a match in name counts three times more than one in description, and fq filters to products in stock without affecting the score:
curl -s -G "http://localhost:8983/solr/products/select" \
--data-urlencode "q=wireless mouse" \
--data-urlencode "defType=edismax" \
--data-urlencode "qf=name^3 description tags^2" \
--data-urlencode "fq=in_stock:true" \
--data-urlencode "fl=id,name,price,score" \
| jq '.response.docs'
Only prod-3 is returned: prod-2 also matches but is out of stock. Filter queries are cached separately from the main query, which makes repeated filters cheap.
Facets count matching documents per value, which is what powers the filter sidebar of a shop. This request counts products per category and tag, and groups prices into ranges of 50:
curl -s -G "http://localhost:8983/solr/products/select" \
--data-urlencode "q=*:*" \
--data-urlencode "rows=0" \
--data-urlencode "facet=true" \
--data-urlencode "facet.mincount=1" \
--data-urlencode "facet.field=category" \
--data-urlencode "facet.field=tags" \
--data-urlencode "facet.range=price" \
--data-urlencode "facet.range.start=0" \
--data-urlencode "facet.range.end=150" \
--data-urlencode "facet.range.gap=50" \
| jq '.facet_counts.facet_fields.category, .facet_counts.facet_ranges.price.counts'
[
"Electronics",
3
]
[
"0.0",
1,
"50.0",
2
]
Solr returns facet values as a flat list of value and count pairs; facet.mincount=1 hides values and ranges with no matches. Finally, highlighting marks the matched words in each result:
curl -s -G "http://localhost:8983/solr/products/select" \
--data-urlencode "q=wrist" \
--data-urlencode "defType=edismax" \
--data-urlencode "qf=name description" \
--data-urlencode "hl=true" \
--data-urlencode "hl.fl=description" \
| jq '.highlighting'
{
"prod-2": {
"description": [
"Wireless mouse designed for <em>wrist</em> comfort"
]
}
}
Troubleshooting
- Solr does not start: read
sudo journalctl -u solr -n 50andsudo tail -n 100 /var/solr/logs/solr.log. A heap larger than the free RAM and a Java version older than 21 are the usual causes. undefined fielderrors when indexing: the document contains a field you did not declare, and automatic field creation is off. Add the field with the Schema API or remove it from the document.- Documents are indexed but not found: they have not been committed yet. Send
commit=trueor wait for the soft commit interval. OutOfMemoryErrorin the log: raiseSOLR_HEAPin/etc/default/solr.in.sh, restart Solr, and make sure the server still has free memory for the operating system cache.- The Admin UI is unreachable: Solr listens on localhost only by design. Use the SSH tunnel from Step 4 instead of opening port 8983 in the firewall.
Conclusion
You installed Solr 10 on Ubuntu 24.04 as a systemd service, created a collection with an explicit schema, indexed JSON documents and ran boosted, filtered, faceted and highlighted searches. Before exposing Solr to other hosts, enable the Basic Authentication plugin and TLS. From here you can tune analyzers for your language in the collection's configset, add backups with the Collections API BACKUP action, and grow into a multi-node SolrCloud cluster with an external three-node ZooKeeper ensemble.
