Apache Spark is an open source engine for large-scale data processing. It splits work into tasks that run in parallel across the CPU cores of one or many machines, and offers APIs for Python (PySpark), SQL, Scala and Java. In this tutorial you will install Spark 4 on Ubuntu 24.04, run its built-in standalone cluster manager (a master and a worker) as systemd services, submit PySpark jobs, and keep a record of finished jobs with the History Server. The last section shows how to add more worker nodes.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS, with at least 4 GB of RAM and 2 vCPUs. Spark keeps data in memory, so more RAM translates directly into larger jobs.
- A non-root user with
sudoprivileges. - UFW enabled with only SSH allowed. The Spark web UIs have no authentication, so this guide keeps them closed and reaches them through an SSH tunnel.
Step 1 - Installing Java
Spark 4 runs on Java 17 or 21. Install the OpenJDK 17 runtime from the Ubuntu repositories, together with Python 3 for PySpark:
sudo apt update
sudo apt install openjdk-17-jre-headless python3
Verify the Java version:
java -version
openjdk version "17.0.16" 2025-07-15
OpenJDK Runtime Environment (build 17.0.16+8-Ubuntu-0ubuntu124.04.1)
OpenJDK 64-Bit Server VM (build 17.0.16+8-Ubuntu-0ubuntu124.04.1, mixed mode, sharing)
The exact build number will differ.
Step 2 - Downloading and installing Spark
Spark is distributed as a prebuilt archive. Check the latest release on the Apache Spark downloads page and set it in a variable; this guide uses 4.0.1 with the package prebuilt for Hadoop 3:
SPARK_VERSION=4.0.1
cd /tmp
curl -fLO "https://archive.apache.org/dist/spark/spark-${SPARK_VERSION}/spark-${SPARK_VERSION}-bin-hadoop3.tgz"
curl -fLO "https://archive.apache.org/dist/spark/spark-${SPARK_VERSION}/spark-${SPARK_VERSION}-bin-hadoop3.tgz.sha512"
Compare the checksum of the archive with the published one. Both hashes must be identical:
sha512sum "spark-${SPARK_VERSION}-bin-hadoop3.tgz"
cat "spark-${SPARK_VERSION}-bin-hadoop3.tgz.sha512"
Extract the archive into /opt and create a version-independent symlink, which makes future upgrades a matter of pointing the link at a new directory:
sudo tar -xzf "spark-${SPARK_VERSION}-bin-hadoop3.tgz" -C /opt
sudo ln -sfn "/opt/spark-${SPARK_VERSION}-bin-hadoop3" /opt/spark
Create a system user to run the Spark daemons, and a directory for the worker's scratch space and the event logs. The event log directory is group writable so that jobs submitted by members of the spark group can write to it:
sudo useradd --system --home-dir /var/lib/spark --shell /usr/sbin/nologin spark
sudo mkdir -p /var/lib/spark/work /var/lib/spark/events
sudo chown -R spark:spark /var/lib/spark
sudo chmod 2775 /var/lib/spark/events
sudo usermod -aG spark "$USER"
Log out and back in so the new group membership applies to your session.
Add Spark to the PATH for all users:
sudo nano /etc/profile.d/spark.sh
export SPARK_HOME=/opt/spark
export PATH="$PATH:$SPARK_HOME/bin"
Load it in the current shell and check the installation:
source /etc/profile.d/spark.sh
spark-submit --version
Welcome to
____ __
/ __/__ ___ _____/ /__
_\ \/ _ \/ _ `/ __/ '_/
/___/ .__/\_,_/_/ /_/\_\ version 4.0.1
/_/
Step 3 - Configuring the standalone cluster
The standalone cluster has two roles. The master accepts applications and assigns resources; each worker offers CPU cores and memory and launches executors that run the tasks. On a single server both run side by side.
Create spark-env.sh, which the daemon scripts read at startup:
sudo nano /opt/spark/conf/spark-env.sh
JAVA_HOME=/usr/lib/jvm/java-17-openjdk-amd64
SPARK_MASTER_HOST=127.0.0.1
SPARK_WORKER_CORES=2
SPARK_WORKER_MEMORY=3g
SPARK_WORKER_DIR=/var/lib/spark/work
SPARK_MASTER_HOSTis the address the master listens on.127.0.0.1is enough for a single server; the last section changes it for multi-node setups.SPARK_WORKER_CORESandSPARK_WORKER_MEMORYare what the worker offers to applications. Leave about 1 GB of RAM for the operating system and the daemons themselves.- On an arm64 server the Java path is
/usr/lib/jvm/java-17-openjdk-arm64.
Next, create spark-defaults.conf, which sets defaults for every application you submit:
sudo nano /opt/spark/conf/spark-defaults.conf
spark.master spark://127.0.0.1:7077
spark.eventLog.enabled true
spark.eventLog.dir file:///var/lib/spark/events
spark.history.fs.logDirectory file:///var/lib/spark/events
spark.executor.memory 2g
spark.driver.memory 1g
With spark.master set here, you no longer need --master on every spark-submit call. The event log settings make each application write a record that the History Server reads in Step 6.
Step 4 - Running the master and worker with systemd
The scripts in /opt/spark/sbin start the daemons in the background by default. Setting SPARK_NO_DAEMONIZE=true keeps them in the foreground, which is what systemd expects, and sends their logs to the journal.
Create the master unit:
sudo nano /etc/systemd/system/spark-master.service
[Unit]
Description=Apache Spark standalone master
After=network-online.target
Wants=network-online.target
[Service]
User=spark
Group=spark
Environment=SPARK_NO_DAEMONIZE=true
ExecStart=/opt/spark/sbin/start-master.sh
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
Create the worker unit. It connects to the master URL passed as an argument:
sudo nano /etc/systemd/system/spark-worker.service
[Unit]
Description=Apache Spark standalone worker
After=spark-master.service
Wants=spark-master.service
[Service]
User=spark
Group=spark
Environment=SPARK_NO_DAEMONIZE=true
ExecStart=/opt/spark/sbin/start-worker.sh spark://127.0.0.1:7077
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
Reload systemd and start both services:
sudo systemctl daemon-reload
sudo systemctl enable --now spark-master spark-worker
Check that the worker registered with the master:
sudo journalctl -u spark-master --no-pager | grep -i "registering worker"
INFO Master: Registering worker 127.0.0.1:41235 with 2 cores, 3.0 GiB RAM
The master also exposes its state as JSON on port 8080, which you can query locally:
curl -s http://127.0.0.1:8080/json/ | python3 -m json.tool | grep -E '"(status|aliveworkers|cores|memory)"'
"aliveworkers": 1,
"cores": 2,
"memory": 3072,
"status": "ALIVE"
Step 5 - Submitting PySpark jobs
Start with the Pi estimation example that ships with Spark. The argument 100 is the number of partitions, which means 100 parallel tasks:
spark-submit /opt/spark/examples/src/main/python/pi.py 100 2>/dev/null
Pi is roughly 3.141640
Redirecting standard error hides Spark's own log lines; remove 2>/dev/null when you need to debug a job.
Now write a job of your own that reads a CSV file and aggregates it with the DataFrame API and with SQL. The executors run as the spark user and read input files themselves, so the data must live in a directory that user can access. Home directories on Ubuntu 24.04 are private (mode 750), so create a shared directory owned by you and the spark group:
sudo install -d -o "$USER" -g spark -m 2775 /srv/spark-demo
Create a sample dataset:
cat > /srv/spark-demo/sales.csv <<'EOF'
product,region,amount
vps,eu,20
vps,us,25
baremetal,eu,180
vps,eu,20
baremetal,us,200
kubernetes,eu,60
EOF
Create the job:
nano /srv/spark-demo/sales_report.py
from pyspark.sql import SparkSession
from pyspark.sql import functions as F
spark = SparkSession.builder.appName("SalesReport").getOrCreate()
spark.sparkContext.setLogLevel("WARN")
df = (
spark.read.option("header", True)
.option("inferSchema", True)
.csv("file:///srv/spark-demo/sales.csv")
)
# DataFrame API
by_product = (
df.groupBy("product")
.agg(F.count("*").alias("orders"), F.sum("amount").alias("revenue"))
.orderBy(F.desc("revenue"))
)
by_product.show()
# The same data through SQL
df.createOrReplaceTempView("sales")
spark.sql(
"SELECT region, SUM(amount) AS revenue FROM sales GROUP BY region ORDER BY region"
).show()
spark.stop()
Submit it:
spark-submit /srv/spark-demo/sales_report.py 2>/dev/null
+----------+------+-------+
| product|orders|revenue|
+----------+------+-------+
| baremetal| 2| 380|
| vps| 3| 65|
|kubernetes| 1| 60|
+----------+------+-------+
+------+-------+
|region|revenue|
+------+-------+
| eu| 280|
| us| 225|
+------+-------+
A file:// path works here because the driver and the only worker share one disk. On a multi-node cluster every executor needs to reach the same data, so jobs read from and write to shared storage (HDFS, S3-compatible object storage or NFS), for example with df.write.parquet("s3a://bucket/path").
For interactive exploration, pyspark opens a Python shell with a ready spark session connected to the cluster:
pyspark
Step 6 - Running the History Server
The application UI on port 4040 only exists while a job is running. The History Server rebuilds the UI of finished applications from the event logs configured in Step 3. Create a unit for it:
sudo nano /etc/systemd/system/spark-history.service
[Unit]
Description=Apache Spark History Server
After=network-online.target
[Service]
User=spark
Group=spark
Environment=SPARK_NO_DAEMONIZE=true
ExecStart=/opt/spark/sbin/start-history-server.sh
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now spark-history
Confirm that it found the jobs from Step 5:
curl -s http://127.0.0.1:18080/api/v1/applications | python3 -m json.tool | grep '"name"'
"name": "SalesReport",
"name": "PythonPi",
To browse the web UIs from your workstation without opening any port, create an SSH tunnel:
ssh -L 8080:127.0.0.1:8080 -L 18080:127.0.0.1:18080 your_user@your_server_ip
Then open http://localhost:8080 for the master UI (workers, running and completed applications) and http://localhost:18080 for the History Server (stages, tasks, executors and SQL plans of each finished job).
Step 7 - Adding more worker nodes
To grow the cluster, prepare each additional server with Steps 1 and 2 and connect them over a private network, which keeps Spark's unauthenticated ports off the internet.
On the master, set SPARK_MASTER_HOST in /opt/spark/conf/spark-env.sh to the master's private_ip, update spark.master in spark-defaults.conf and the URL in spark-worker.service to spark://private_ip:7077, then restart:
sudo systemctl daemon-reload
sudo systemctl restart spark-master spark-worker
Spark nodes talk to each other on several ports, some of them random, so allow all traffic from the private subnet on every node. Replace 10.0.0.0/24 with your private range:
sudo ufw allow from 10.0.0.0/24
On each new worker, copy spark-env.sh without the SPARK_MASTER_HOST line (adjust cores and memory to that machine) and create spark-worker.service pointing to spark://private_ip:7077. After enabling it, the master's JSON endpoint shows the new total of aliveworkers.
Troubleshooting
The worker does not register with the master. Read sudo journalctl -u spark-worker -n 50. Connection refused means the master URL in the unit does not match SPARK_MASTER_HOST; the host in both must be identical, including 127.0.0.1 versus a hostname.
The job waits with "Initial job has not accepted any resources". The application asks for more than a worker can offer. Make sure spark.executor.memory is lower than SPARK_WORKER_MEMORY, or pass a smaller value with spark-submit --executor-memory 1g.
Executors fail with java.lang.OutOfMemoryError. Give executors more memory, or split the data into more partitions with spark.sql.shuffle.partitions so each task handles less data at a time.
Conclusion
You now have Apache Spark 4 running as a standalone cluster on Ubuntu 24.04, with systemd managing the master, worker and History Server, and PySpark jobs submitted with spark-submit. Next, connect Spark to your real data by adding the S3A or JDBC connectors with --packages, schedule recurring jobs with a workflow tool such as Apache Airflow, and enable Spark's authentication and encryption options (spark.authenticate, spark.network.crypto.enabled) before extending the cluster beyond a trusted private network.
