MinIO is an S3-compatible object storage server that you can run on your own hardware. For machine learning work it gives you one place to keep raw datasets, trained models and experiment artifacts, reachable from any tool that speaks the S3 API. In this tutorial you will install MinIO on Ubuntu 24.04 as a systemd service, create buckets with versioning for models, give your applications their own credentials, and read and write objects from Python with boto3 and from DVC.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS (x86_64), for example a CubePath VPS, with at least 2 GB of RAM.
- A non-root user with
sudoprivileges. - Enough disk for your data. Ideally mount a dedicated volume at
/data; MinIO recommends XFS for its data drives. - UFW enabled with
OpenSSHallowed. - Python 3 on the machine that will run your training code.
NoteIn 2025 MinIO changed how it distributes its community edition: the web console was reduced to an object browser, and binary builds are no longer published on a regular schedule. The server works as shown here, but check the project's GitHub README for the current distribution status before deploying. All administration in this guide uses the
mccommand line client, which does not depend on the console.
Step 1 - Installing the MinIO server and client
Download the server binary and install it into /usr/local/bin:
wget https://dl.min.io/server/minio/release/linux-amd64/minio -O /tmp/minio
sudo install -m 0755 /tmp/minio /usr/local/bin/minio
Download the mc client the same way:
wget https://dl.min.io/client/mc/release/linux-amd64/mc -O /tmp/mc
sudo install -m 0755 /tmp/mc /usr/local/bin/mc
Verify both:
minio --version
mc --version
minio version RELEASE.2025-xx-xxTxx-xx-xxZ ...
mc version RELEASE.2025-xx-xxTxx-xx-xxZ ...
Step 2 - Creating the user, data directory and configuration
MinIO should run as an unprivileged system user that owns only its data directory:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin minio-user
sudo mkdir -p /data/minio
sudo chown -R minio-user:minio-user /data/minio
The systemd unit you will create reads its settings from /etc/default/minio. Create it:
sudo nano /etc/default/minio
MINIO_ROOT_USER=minioadmin
MINIO_ROOT_PASSWORD=your_strong_password
MINIO_VOLUMES="/data/minio"
MINIO_OPTS="--console-address :9001"
Replace your_strong_password with a random value of at least 16 characters, for example the output of openssl rand -base64 24. The root account is for administration only; applications get their own keys in step 5. Restrict the file, since it contains the root password:
sudo chmod 600 /etc/default/minio
Step 3 - Running MinIO as a systemd service
Create the unit file:
sudo nano /etc/systemd/system/minio.service
[Unit]
Description=MinIO object storage
Wants=network-online.target
After=network-online.target
AssertFileIsExecutable=/usr/local/bin/minio
[Service]
Type=notify
User=minio-user
Group=minio-user
WorkingDirectory=/usr/local
EnvironmentFile=/etc/default/minio
ExecStart=/usr/local/bin/minio server $MINIO_OPTS $MINIO_VOLUMES
Restart=always
LimitNOFILE=1048576
TasksMax=infinity
TimeoutStopSec=infinity
SendSIGKILL=no
[Install]
WantedBy=multi-user.target
Type=notify lets systemd wait until MinIO reports it is ready, and the high LimitNOFILE avoids running out of file descriptors with many small objects such as image datasets. Start the service:
sudo systemctl daemon-reload
sudo systemctl enable --now minio
sudo systemctl status minio --no-pager
● minio.service - MinIO object storage
Loaded: loaded (/etc/systemd/system/minio.service; enabled; preset: enabled)
Active: active (running) since ...
The S3 API listens on port 9000 and the console on 9001. Check that the API is healthy:
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:9000/minio/health/live
200
Step 4 - Configuring mc and creating ML buckets
Register the local server in mc under the alias local, using the root credentials from /etc/default/minio:
mc alias set local http://127.0.0.1:9000 minioadmin your_strong_password
mc admin info local
● 127.0.0.1:9000
Uptime: 2 minutes
Version: ...
Network: 1/1 OK
Drives: 1/1 OK
Separate buckets by lifecycle, because versioning and retention rules apply per bucket. A practical layout for ML work is:
| Bucket | Contents | Versioning |
|---|---|---|
ml-datasets | Raw and processed datasets, tracked by DVC | Off (DVC keeps history) |
ml-models | Exported models, by name and stage | On |
ml-experiments | Checkpoints, logs and metrics from runs | Off, with expiry |
Create the buckets:
mc mb local/ml-datasets local/ml-models local/ml-experiments
Enable versioning on the models bucket so that overwriting or deleting a model never loses the previous file:
mc version enable local/ml-models
mc version info local/ml-models
local/ml-models versioning is enabled
Old versions still use disk space. Add a lifecycle rule that removes non-current versions after 90 days, and one that expires experiment artifacts after 30 days:
mc ilm rule add local/ml-models --noncurrent-expire-days 90
mc ilm rule add local/ml-experiments --expire-days 30
mc ilm rule ls local/ml-models
Test versioning by uploading a file twice to the same key:
echo "v1" > /tmp/model.txt
mc cp /tmp/model.txt local/ml-models/classifier/production/model.txt
echo "v2" > /tmp/model.txt
mc cp /tmp/model.txt local/ml-models/classifier/production/model.txt
mc ls --versions local/ml-models/classifier/production/
[2026-09-25 10:02:11 UTC] 3B STANDARD 8b1f... v2 PUT model.txt
[2026-09-25 10:01:58 UTC] 3B STANDARD 3c7a... v1 PUT model.txt
To restore an older version, copy it by its version ID:
mc cp --version-id VERSION_ID local/ml-models/classifier/production/model.txt ./model-restored.txt
Step 5 - Creating credentials for applications
Training jobs and services should not use the root account. Create a user with a scoped policy instead. First write a policy that allows reading and writing only the three ML buckets:
nano /tmp/ml-readwrite.json
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:ListBucket", "s3:GetBucketLocation"],
"Resource": [
"arn:aws:s3:::ml-datasets",
"arn:aws:s3:::ml-models",
"arn:aws:s3:::ml-experiments"
]
},
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
"Resource": [
"arn:aws:s3:::ml-datasets/*",
"arn:aws:s3:::ml-models/*",
"arn:aws:s3:::ml-experiments/*"
]
}
]
}
Load the policy, create the user and attach the policy. Replace app_secret_key with a random string of at least 8 characters:
mc admin policy create local ml-readwrite /tmp/ml-readwrite.json
mc admin user add local ml-app app_secret_key
mc admin policy attach local ml-readwrite --user ml-app
Verify the user and its policy:
mc admin user info local ml-app
AccessKey: ml-app
Status: enabled
PolicyName: ml-readwrite
...
Step 6 - Opening the API to your training machines
Only the machines that run training or inference need the S3 API. Allow port 9000 from their IP addresses rather than from everywhere, replacing your_client_ip:
sudo ufw allow from your_client_ip to any port 9000 proto tcp
Leave port 9001 closed and reach the console through an SSH tunnel from your workstation when you need it:
ssh -L 9001:127.0.0.1:9001 your_user@your_server_ip
Then open http://localhost:9001 in your browser. If the storage must be reachable over the Internet, put Nginx with a TLS certificate in front of port 9000 and use https endpoints in your clients, since S3 credentials and data otherwise travel in clear text.
Step 7 - Accessing MinIO from Python with boto3
On the client machine, install boto3 in a virtual environment:
python3 -m venv ~/ml-env
source ~/ml-env/bin/activate
pip install boto3
Create a script that uploads a model file, lists the models and generates a temporary download link:
nano ~/minio_models.py
import os
import boto3
from botocore.config import Config
s3 = boto3.client(
"s3",
endpoint_url=os.environ["S3_ENDPOINT"],
aws_access_key_id=os.environ["S3_ACCESS_KEY"],
aws_secret_access_key=os.environ["S3_SECRET_KEY"],
region_name="us-east-1",
config=Config(signature_version="s3v4", s3={"addressing_style": "path"}),
)
# Upload a trained model; boto3 switches to multipart upload for large files
with open("model.onnx", "wb") as f:
f.write(os.urandom(1024))
s3.upload_file("model.onnx", "ml-models", "classifier/v3/model.onnx")
# List everything under the classifier prefix
resp = s3.list_objects_v2(Bucket="ml-models", Prefix="classifier/")
for obj in resp.get("Contents", []):
print(obj["Key"], obj["Size"])
# Temporary link, valid for one hour, for a service that should not hold credentials
url = s3.generate_presigned_url(
"get_object",
Params={"Bucket": "ml-models", "Key": "classifier/v3/model.onnx"},
ExpiresIn=3600,
)
print(url)
The model.onnx written here is placeholder data; in a real pipeline you upload the file your training job produced. Path-style addressing is used because MinIO serves buckets under the endpoint path rather than as subdomains. Run the script with the application credentials from step 5:
export S3_ENDPOINT=http://your_server_ip:9000
export S3_ACCESS_KEY=ml-app
export S3_SECRET_KEY=app_secret_key
python ~/minio_models.py
classifier/production/model.txt 3
classifier/v3/model.onnx 1024
http://your_server_ip:9000/ml-models/classifier/v3/model.onnx?X-Amz-Algorithm=AWS4-HMAC-SHA256&...
Most ML libraries that read from S3, such as pandas with s3fs or PyTorch data loaders built on boto3, accept the same endpoint URL and keys.
Step 8 - Tracking datasets with DVC
DVC versions large datasets alongside your Git repository by storing the data in a remote and committing only small pointer files. Install it with S3 support in the same environment:
pip install "dvc[s3]"
Inside an existing Git repository for your project, initialize DVC and configure MinIO as the default remote. The --local flag stores the keys in .dvc/config.local, which is not committed:
dvc init
dvc remote add -d minio s3://ml-datasets/dvc
dvc remote modify minio endpointurl http://your_server_ip:9000
dvc remote modify --local minio access_key_id ml-app
dvc remote modify --local minio secret_access_key app_secret_key
Track a dataset directory, commit the pointer file and push the data:
dvc add data/train
git add data/train.dvc data/.gitignore .dvc/config
git commit -m "Track training data with DVC"
dvc push
Verify that the objects arrived in the bucket:
mc ls --recursive local/ml-datasets/dvc | head -n 3
On another machine, git clone the repository, set the same local credentials and run dvc pull to download exactly the dataset version referenced by the current commit.
Troubleshooting
The service fails to start and journalctl -u minio shows a permission error on /data/minio. The data directory is not owned by minio-user. Run sudo chown -R minio-user:minio-user /data/minio and restart with sudo systemctl restart minio.
SignatureDoesNotMatch from boto3 or DVC. The secret key is wrong, or the endpoint URL used by the client differs from the one it signs for, which happens behind a proxy that rewrites the Host header. Check the keys with mc alias set and make sure your proxy passes Host unchanged.
AccessDenied for the application user. The bucket or action is not in the policy. Review it with mc admin policy info local ml-readwrite and remember that listing needs s3:ListBucket on the bucket ARN, while reads and writes need object ARNs ending in /*.
Clients get Connection refused or time out. Confirm MinIO is listening with sudo ss -tlnp | grep 9000 and that UFW allows the client's IP with sudo ufw status.
Conclusion
MinIO is now running on Ubuntu 24.04 as S3-compatible storage for your ML work, with versioned model storage, lifecycle rules that control disk usage, restricted application credentials and working access from boto3 and DVC. Next, consider placing Nginx with TLS in front of the API, replicating the ml-models bucket to a second server with mc mirror for backups, and using the bucket as the artifact store for an experiment tracker such as MLflow.
