TensorFlow Serving is Google's production server for TensorFlow models. It loads models in the SavedModel format, exposes them through a REST API and a gRPC API, and picks up new model versions from disk without a restart. In this tutorial you will run TensorFlow Serving in Docker on Ubuntu 24.04, export a small model, send predictions to it, roll out a second version and serve several models from one container.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS (x86_64), for example a CubePath VPS, with at least 2 GB of RAM. Real models usually need 8 GB or more.
- A non-root user with
sudoprivileges. - Docker Engine installed from Docker's official repository, with your user in the
dockergroup. - Optional, for GPU inference: an NVIDIA GPU with the driver and the NVIDIA Container Toolkit installed.
The official image is the supported way to run TensorFlow Serving today. It bundles the correct tensorflow_model_server build, so you do not need to manage the apt repository or match CPU instruction sets by hand.
Step 1 - Pulling the TensorFlow Serving image
Download the CPU image from Docker Hub:
docker pull tensorflow/serving:latest
The image's entrypoint starts the server with preset ports and paths. To check the bundled version, override the entrypoint and call the binary directly:
docker run --rm --entrypoint tensorflow_model_server tensorflow/serving:latest --version
TensorFlow ModelServer: 2.x.x-rc0+dev.sha.xxxxxxx
TensorFlow Library: 2.x.x
Step 2 - Exporting a model in the SavedModel format
TensorFlow Serving expects one directory per model, with one numbered subdirectory per version:
| Path | Contents |
|---|---|
/opt/models/half_plus_two/1/ | Version 1 (saved_model.pb and variables/) |
/opt/models/half_plus_two/2/ | Version 2, loaded automatically when it appears |
Create the model directory and give your user ownership of it:
sudo mkdir -p /opt/models
sudo chown "$USER":"$USER" /opt/models
You need TensorFlow on the machine that exports the model. Install it in a Python virtual environment so it does not interfere with system packages:
sudo apt update
sudo apt install -y python3-venv
python3 -m venv ~/tf-env
source ~/tf-env/bin/activate
pip install --upgrade pip
pip install tensorflow
Create a small export script. The model computes y = 0.5 * x + 2 and declares an explicit signature named serving_default, which is what TensorFlow Serving calls by default:
nano ~/export_model.py
import sys
import tensorflow as tf
class HalfPlusTwo(tf.Module):
def __init__(self, offset):
super().__init__()
self.offset = tf.Variable(offset, dtype=tf.float32)
@tf.function(input_signature=[tf.TensorSpec(shape=[None, 1], dtype=tf.float32, name="x")])
def __call__(self, x):
return {"y": 0.5 * x + self.offset}
version = sys.argv[1]
offset = float(sys.argv[2])
module = HalfPlusTwo(offset)
tf.saved_model.save(
module,
f"/opt/models/half_plus_two/{version}",
signatures={"serving_default": module.__call__},
)
print(f"Exported version {version}")
Export version 1 with an offset of 2:
python ~/export_model.py 1 2.0
For a Keras 3 model, the equivalent call is model.export("/opt/models/my_model/1"), which writes a SavedModel with a serving_default signature.
Inspect the signature with saved_model_cli, which is installed with TensorFlow. You need the input and output names later for gRPC requests:
saved_model_cli show --dir /opt/models/half_plus_two/1 --tag_set serve --signature_def serving_default
The given SavedModel SignatureDef contains the following input(s):
inputs['x'] tensor_info:
dtype: DT_FLOAT
shape: (-1, 1)
name: serving_default_x:0
The given SavedModel SignatureDef contains the following output(s):
outputs['y'] tensor_info:
dtype: DT_FLOAT
shape: (-1, 1)
name: StatefulPartitionedCall:0
Method name is: tensorflow/serving/predict
Step 3 - Starting the model server
Run the container with the model directory mounted under /models. The MODEL_NAME variable tells the entrypoint which model to load, port 8501 is the REST API and port 8500 is gRPC:
docker run -d \
--name tf-serving \
--restart unless-stopped \
-p 127.0.0.1:8501:8501 \
-p 127.0.0.1:8500:8500 \
-v /opt/models/half_plus_two:/models/half_plus_two:ro \
-e MODEL_NAME=half_plus_two \
tensorflow/serving:latest
The ports are published on 127.0.0.1 only. Docker writes its own iptables rules and bypasses UFW, so a port published on all interfaces would be reachable from the Internet even with a firewall enabled. Keep the API private and put an application or a reverse proxy in front of it.
Check the logs:
docker logs tf-serving
... Successfully loaded servable version {name: half_plus_two version: 1}
... Running gRPC ModelServer at 0.0.0.0:8500 ...
... Exporting HTTP/REST API at:localhost:8501 ...
Ask the server for the model status:
curl http://localhost:8501/v1/models/half_plus_two
{
"model_version_status": [
{
"version": "1",
"state": "AVAILABLE",
"status": {
"error_code": "OK",
"error_message": ""
}
}
]
}
Step 4 - Sending predictions to the REST API
The :predict endpoint accepts a JSON body with an instances list, one entry per input row:
curl -s -X POST http://localhost:8501/v1/models/half_plus_two:predict \
-H "Content-Type: application/json" \
-d '{"instances": [[1.0], [2.0], [5.0]]}'
{
"predictions": [[2.5], [3.0], [4.5]
]
}
To query a specific version instead of the latest one, add versions/<n> to the path:
curl -s -X POST http://localhost:8501/v1/models/half_plus_two/versions/1:predict \
-H "Content-Type: application/json" \
-d '{"instances": [[10.0]]}'
The same request from Python only needs the requests library:
import requests
resp = requests.post(
"http://localhost:8501/v1/models/half_plus_two:predict",
json={"instances": [[1.0], [2.0]]},
timeout=5,
)
resp.raise_for_status()
print(resp.json()["predictions"])
Step 5 - Using the gRPC API
gRPC sends tensors in binary form, which is faster than JSON for large inputs such as images. Install the client package in the same virtual environment:
pip install tensorflow-serving-api
Create a client that uses the input name x and the output name y reported by saved_model_cli:
nano ~/grpc_client.py
import grpc
import numpy as np
import tensorflow as tf
from tensorflow_serving.apis import predict_pb2, prediction_service_pb2_grpc
channel = grpc.insecure_channel("localhost:8500")
stub = prediction_service_pb2_grpc.PredictionServiceStub(channel)
request = predict_pb2.PredictRequest()
request.model_spec.name = "half_plus_two"
request.model_spec.signature_name = "serving_default"
request.inputs["x"].CopyFrom(
tf.make_tensor_proto(np.array([[1.0], [2.0]], dtype=np.float32))
)
response = stub.Predict(request, timeout=5.0)
print(tf.make_ndarray(response.outputs["y"]))
Run it:
python ~/grpc_client.py
[[2.5]
[3. ]]
Step 6 - Rolling out a new model version
TensorFlow Serving polls the model directory and, by default, serves only the highest version number. Export version 2 with an offset of 3:
python ~/export_model.py 2 3.0
Within a few seconds the server loads version 2 and unloads version 1. Confirm it in the logs and with a prediction:
docker logs --tail 5 tf-serving
curl -s -X POST http://localhost:8501/v1/models/half_plus_two:predict \
-H "Content-Type: application/json" \
-d '{"instances": [[1.0]]}'
{
"predictions": [[3.5]
]
}
To roll back, delete or move the 2/ directory and the server goes back to version 1. Always write a new version to a temporary directory first and rename it into place when the export is complete, so the server never sees a half written model.
Step 7 - Serving several models with a config file
To host more than one model in a single server, describe them in a model config file. Create it in /opt/models:
nano /opt/models/models.config
model_config_list {
config {
name: "half_plus_two"
base_path: "/models/half_plus_two"
model_platform: "tensorflow"
model_version_policy {
specific {
versions: 1
versions: 2
}
}
}
config {
name: "half_plus_three"
base_path: "/models/half_plus_three"
model_platform: "tensorflow"
}
}
The specific policy keeps both versions of half_plus_two loaded so clients can pick one by URL. The second model uses the default policy (latest version only). Export the second model, reusing the script with a different path:
sed 's#half_plus_two#half_plus_three#' ~/export_model.py > ~/export_model3.py
python ~/export_model3.py 1 3.0
Replace the container, this time mounting the whole /opt/models directory and passing the config file. Extra arguments after the image name are appended to the server command line:
docker rm -f tf-serving
docker run -d \
--name tf-serving \
--restart unless-stopped \
-p 127.0.0.1:8501:8501 \
-p 127.0.0.1:8500:8500 \
-v /opt/models:/models:ro \
tensorflow/serving:latest \
--model_config_file=/models/models.config \
--model_config_file_poll_wait_seconds=60
With --model_config_file_poll_wait_seconds=60 the server rereads the config file every minute, so you can add or remove models without restarting the container. Verify both models:
curl -s http://localhost:8501/v1/models/half_plus_two/versions/1
curl -s http://localhost:8501/v1/models/half_plus_three
Each command returns a model_version_status block with "state": "AVAILABLE".
Step 8 - Enabling GPU inference (optional)
On a server with an NVIDIA GPU, the driver and the NVIDIA Container Toolkit, use the latest-gpu image and pass the GPU to the container:
docker rm -f tf-serving
docker run -d \
--name tf-serving \
--restart unless-stopped \
--gpus all \
-p 127.0.0.1:8501:8501 \
-p 127.0.0.1:8500:8500 \
-v /opt/models:/models:ro \
tensorflow/serving:latest-gpu \
--model_config_file=/models/models.config
Confirm that the process uses the GPU while you send requests:
nvidia-smi
The tensorflow_model_server process appears in the process list with the memory it has reserved. To pin the container to one card, use --gpus '"device=0"' instead of --gpus all.
Troubleshooting
The model stays in state LOADING or the server logs No versions of servable ... found. The mounted directory must contain numbered subdirectories directly. Check the layout:
find /opt/models -name saved_model.pb
Each result must look like /opt/models/<model>/<number>/saved_model.pb.
Serving signature name: "serving_default" not found in signature def. The model was exported without a default signature. Re-export it with an explicit signatures={"serving_default": ...} argument or with model.export() for Keras models, then check again with saved_model_cli show.
Connection refused from another host. The container publishes the ports on 127.0.0.1 on purpose. Call the API from the same server, over a private network, or through a reverse proxy with authentication.
The first request is slow. TensorFlow initializes kernels on the first call. You can ship warmup requests in assets.extra/tf_serving_warmup_requests inside each version directory so the server runs them before marking the version as available.
Conclusion
You now have TensorFlow Serving running in Docker on Ubuntu 24.04, answering REST and gRPC requests, rolling out new versions from disk and hosting several models from one config file. Next, consider adding request batching with --enable_batching, placing an authenticated reverse proxy such as Nginx in front of the REST port, and automating model exports from your training pipeline into /opt/models.
