Envoy is an open source Layer 7 proxy, originally built at Lyft and now a graduated CNCF project, that powers service meshes such as Istio and many API gateways. It can also run on its own as an edge proxy in front of your applications. In this tutorial you will install the official Envoy binary on Ubuntu 24.04, run it as a systemd service, and build a static configuration that load balances two backends with active health checks, retries, outlier detection and local rate limiting.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS (x86_64 or arm64), for example a CubePath VPS, with at least 1 GB of RAM.
- A non-root user with
sudoprivileges. - Port 80 open if you want to reach Envoy from outside the server.
You will run two small demo backends on the same server with Python, which is installed by default on Ubuntu 24.04. Replace them with your own services when you move to production.
Step 1 - Installing the Envoy binary
The Envoy project publishes statically linked binaries for each release on GitHub, together with a signed checksum file. This guide uses version 1.39.1; check the releases page for the latest version and adjust the variable.
Download the binary and the checksum file:
ENVOY_VERSION=1.39.1
cd /tmp
curl -fLO "https://github.com/envoyproxy/envoy/releases/download/v${ENVOY_VERSION}/envoy-${ENVOY_VERSION}-linux-x86_64"
curl -fLO "https://github.com/envoyproxy/envoy/releases/download/v${ENVOY_VERSION}/checksums.txt.asc"
On an arm64 server, replace linux-x86_64 with linux-aarch_64 in this step.
Verify the download against the published SHA-256 checksum:
grep "envoy-${ENVOY_VERSION}-linux-x86_64$" checksums.txt.asc | awk '{print $1" envoy-'"${ENVOY_VERSION}"'-linux-x86_64"}' | sha256sum -c
envoy-1.39.1-linux-x86_64: OK
Install it to /usr/local/bin and check the version:
sudo install -m 755 "envoy-${ENVOY_VERSION}-linux-x86_64" /usr/local/bin/envoy
envoy --version
envoy version: 5c8f3d6c.../1.39.1/Clean/RELEASE/BoringSSL
Step 2 - Creating a user and directories
Envoy does not need root privileges. Create a system user without a login shell and a configuration directory:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin envoy
sudo install -d -m 755 /etc/envoy
Step 3 - Starting two demo backends
To see load balancing and health checks at work, start two simple HTTP servers on ports 8081 and 8082, bound to localhost. Each one serves an index.html that identifies it and a /health file for health checks:
for n in 1 2; do
sudo install -d -m 755 "/srv/demo$n"
echo "Hello from backend $n" | sudo tee "/srv/demo$n/index.html" > /dev/null
echo "ok" | sudo tee "/srv/demo$n/health" > /dev/null
done
Run them as transient systemd services so they keep running after you log out:
sudo systemd-run --unit=demo-backend-1 -p DynamicUser=yes python3 -m http.server 8081 --bind 127.0.0.1 --directory /srv/demo1
sudo systemd-run --unit=demo-backend-2 -p DynamicUser=yes python3 -m http.server 8082 --bind 127.0.0.1 --directory /srv/demo2
Confirm both respond:
curl http://127.0.0.1:8081/
curl http://127.0.0.1:8082/
Hello from backend 1
Hello from backend 2
Step 4 - Writing the Envoy configuration
An Envoy configuration has three main building blocks:
- Listeners accept connections on an address and port and pass them through a chain of filters. For HTTP traffic, the key filter is the HTTP connection manager, which contains the routing table and the HTTP filters.
- Routes match requests by domain and path and send them to a cluster, with per-route timeouts and retry policies.
- Clusters are groups of upstream endpoints, with a load balancing policy, health checks and outlier detection.
The admin section exposes an HTTP interface for statistics and debugging.
Create the configuration file:
sudo nano /etc/envoy/envoy.yaml
admin:
address:
socket_address:
address: 127.0.0.1
port_value: 9901
static_resources:
listeners:
- name: http_listener
address:
socket_address:
address: 0.0.0.0
port_value: 80
filter_chains:
- filters:
- name: envoy.filters.network.http_connection_manager
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
stat_prefix: ingress_http
access_log:
- name: envoy.access_loggers.stdout
typed_config:
"@type": type.googleapis.com/envoy.extensions.access_loggers.stream.v3.StdoutAccessLog
route_config:
name: local_route
virtual_hosts:
- name: app
domains: ["*"]
routes:
- match:
prefix: "/"
route:
cluster: app_backend
timeout: 15s
retry_policy:
retry_on: "connect-failure,refused-stream,5xx"
num_retries: 2
per_try_timeout: 5s
http_filters:
- name: envoy.filters.http.local_ratelimit
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.http.local_ratelimit.v3.LocalRateLimit
stat_prefix: http_local_rate_limiter
token_bucket:
max_tokens: 20
tokens_per_fill: 20
fill_interval: 1s
filter_enabled:
runtime_key: local_rate_limit_enabled
default_value:
numerator: 100
denominator: HUNDRED
filter_enforced:
runtime_key: local_rate_limit_enforced
default_value:
numerator: 100
denominator: HUNDRED
- name: envoy.filters.http.router
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router
clusters:
- name: app_backend
type: STATIC
connect_timeout: 1s
lb_policy: ROUND_ROBIN
load_assignment:
cluster_name: app_backend
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address:
address: 127.0.0.1
port_value: 8081
- endpoint:
address:
socket_address:
address: 127.0.0.1
port_value: 8082
health_checks:
- timeout: 1s
interval: 5s
unhealthy_threshold: 2
healthy_threshold: 2
http_health_check:
path: /health
outlier_detection:
consecutive_5xx: 5
interval: 10s
base_ejection_time: 30s
max_ejection_percent: 50
What each part does:
- The admin interface listens only on
127.0.0.1:9901. It can change runtime settings and shut Envoy down, so never expose it publicly. - The listener accepts HTTP on port 80 and logs every request to standard output, which systemd sends to the journal.
- The route sends all paths to
app_backend, gives each request 15 seconds in total and retries up to twice on connection failures and5xxresponses, with 5 seconds per attempt. - The local rate limit filter is a token bucket shared by the whole listener: 20 requests per second, answering
429when the bucket is empty. It runs before the router, which must always be the last HTTP filter. - The cluster is
STATICbecause the endpoints are IP addresses. UseSTRICT_DNSwith hostnames if you want Envoy to resolve them and follow DNS changes. - Active health checks request
/healthevery 5 seconds and remove an endpoint after 2 failures. - Outlier detection is passive: an endpoint that returns 5 consecutive
5xxresponses to real traffic is ejected for 30 seconds, and never more than half of the cluster at once.
Validate the file before starting Envoy:
envoy --mode validate -c /etc/envoy/envoy.yaml
configuration '/etc/envoy/envoy.yaml' OK
Step 5 - Running Envoy as a systemd service
Create a unit file:
sudo nano /etc/systemd/system/envoy.service
[Unit]
Description=Envoy Proxy
Documentation=https://www.envoyproxy.io/docs
After=network-online.target
Wants=network-online.target
[Service]
User=envoy
Group=envoy
ExecStartPre=/usr/local/bin/envoy --mode validate -c /etc/envoy/envoy.yaml
ExecStart=/usr/local/bin/envoy -c /etc/envoy/envoy.yaml --log-level info
AmbientCapabilities=CAP_NET_BIND_SERVICE
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=yes
LimitNOFILE=65536
Restart=on-failure
RestartSec=5s
[Install]
WantedBy=multi-user.target
CAP_NET_BIND_SERVICE lets the unprivileged envoy user bind to port 80, and ExecStartPre refuses to start with an invalid configuration. LimitNOFILE raises the open file limit, since each connection uses a file descriptor.
Load the unit, then enable and start Envoy:
sudo systemctl daemon-reload
sudo systemctl enable --now envoy
systemctl status envoy --no-pager
● envoy.service - Envoy Proxy
Loaded: loaded (/etc/systemd/system/envoy.service; enabled; preset: enabled)
Active: active (running) since ...
If you use UFW, allow HTTP traffic:
sudo ufw allow 80/tcp
Step 6 - Testing load balancing and rate limiting
Send four requests through Envoy. Round robin alternates between the two backends:
for i in 1 2 3 4; do curl -s http://127.0.0.1/; done
Hello from backend 1
Hello from backend 2
Hello from backend 1
Hello from backend 2
The access log entries appear in the journal:
sudo journalctl -u envoy -n 3 --no-pager
Sep 25 11:02:13 envoy-01 envoy[4121]: [2026-09-25T11:02:13.418Z] "GET / HTTP/1.1" 200 - 0 21 2 1 "-" "curl/8.5.0" "6a1d..." "127.0.0.1" "127.0.0.1:8082"
The last field is the upstream endpoint that served the request.
Now send a burst of 50 requests as fast as possible and count the status codes. With a bucket of 20 tokens, most of the burst is rejected:
for i in $(seq 1 50); do curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1/; done | sort | uniq -c
20 200
30 429
The exact split depends on how fast your server sends the requests, since the bucket refills every second.
Step 7 - Watching health checks in the admin interface
The /clusters endpoint of the admin interface shows every endpoint with its health flags:
curl -s http://127.0.0.1:9901/clusters | grep health_flags
app_backend::127.0.0.1:8081::health_flags::healthy
app_backend::127.0.0.1:8082::health_flags::healthy
Stop the second backend to simulate a failure:
sudo systemctl stop demo-backend-2
After about 10 seconds (two failed checks at a 5 second interval), Envoy marks it as failed:
curl -s http://127.0.0.1:9901/clusters | grep health_flags
app_backend::127.0.0.1:8081::health_flags::healthy
app_backend::127.0.0.1:8082::health_flags::/failed_active_hc
All traffic now goes to backend 1, and the retry policy covers requests that were in flight when the backend went down:
for i in 1 2 3 4; do curl -s http://127.0.0.1/; done
Hello from backend 1
Hello from backend 1
Hello from backend 1
Hello from backend 1
Other useful admin endpoints:
/stats?filter=app_backendshows counters for the cluster, such asupstream_rq_retryandhealth_check.failure./stats/prometheusexposes every metric in Prometheus format for scraping./config_dumpreturns the full configuration Envoy is running, including defaults./readyreturnsLIVEwhen Envoy is ready to serve traffic, which is useful for external health checks.
When you are done testing, remove the demo backends:
sudo systemctl stop demo-backend-1
Then replace the two 127.0.0.1 endpoints in /etc/envoy/envoy.yaml with your real services and restart Envoy:
sudo systemctl restart envoy
Troubleshooting
- The service fails with
cannot bind '0.0.0.0:80': Permission denied: the capability lines are missing from the unit, or another service already uses port 80. Check withsudo ss -ltnp 'sport = :80'. - Validation fails with
Protobuf message ... has unknown fields: a key is misspelled or indented under the wrong parent. The error names the field; compare it with the Envoy API reference for your version. - Requests return
503withno healthy upstream: every endpoint failed its health check. Check the flags in/clustersand make sure the health check path returns200from the Envoy server itself. - More detail is needed: raise the log level at runtime without a restart, for example
curl -X POST "http://127.0.0.1:9901/logging?level=debug", and set it back toinfowhen you are done.
Conclusion
You installed Envoy from the official release binaries, ran it as a hardened systemd service and configured a listener, routes and a cluster with active health checks, outlier detection, retries and local rate limiting. Good next steps are terminating TLS in the listener with a DownstreamTlsContext, scraping /stats/prometheus with Prometheus, and moving from static configuration to dynamic xDS if you manage many services.
