Envoy is an open source Layer 7 proxy, originally built at Lyft and now a graduated CNCF project, that powers service meshes such as Istio and many API gateways. It can also run on its own as an edge proxy in front of your applications. In this tutorial you will install the official Envoy binary on Ubuntu 24.04, run it as a systemd service, and build a static configuration that load balances two backends with active health checks, retries, outlier detection and local rate limiting.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS (x86_64 or arm64), for example a CubePath VPS, with at least 1 GB of RAM.
  • A non-root user with sudo privileges.
  • Port 80 open if you want to reach Envoy from outside the server.

You will run two small demo backends on the same server with Python, which is installed by default on Ubuntu 24.04. Replace them with your own services when you move to production.

Step 1 - Installing the Envoy binary

The Envoy project publishes statically linked binaries for each release on GitHub, together with a signed checksum file. This guide uses version 1.39.1; check the releases page for the latest version and adjust the variable.

Download the binary and the checksum file:

ENVOY_VERSION=1.39.1
cd /tmp
curl -fLO "https://github.com/envoyproxy/envoy/releases/download/v${ENVOY_VERSION}/envoy-${ENVOY_VERSION}-linux-x86_64"
curl -fLO "https://github.com/envoyproxy/envoy/releases/download/v${ENVOY_VERSION}/checksums.txt.asc"

On an arm64 server, replace linux-x86_64 with linux-aarch_64 in this step.

Verify the download against the published SHA-256 checksum:

grep "envoy-${ENVOY_VERSION}-linux-x86_64$" checksums.txt.asc | awk '{print $1"  envoy-'"${ENVOY_VERSION}"'-linux-x86_64"}' | sha256sum -c
envoy-1.39.1-linux-x86_64: OK

Install it to /usr/local/bin and check the version:

sudo install -m 755 "envoy-${ENVOY_VERSION}-linux-x86_64" /usr/local/bin/envoy
envoy --version
envoy  version: 5c8f3d6c.../1.39.1/Clean/RELEASE/BoringSSL

Step 2 - Creating a user and directories

Envoy does not need root privileges. Create a system user without a login shell and a configuration directory:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin envoy
sudo install -d -m 755 /etc/envoy

Step 3 - Starting two demo backends

To see load balancing and health checks at work, start two simple HTTP servers on ports 8081 and 8082, bound to localhost. Each one serves an index.html that identifies it and a /health file for health checks:

for n in 1 2; do
  sudo install -d -m 755 "/srv/demo$n"
  echo "Hello from backend $n" | sudo tee "/srv/demo$n/index.html" > /dev/null
  echo "ok" | sudo tee "/srv/demo$n/health" > /dev/null
done

Run them as transient systemd services so they keep running after you log out:

sudo systemd-run --unit=demo-backend-1 -p DynamicUser=yes python3 -m http.server 8081 --bind 127.0.0.1 --directory /srv/demo1
sudo systemd-run --unit=demo-backend-2 -p DynamicUser=yes python3 -m http.server 8082 --bind 127.0.0.1 --directory /srv/demo2

Confirm both respond:

curl http://127.0.0.1:8081/
curl http://127.0.0.1:8082/
Hello from backend 1
Hello from backend 2

Step 4 - Writing the Envoy configuration

An Envoy configuration has three main building blocks:

  • Listeners accept connections on an address and port and pass them through a chain of filters. For HTTP traffic, the key filter is the HTTP connection manager, which contains the routing table and the HTTP filters.
  • Routes match requests by domain and path and send them to a cluster, with per-route timeouts and retry policies.
  • Clusters are groups of upstream endpoints, with a load balancing policy, health checks and outlier detection.

The admin section exposes an HTTP interface for statistics and debugging.

Create the configuration file:

sudo nano /etc/envoy/envoy.yaml
admin:
  address:
    socket_address:
      address: 127.0.0.1
      port_value: 9901

static_resources:
  listeners:
    - name: http_listener
      address:
        socket_address:
          address: 0.0.0.0
          port_value: 80
      filter_chains:
        - filters:
            - name: envoy.filters.network.http_connection_manager
              typed_config:
                "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
                stat_prefix: ingress_http
                access_log:
                  - name: envoy.access_loggers.stdout
                    typed_config:
                      "@type": type.googleapis.com/envoy.extensions.access_loggers.stream.v3.StdoutAccessLog
                route_config:
                  name: local_route
                  virtual_hosts:
                    - name: app
                      domains: ["*"]
                      routes:
                        - match:
                            prefix: "/"
                          route:
                            cluster: app_backend
                            timeout: 15s
                            retry_policy:
                              retry_on: "connect-failure,refused-stream,5xx"
                              num_retries: 2
                              per_try_timeout: 5s
                http_filters:
                  - name: envoy.filters.http.local_ratelimit
                    typed_config:
                      "@type": type.googleapis.com/envoy.extensions.filters.http.local_ratelimit.v3.LocalRateLimit
                      stat_prefix: http_local_rate_limiter
                      token_bucket:
                        max_tokens: 20
                        tokens_per_fill: 20
                        fill_interval: 1s
                      filter_enabled:
                        runtime_key: local_rate_limit_enabled
                        default_value:
                          numerator: 100
                          denominator: HUNDRED
                      filter_enforced:
                        runtime_key: local_rate_limit_enforced
                        default_value:
                          numerator: 100
                          denominator: HUNDRED
                  - name: envoy.filters.http.router
                    typed_config:
                      "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

  clusters:
    - name: app_backend
      type: STATIC
      connect_timeout: 1s
      lb_policy: ROUND_ROBIN
      load_assignment:
        cluster_name: app_backend
        endpoints:
          - lb_endpoints:
              - endpoint:
                  address:
                    socket_address:
                      address: 127.0.0.1
                      port_value: 8081
              - endpoint:
                  address:
                    socket_address:
                      address: 127.0.0.1
                      port_value: 8082
      health_checks:
        - timeout: 1s
          interval: 5s
          unhealthy_threshold: 2
          healthy_threshold: 2
          http_health_check:
            path: /health
      outlier_detection:
        consecutive_5xx: 5
        interval: 10s
        base_ejection_time: 30s
        max_ejection_percent: 50

What each part does:

  • The admin interface listens only on 127.0.0.1:9901. It can change runtime settings and shut Envoy down, so never expose it publicly.
  • The listener accepts HTTP on port 80 and logs every request to standard output, which systemd sends to the journal.
  • The route sends all paths to app_backend, gives each request 15 seconds in total and retries up to twice on connection failures and 5xx responses, with 5 seconds per attempt.
  • The local rate limit filter is a token bucket shared by the whole listener: 20 requests per second, answering 429 when the bucket is empty. It runs before the router, which must always be the last HTTP filter.
  • The cluster is STATIC because the endpoints are IP addresses. Use STRICT_DNS with hostnames if you want Envoy to resolve them and follow DNS changes.
  • Active health checks request /health every 5 seconds and remove an endpoint after 2 failures.
  • Outlier detection is passive: an endpoint that returns 5 consecutive 5xx responses to real traffic is ejected for 30 seconds, and never more than half of the cluster at once.

Validate the file before starting Envoy:

envoy --mode validate -c /etc/envoy/envoy.yaml
configuration '/etc/envoy/envoy.yaml' OK

Step 5 - Running Envoy as a systemd service

Create a unit file:

sudo nano /etc/systemd/system/envoy.service
[Unit]
Description=Envoy Proxy
Documentation=https://www.envoyproxy.io/docs
After=network-online.target
Wants=network-online.target

[Service]
User=envoy
Group=envoy
ExecStartPre=/usr/local/bin/envoy --mode validate -c /etc/envoy/envoy.yaml
ExecStart=/usr/local/bin/envoy -c /etc/envoy/envoy.yaml --log-level info
AmbientCapabilities=CAP_NET_BIND_SERVICE
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=yes
LimitNOFILE=65536
Restart=on-failure
RestartSec=5s

[Install]
WantedBy=multi-user.target

CAP_NET_BIND_SERVICE lets the unprivileged envoy user bind to port 80, and ExecStartPre refuses to start with an invalid configuration. LimitNOFILE raises the open file limit, since each connection uses a file descriptor.

Load the unit, then enable and start Envoy:

sudo systemctl daemon-reload
sudo systemctl enable --now envoy
systemctl status envoy --no-pager
● envoy.service - Envoy Proxy
     Loaded: loaded (/etc/systemd/system/envoy.service; enabled; preset: enabled)
     Active: active (running) since ...

If you use UFW, allow HTTP traffic:

sudo ufw allow 80/tcp

Step 6 - Testing load balancing and rate limiting

Send four requests through Envoy. Round robin alternates between the two backends:

for i in 1 2 3 4; do curl -s http://127.0.0.1/; done
Hello from backend 1
Hello from backend 2
Hello from backend 1
Hello from backend 2

The access log entries appear in the journal:

sudo journalctl -u envoy -n 3 --no-pager
Sep 25 11:02:13 envoy-01 envoy[4121]: [2026-09-25T11:02:13.418Z] "GET / HTTP/1.1" 200 - 0 21 2 1 "-" "curl/8.5.0" "6a1d..." "127.0.0.1" "127.0.0.1:8082"

The last field is the upstream endpoint that served the request.

Now send a burst of 50 requests as fast as possible and count the status codes. With a bucket of 20 tokens, most of the burst is rejected:

for i in $(seq 1 50); do curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1/; done | sort | uniq -c
     20 200
     30 429

The exact split depends on how fast your server sends the requests, since the bucket refills every second.

Step 7 - Watching health checks in the admin interface

The /clusters endpoint of the admin interface shows every endpoint with its health flags:

curl -s http://127.0.0.1:9901/clusters | grep health_flags
app_backend::127.0.0.1:8081::health_flags::healthy
app_backend::127.0.0.1:8082::health_flags::healthy

Stop the second backend to simulate a failure:

sudo systemctl stop demo-backend-2

After about 10 seconds (two failed checks at a 5 second interval), Envoy marks it as failed:

curl -s http://127.0.0.1:9901/clusters | grep health_flags
app_backend::127.0.0.1:8081::health_flags::healthy
app_backend::127.0.0.1:8082::health_flags::/failed_active_hc

All traffic now goes to backend 1, and the retry policy covers requests that were in flight when the backend went down:

for i in 1 2 3 4; do curl -s http://127.0.0.1/; done
Hello from backend 1
Hello from backend 1
Hello from backend 1
Hello from backend 1

Other useful admin endpoints:

  • /stats?filter=app_backend shows counters for the cluster, such as upstream_rq_retry and health_check.failure.
  • /stats/prometheus exposes every metric in Prometheus format for scraping.
  • /config_dump returns the full configuration Envoy is running, including defaults.
  • /ready returns LIVE when Envoy is ready to serve traffic, which is useful for external health checks.

When you are done testing, remove the demo backends:

sudo systemctl stop demo-backend-1

Then replace the two 127.0.0.1 endpoints in /etc/envoy/envoy.yaml with your real services and restart Envoy:

sudo systemctl restart envoy

Troubleshooting

  • The service fails with cannot bind '0.0.0.0:80': Permission denied: the capability lines are missing from the unit, or another service already uses port 80. Check with sudo ss -ltnp 'sport = :80'.
  • Validation fails with Protobuf message ... has unknown fields: a key is misspelled or indented under the wrong parent. The error names the field; compare it with the Envoy API reference for your version.
  • Requests return 503 with no healthy upstream: every endpoint failed its health check. Check the flags in /clusters and make sure the health check path returns 200 from the Envoy server itself.
  • More detail is needed: raise the log level at runtime without a restart, for example curl -X POST "http://127.0.0.1:9901/logging?level=debug", and set it back to info when you are done.

Conclusion

You installed Envoy from the official release binaries, ran it as a hardened systemd service and configured a listener, routes and a cluster with active health checks, outlier detection, retries and local rate limiting. Good next steps are terminating TLS in the listener with a DownstreamTlsContext, scraping /stats/prometheus with Prometheus, and moving from static configuration to dynamic xDS if you manage many services.