HAProxy includes a small in-memory HTTP cache that can answer repeated GET requests without contacting your backend servers. It is not a replacement for Varnish, but when HAProxy already sits in front of your application it can absorb a large share of traffic for static files and public API responses with a few lines of configuration. In this tutorial you will enable the cache on Ubuntu 24.04, restrict it to safe paths, add an X-Cache-Status header and verify hits from the command line and the runtime API.

Prerequisites

To follow this guide you need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS, with a non-root user with sudo privileges.
  • A web application listening on 127.0.0.1:8080. If you do not have one yet, Step 2 sets up Nginx on that port as a test backend.
  • Port 80 open in your firewall (sudo ufw allow 80/tcp if you use UFW).

This guide uses plain HTTP to keep the focus on caching. The cache configuration is the same when your frontend also terminates TLS.

Step 1 - Installing HAProxy

Ubuntu 24.04 ships HAProxy 2.8, a long-term support branch that includes every cache feature used here, including process-vary. Install it from the Ubuntu repositories:

sudo apt update
sudo apt install haproxy socat

socat is used later to query the HAProxy runtime API. Check the installed version:

haproxy -v
HAProxy version 2.8.x-1ubuntu0.x 2024/xx/xx - https://haproxy.org/

Before editing anything, keep a copy of the default configuration:

sudo cp /etc/haproxy/haproxy.cfg /etc/haproxy/haproxy.cfg.orig

Step 2 - Preparing a test backend (optional)

If your application already listens on 127.0.0.1:8080, skip to Step 3. Otherwise, install Nginx and move it off port 80 so HAProxy can take that port:

sudo apt install nginx
sudo sed -i 's/listen 80 default_server;/listen 127.0.0.1:8080 default_server;/; /listen \[::\]:80 default_server;/d' /etc/nginx/sites-available/default
sudo nginx -t && sudo systemctl restart nginx

Nginx serves its default page with Last-Modified and ETag headers, which is enough for HAProxy to consider it cacheable. Confirm it answers on the new port:

curl -sI http://127.0.0.1:8080/ | head -n 1
HTTP/1.1 200 OK

Step 3 - Defining the cache

A cache in HAProxy is a named section that reserves a block of shared memory. Open the configuration file:

sudo nano /etc/haproxy/haproxy.cfg

Leave the existing global and defaults sections as they are and add this section at the end of the file:

cache app_cache
    total-max-size 256
    max-object-size 1048576
    max-age 300
    process-vary on
    max-secondary-entries 10

What each line does:

  • total-max-size 256: memory reserved for the cache, in megabytes. The maximum is 4095.
  • max-object-size 1048576: largest object that will be stored, in bytes (1 MB here). It must be smaller than half of total-max-size. Larger responses are passed through without being cached.
  • max-age 300: upper limit, in seconds, for how long an object stays in the cache. If the backend sends a shorter Cache-Control: max-age or s-maxage, the shorter value wins.
  • process-vary on: store separate variants for responses that carry a Vary header. Without it, any response with Vary is never cached.
  • max-secondary-entries 10: how many variants of the same URL can be stored when process-vary is on.

Size the cache according to the memory you can spare on the server. A 256 MB cache holding objects that average 50 KB fits roughly 5,000 objects.

Step 4 - Enabling the cache in the frontend and backend

A cache does nothing until a proxy uses it. Two directives do the work:

  • http-request cache-use looks the request up in the cache and, on a hit, answers directly.
  • http-response cache-store saves eligible responses coming back from the server.

Add a frontend and a backend below the cache section in /etc/haproxy/haproxy.cfg:

frontend web
    bind :80
    default_backend app

    http-response set-header X-Cache-Status HIT if !{ srv_id -m found }
    http-response set-header X-Cache-Status MISS if { srv_id -m found }

backend app
    acl cacheable_path path_beg /static/ /assets/ /api/public/
    acl cacheable_path path /
    acl has_cookie req.hdr(Cookie) -m found

    http-request cache-use app_cache if cacheable_path !has_cookie
    http-response cache-store app_cache if cacheable_path

    server app1 127.0.0.1:8080 check

The X-Cache-Status rules rely on a simple fact: when a response comes from the cache, no server was involved, so the srv_id sample is empty. On a miss, the response comes from app1 and srv_id is set.

The ACLs limit caching to paths you know are public (/, /static/, /assets/ and /api/public/ in this example; adjust them to your application) and skip the cache lookup for clients that send cookies. Caching only what you explicitly allow is safer than caching everything and trying to exclude private pages later.

Validate the file and reload HAProxy:

sudo haproxy -c -f /etc/haproxy/haproxy.cfg
sudo systemctl reload haproxy
Configuration file is valid

If the check reports an error, it prints the file and line number. Fix it before reloading, otherwise HAProxy keeps running with the previous configuration.

Step 5 - Verifying cache hits

Request the same URL twice through HAProxy. Replace your_server_ip with the public IP of your server, or use 127.0.0.1 from the server itself:

curl -sI http://your_server_ip/ | grep -iE '^(x-cache-status|age)'
curl -sI http://your_server_ip/ | grep -iE '^(x-cache-status|age)'
x-cache-status: MISS
x-cache-status: HIT
age: 2

The first request went to the backend and was stored, the second one was served from memory. HAProxy adds an Age header to cached responses with the number of seconds the object has been in the cache.

You can also inspect the cache through the runtime API. The default Ubuntu configuration already exposes an admin socket at /run/haproxy/admin.sock:

echo "show cache" | sudo socat stdio /run/haproxy/admin.sock

The output shows the app_cache section with its number of available blocks, followed by one line per stored object with its size, key and remaining lifetime. When the list is empty after several requests, nothing is being stored; see the Troubleshooting section.

Step 6 - Controlling what gets cached and for how long

HAProxy respects standard HTTP caching rules. It does not store a response when:

  • The request method is not GET (HEAD, POST and others are never cached).
  • The request contains an Authorization header.
  • The response status is not 200.
  • The response is marked Cache-Control: private, no-cache or no-store.
  • The response has no expiration (Cache-Control: max-age, s-maxage or Expires) and no validator (ETag or Last-Modified).
  • The response carries a Vary header on something other than Accept-Encoding, Referer or Origin, or process-vary is off.
  • The response is larger than max-object-size.

A Set-Cookie header also usually makes the response uncacheable, which is what you want for pages tied to a session.

The best place to set lifetimes is the application itself, because it knows which content is public. When you cannot change the application, set the header in HAProxy for specific paths before the response is stored. For example, give static assets a long browser lifetime while the HAProxy cache still keeps them for at most max-age (300 seconds):

backend app
    acl cacheable_path path_beg /static/ /assets/ /api/public/
    acl cacheable_path path /
    acl has_cookie req.hdr(Cookie) -m found
    acl static_path path_beg /static/ /assets/

    http-response set-header Cache-Control "public, max-age=86400" if static_path
    http-request cache-use app_cache if cacheable_path !has_cookie
    http-response cache-store app_cache if cacheable_path

    server app1 127.0.0.1:8080 check

http-response rules run in the order they appear, so the set-header rule must come before cache-store. Only do this for content that is identical for every visitor.

Step 7 - Normalizing headers for Vary

When your application sends Vary: Accept-Encoding, HAProxy stores one copy per distinct Accept-Encoding value. Browsers send many different combinations (gzip, deflate, br, br, gzip, gzip...), which splits the same object into several entries and lowers the hit rate.

If your backend only compresses with gzip, collapse the header to two values before the cache lookup. Add these lines to the backend app section, above http-request cache-use:

    http-request set-header Accept-Encoding gzip if { req.hdr(Accept-Encoding) -m sub -i gzip }
    http-request del-header Accept-Encoding unless { req.hdr(Accept-Encoding) -m sub -i gzip }

Validate and reload again:

sudo haproxy -c -f /etc/haproxy/haproxy.cfg && sudo systemctl reload haproxy

Now each URL has at most two variants: gzip and uncompressed. Test both:

curl -sI -H 'Accept-Encoding: gzip, deflate, br' http://your_server_ip/ | grep -i x-cache-status
curl -sI -H 'Accept-Encoding: br, gzip' http://your_server_ip/ | grep -i x-cache-status
x-cache-status: MISS
x-cache-status: HIT

The second request is a hit even though the client sent a different header, because both were normalized to gzip.

Troubleshooting

Every request shows X-Cache-Status: MISS. Look at the headers your backend sends:

curl -sI http://127.0.0.1:8080/your/path

Check for Cache-Control: private or no-store, a Set-Cookie header, a missing ETag/Last-Modified/max-age, or a Vary header on an unsupported field such as Cookie or User-Agent. Also make sure you are testing with GET and without cookies or an Authorization header.

The configuration check fails with an error about max-object-size. The value is in bytes and must be smaller than half of total-max-size (which is in megabytes). A value such as max-object-size 10 means 10 bytes, not 10 MB.

Stale content after a deployment. HAProxy has no command to purge a single URL. Objects expire after max-age, and the cache lives in the memory of the running process, so sudo systemctl reload haproxy starts the new process with an empty cache. Keep max-age short for content that changes often, or use versioned file names for static assets.

HAProxy fails to start after a reboot. Check the logs:

sudo journalctl -u haproxy -n 50 --no-pager

A cache larger than the memory available is a common cause on small servers; lower total-max-size.

Conclusion

HAProxy is now caching public responses in memory, tagging each one with X-Cache-Status, and storing gzip and uncompressed variants separately. Because the cache follows standard Cache-Control rules, most tuning happens in your application's headers rather than in HAProxy. As next steps, add TLS termination to the web frontend with a Let's Encrypt certificate, enable the HAProxy stats page restricted to your IP to watch backend load drop, or move to Varnish if you need purging and larger on-disk caches.