Rate limiting caps how many requests a client can make in a period of time. It protects an API from abusive clients, brute-force attempts on login endpoints, and runaway scripts that would otherwise exhaust your backend. In this tutorial you will compare the common algorithms, then use Nginx on Ubuntu 24.04 as a reverse proxy to apply per-IP, per-API-key and per-endpoint limits that return proper 429 Too Many Requests responses, and finally implement the same policy in HAProxy.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS, with a non-root user with
sudoprivileges. - An API listening on the server, or reachable from it. The examples use a backend at
127.0.0.1:8000; Step 1 creates a stand-in if you do not have one. - Nginx installed (
sudo apt install -y nginx). The HAProxy section is independent and usessudo apt install -y haproxy.
Rate limiting algorithms
Every rate limiter answers the same question, "has this client exceeded its budget?", with a different trade-off between precision, memory and burst handling.
| Algorithm | How it works | Bursts | Used by |
|---|---|---|---|
| Fixed window | Counts requests per calendar interval (for example per minute) and resets the counter at the boundary | Allows up to 2x the limit around the boundary | Simple app-level counters in Redis |
| Sliding window | Counts requests in the last N seconds, rolling continuously | Smooth, no boundary spikes | HAProxy http_req_rate() |
| Token bucket | A bucket refills at a fixed rate; each request takes a token; an empty bucket rejects | Allows short bursts up to the bucket size | Many API gateways and cloud APIs |
| Leaky bucket | Requests enter a queue drained at a fixed rate; overflow is rejected | Smooths traffic to a steady rate | Nginx limit_req |
Nginx's limit_req is a leaky bucket with a burst queue. With nodelay, queued requests are served immediately instead of being spaced out, which makes it behave like a token bucket: a client can burst up to burst requests, then is held to the base rate.
What you key the limit on matters as much as the algorithm:
- Client IP: works for anonymous traffic, but many users behind one NAT share a budget.
- API key or user ID: fair per customer, but only once you can identify the client.
- Endpoint: expensive or sensitive routes (login, search, exports) deserve their own, stricter limits.
Step 1 - Preparing a test backend
If you already have an API on 127.0.0.1:8000, skip this step. Otherwise, create a small Nginx server block that stands in for it, so you can test limits without an application:
sudo nano /etc/nginx/sites-available/fake-api
server {
listen 127.0.0.1:8000;
location / {
default_type application/json;
return 200 '{"status":"ok"}\n';
}
}
Enable it, test the configuration and reload:
sudo ln -s /etc/nginx/sites-available/fake-api /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
curl -s http://127.0.0.1:8000/
{"status":"ok"}
The rate limits go in a separate public server block that proxies to this backend. They cannot live in the same location as a return directive, because return answers the request before Nginx evaluates limit_req.
Step 2 - Limiting requests per client IP
A limit in Nginx has two parts: limit_req_zone, declared once in the http context, defines the key, a shared memory zone and the rate; limit_req, inside a server or location, applies it.
Ubuntu's /etc/nginx/nginx.conf includes every file in /etc/nginx/conf.d/ inside the http block, so create the zones there:
sudo nano /etc/nginx/conf.d/rate-limits.conf
# 10 requests per second per client IP, 10 MB of state (about 160,000 IPs)
limit_req_zone $binary_remote_addr zone=api_per_ip:10m rate=10r/s;
# Answer rejected requests with 429 instead of the default 503
limit_req_status 429;
limit_conn_status 429;
$binary_remote_addr stores the address in 4 bytes (16 for IPv6), using less memory than the text form $remote_addr.
Now create the public API server block. Replace your_domain with your domain or server IP:
sudo nano /etc/nginx/sites-available/api
server {
listen 80;
server_name your_domain;
location /api/ {
limit_req zone=api_per_ip burst=20 nodelay;
proxy_pass http://127.0.0.1:8000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
}
With rate=10r/s, burst=20 and nodelay, a client can send 20 requests at once above the base rate, served immediately, and after that 1 request every 100 ms. Without nodelay, excess requests would wait in the queue instead, which adds latency to every burst.
Enable the site and reload:
sudo ln -s /etc/nginx/sites-available/api /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
Send 50 requests as fast as possible and count the status codes:
for i in $(seq 1 50); do
curl -s -o /dev/null -w '%{http_code}\n' -H 'Host: your_domain' http://127.0.0.1/api/
done | sort | uniq -c
About 21 requests (1 plus the burst of 20) succeed, plus one or two more as the bucket refills during the loop:
22 200
28 429
Every rejection is logged in the Nginx error log:
sudo grep "limiting requests" /var/log/nginx/error.log | tail -n 2
2026/09/25 10:12:01 [error] 1234#1234: *57 limiting requests, excess: 20.970 by zone "api_per_ip", client: 127.0.0.1, server: your_domain, request: "GET /api/ HTTP/1.1", host: "your_domain"
TipWhen adding limits to an API that is already in production, add
limit_req_dry_run on;next tolimit_reqfirst. Nginx then only logs the requests it would reject, and you can tune the numbers against real traffic before enforcing them.
Step 3 - Returning a useful 429 response
API clients handle rate limits better when the response says so in a machine-readable way. Send a JSON body and a Retry-After header by routing 429 errors to a named location. Edit the API server block:
sudo nano /etc/nginx/sites-available/api
Add the error_page line and the named location inside the server block:
server {
listen 80;
server_name your_domain;
error_page 429 @rate_limited;
location @rate_limited {
default_type application/json;
add_header Retry-After 1 always;
return 429 '{"detail":"Too many requests, retry later"}\n';
}
location /api/ {
limit_req zone=api_per_ip burst=20 nodelay;
proxy_pass http://127.0.0.1:8000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
}
Retry-After is in seconds. With a rate of 10 requests per second, one second is enough for the bucket to accept requests again. Reload and trigger the limit, then look at one rejected response:
sudo nginx -t && sudo systemctl reload nginx
for i in $(seq 1 30); do curl -s -o /dev/null -H 'Host: your_domain' http://127.0.0.1/api/; done
curl -si -H 'Host: your_domain' http://127.0.0.1/api/
HTTP/1.1 429 Too Many Requests
Server: nginx/1.24.0 (Ubuntu)
Content-Type: application/json
Retry-After: 1
...
{"detail":"Too many requests, retry later"}
Step 4 - Adding per-endpoint and per-API-key limits
A single limit per IP is rarely enough. Login endpoints need a much stricter limit against password guessing, and authenticated customers should be limited by their API key rather than their IP. Nginx can apply several limit_req directives to the same request; a request is rejected if any of them is exceeded.
Nginx does not count requests whose key is empty. The following map uses that to split clients: requests with an X-API-Key header are counted per key, requests without one are counted per IP. Replace the contents of the zones file:
sudo nano /etc/nginx/conf.d/rate-limits.conf
# Requests without an API key are keyed by IP; requests with one get an empty key here
map $http_x_api_key $anon_client {
"" $binary_remote_addr;
default "";
}
# Anonymous clients: 5 r/s per IP
limit_req_zone $anon_client zone=api_anon:10m rate=5r/s;
# Authenticated clients: 50 r/s per API key
limit_req_zone $http_x_api_key zone=api_per_key:10m rate=50r/s;
# Ceiling for everyone, so rotating fake keys does not bypass limits
limit_req_zone $binary_remote_addr zone=api_per_ip:10m rate=100r/s;
# Login: 5 attempts per minute per IP
limit_req_zone $binary_remote_addr zone=login:10m rate=5r/m;
# Concurrent connections per IP
limit_conn_zone $binary_remote_addr zone=conn_per_ip:10m;
limit_req_status 429;
limit_conn_status 429;
Then update the location blocks in /etc/nginx/sites-available/api, keeping the error_page and @rate_limited blocks from Step 3:
location = /api/login {
limit_req zone=login burst=5 nodelay;
proxy_pass http://127.0.0.1:8000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
location /api/ {
limit_req zone=api_anon burst=10 nodelay;
limit_req zone=api_per_key burst=100 nodelay;
limit_req zone=api_per_ip burst=200 nodelay;
limit_conn conn_per_ip 50;
proxy_pass http://127.0.0.1:8000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
The location = /api/login block is an exact match, so it takes priority over the /api/ prefix. Its burst=5 allows a few quick retries after a typo, then one attempt every 12 seconds.
ImportantNginx does not validate API keys. A client that sends random keys gets a fresh
api_per_keybucket each time, which is why theapi_per_ipceiling applies to every request. Your application must still reject unknown keys.
Reload and compare anonymous and authenticated traffic:
sudo nginx -t && sudo systemctl reload nginx
for i in $(seq 1 30); do curl -s -o /dev/null -w '%{http_code}\n' -H 'Host: your_domain' http://127.0.0.1/api/; done | sort | uniq -c
for i in $(seq 1 30); do curl -s -o /dev/null -w '%{http_code}\n' -H 'Host: your_domain' -H 'X-API-Key: test-key-1' http://127.0.0.1/api/; done | sort | uniq -c
The anonymous loop hits its 5 r/s limit after about 11 requests; the loop with a key stays within its budget:
11 200
19 429
30 200
Test the login limit the same way:
for i in $(seq 1 10); do curl -s -o /dev/null -w '%{http_code}\n' -H 'Host: your_domain' http://127.0.0.1/api/login; done | sort | uniq -c
6 200
4 429
Exempting trusted networks
Internal services and monitoring should not be throttled. Use geo to flag trusted addresses and turn the key empty for them. Add this to rate-limits.conf, replacing 10.0.0.0/24 with your internal network:
geo $rate_limit_exempt {
default 0;
10.0.0.0/24 1;
}
map $rate_limit_exempt $limited_ip {
0 $binary_remote_addr;
1 "";
}
Then use $limited_ip instead of $binary_remote_addr as the key of the zones you want to skip for those networks, for example limit_req_zone $limited_ip zone=api_per_ip:10m rate=100r/s;.
Behind a load balancer or CDN
If Nginx receives traffic from a load balancer or CDN, $binary_remote_addr is the proxy's address and every user shares one budget. Restore the real client IP with the realip module, trusting only your proxies' addresses. Add to the server block:
set_real_ip_from 10.0.0.0/24;
real_ip_header X-Forwarded-For;
real_ip_recursive on;
Never trust X-Forwarded-For from arbitrary sources, because clients can set it to any value and escape their limits.
Step 5 - Rate limiting in HAProxy
HAProxy tracks request rates in stick tables with a sliding window, which avoids the boundary spikes of fixed windows. The following frontend rejects clients that send more than 100 requests in 10 seconds. Open the HAProxy configuration:
sudo nano /etc/haproxy/haproxy.cfg
Add a frontend and backend at the end of the file:
frontend api_in
bind *:8080
stick-table type ipv6 size 200k expire 30s store http_req_rate(10s)
http-request track-sc0 src
http-request deny deny_status 429 if { sc_http_req_rate(0) gt 100 }
default_backend api_servers
backend api_servers
server api1 127.0.0.1:8000 check
stick-table type ipv6: stores one entry per client address; IPv4 clients are stored as IPv4-mapped IPv6 addresses, so one table covers both. Entries expire 30 seconds after the client goes quiet.store http_req_rate(10s): keeps each client's request rate over a rolling 10-second window.http-request track-sc0 src: starts tracking the client's source address in sticky counter 0.http-request deny deny_status 429 if ...: rejects the request with 429 while the rate is above 100.
To limit per API key instead, track a header in a second table defined in a dedicated backend:
backend api_keys
stick-table type string len 64 size 100k expire 60s store http_req_rate(60s)
And add these lines to frontend api_in, after the existing track-sc0 rule:
http-request track-sc1 req.hdr(x-api-key) table api_keys if { req.hdr(x-api-key) -m found }
http-request deny deny_status 429 if { sc_http_req_rate(1,api_keys) gt 1000 }
Validate and reload:
sudo haproxy -c -f /etc/haproxy/haproxy.cfg
sudo systemctl reload haproxy
Configuration file is valid
Send 120 requests and count the results:
for i in $(seq 1 120); do curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8080/api/; done | sort | uniq -c
100 200
20 429
Inspect the table to see each client and its current rate:
echo "show table api_in" | sudo socat stdio /run/haproxy/admin.sock
# table: api_in, type: ipv6, size:204800, used:1
0x...: key=::ffff:127.0.0.1 use=0 exp=29120 http_req_rate(10000)=120
Install socat with sudo apt install -y socat if the command is missing.
Running limits across several servers
Nginx zones and HAProxy stick tables live in the memory of one process. With three load balancers behind DNS or an anycast address, each one enforces its own copy, so a client can get up to three times the limit. You can accept that and divide the configured rates by the number of nodes, synchronize HAProxy stick tables between nodes with a peers section, or enforce exact per-customer quotas in the application with a shared store such as Redis. A common setup combines both: coarse per-IP limits at the proxy to absorb floods cheaply, and precise per-key quotas in the application.
Troubleshooting
Every client is limited at once. The proxy sees one address for everyone, usually a load balancer or CDN. Configure the realip module (Nginx) or track the right header (HAProxy) as shown above.
nginx: [emerg] zero size shared memory zone. A limit_req directive references a zone name that no limit_req_zone defines. Check the spelling in both files.
Limits do not apply at all. The location answers with return, or another, more specific location block handles the request without a limit_req. Run sudo nginx -T | less to see the effective configuration.
Legitimate bursts are rejected. Page loads and SDKs often send several requests in parallel. Raise burst rather than the base rate, and keep nodelay so the burst is not slowed down.
Conclusion
You compared the main rate limiting algorithms, configured Nginx to limit requests per IP, per API key and per endpoint with JSON 429 responses and Retry-After, and built an equivalent sliding window limit in HAProxy. Next, roll new limits out with limit_req_dry_run first, graph the limiting requests log lines to see who hits them, and document your limits for API consumers so they can back off correctly.
