wrk and siege are two open source HTTP load testing tools that go further than Apache Bench. wrk is multi-threaded and event driven, so a single client machine can generate very high load, and it can be scripted in Lua. siege is better at simulating users: it replays a list of URLs, adds think time between requests, and runs for a fixed duration. In this tutorial you will install both on Ubuntu 24.04, run baseline tests, script a POST request and custom latency report with wrk, and replay a realistic URL list with siege.

Prerequisites

To follow this tutorial, you will need:

  • A client machine running Ubuntu 24.04 LTS with a non-root sudo user and at least 2 CPU cores, since wrk uses one thread per core.
  • A target web server you are allowed to load test, reachable at your_server_ip or your_domain, for example a CubePath VPS running Nginx.
  • Ideally, client and target on separate machines in the same region, so the load generator does not compete with the server for CPU.

Step 1 - Installing wrk and siege

Both tools are packaged in Ubuntu's universe repository:

sudo apt update
sudo apt install wrk siege

Check the installed versions:

wrk --version | head -n 1
siege --version 2>&1 | head -n 1
wrk 4.1.0 [epoll] Copyright (C) 2012 Will Glozer
SIEGE 4.0.7

High-concurrency tests need many open sockets. Raise the file descriptor limit for your current shell before testing:

ulimit -n 65535

Step 2 - Running a baseline test with wrk

wrk takes three main options: -t (threads), -c (total open connections, shared across threads) and -d (duration). Use as many threads as the client has CPU cores, and add --latency to print the latency distribution:

wrk -t 2 -c 100 -d 30s --latency http://your_server_ip/
Running 30s test @ http://203.0.113.10/
  2 threads and 100 connections
  Thread Stats   Avg      Stdev     Max   +/- Stdev
    Latency     8.12ms    3.05ms  61.40ms   82.14%
    Req/Sec     6.21k   512.33     7.40k    71.83%
  Latency Distribution
     50%    7.61ms
     75%    9.35ms
     90%   11.72ms
     99%   19.88ms
  370912 requests in 30.02s, 300.72MB read
Requests/sec:  12355.49
Transfer/sec:     10.02MB

How to read it:

  • Requests/sec is the total throughput, the number to compare between configurations.
  • Latency Distribution shows percentiles. The 99% value is the latency the slowest 1% of requests experience.
  • Req/Sec in Thread Stats is per thread. If one thread is far below the others, the client may be the bottleneck.
  • If there are errors, wrk adds a line such as Socket errors: connect 0, read 12, write 0, timeout 45 and Non-2xx or 3xx responses: 230. Any of these means the server could not keep up at this load.

wrk always uses keep-alive connections, like a browser. It speaks HTTP/1.1 only.

Step 3 - Scripting requests with Lua

wrk exposes a Lua table called wrk that you can change before the test starts, plus hooks such as request(), response() and done(). This is how you send POST requests, add headers or produce custom reports.

Sending a JSON POST request

Create a script:

nano post.lua
wrk.method = "POST"
wrk.body   = '{"name": "benchmark", "email": "[email protected]"}'
wrk.headers["Content-Type"] = "application/json"
wrk.headers["Authorization"] = "Bearer your_api_token"

Run it with -s:

wrk -t 2 -c 50 -d 30s -s post.lua http://your_server_ip/api/items

Rotating between several endpoints

A single URL is rarely representative. The request() function is called for every request and can build a different one each time:

nano paths.lua
local paths = { "/", "/about", "/blog/", "/api/status" }
local i = 0

request = function()
  i = i % #paths + 1
  return wrk.format("GET", paths[i])
end
wrk -t 2 -c 100 -d 30s -s paths.lua http://your_server_ip

Each thread cycles through the paths in order. The host part of the URL is taken from the command line.

Printing a custom latency report

The done() hook runs once at the end and receives summary and latency objects. Latency values are in microseconds. This script prints the percentiles and error counts in a compact, machine-readable form:

nano report.lua
done = function(summary, latency, requests)
  io.write("p50_ms,p90_ms,p99_ms,p999_ms,requests,errors\n")
  local errors = summary.errors.connect + summary.errors.read
               + summary.errors.write + summary.errors.status
               + summary.errors.timeout
  io.write(string.format("%.2f,%.2f,%.2f,%.2f,%d,%d\n",
    latency:percentile(50) / 1000,
    latency:percentile(90) / 1000,
    latency:percentile(99) / 1000,
    latency:percentile(99.9) / 1000,
    summary.requests, errors))
end
wrk -t 2 -c 100 -d 30s -s report.lua http://your_server_ip/ | tail -n 2
p50_ms,p90_ms,p99_ms,p999_ms,requests,errors
7.58,11.69,19.91,34.20,371204,0

You can append that last line to a CSV file after every run to track results over time.

Step 4 - Running a baseline test with siege

siege simulates a number of concurrent users (-c) for a given time (-t, with S, M or H as the unit). The -b (benchmark) flag removes the default random delay between requests, so the test is comparable to wrk:

siege -b -c 50 -t 30S http://your_server_ip/

When the test ends, siege prints a summary. The exact layout depends on the version (newer releases print it as JSON), but the fields are the same:

Transactions:                 290521 hits
Availability:                 100.00 %
Elapsed time:                  29.63 secs
Data transferred:             170.36 MB
Response time:                  0.01 secs
Transaction rate:            9804.96 trans/sec
Throughput:                     5.75 MB/sec
Concurrency:                   49.52
Successful transactions:      290521
Failed transactions:               0
Longest transaction:            0.21
Shortest transaction:           0.00

The most useful fields are:

  • Availability: percentage of requests that succeeded. Anything under 100% means errors or dropped connections.
  • Transaction rate: requests per second.
  • Response time: the average response time. siege does not report percentiles, so use wrk when you need tail latency.
  • Concurrency: the average number of simultaneous connections. If it is far below -c, the client could not open connections fast enough.

siege refuses to go above 255 concurrent users by default. That limit is the limit setting in its configuration file, which you can create in your home directory with siege.config and then edit in ~/.siege/siege.conf.

Step 5 - Replaying a URL list with siege

The main strength of siege is testing many URLs at once. Create a file with one URL per line, in rough proportion to your real traffic (a URL that gets twice the traffic can appear twice):

nano urls.txt
http://your_server_ip/
http://your_server_ip/
http://your_server_ip/about
http://your_server_ip/blog/
http://your_server_ip/api/status
http://your_server_ip/api/items POST {"name": "benchmark"}

The last line shows the siege syntax for POST requests: the URL, the word POST and the body.

Run a test that picks URLs at random (-i, internet mode), with up to 2 seconds of think time per user (-d 2), for 5 minutes:

siege -i -c 100 -d 2 -t 5M -f urls.txt --content-type "application/json"

This does not measure maximum throughput. It shows how the server behaves with 100 users browsing at a human pace, which is closer to production traffic. To send a header on every request, add -H 'Authorization: Bearer your_api_token'.

When logging = true is set in ~/.siege/siege.conf, siege also appends a one-line summary of every run to the log file defined by logfile in the same file, which is convenient for comparing runs.

Step 6 - Finding the server's limit

Run wrk at increasing connection counts and watch where throughput stops growing while latency keeps rising:

for c in 50 100 200 400 800; do
  echo "== $c connections"
  wrk -t 2 -c "$c" -d 30s --latency http://your_server_ip/ | grep -E 'Requests/sec|99%|Socket errors'
done
== 50 connections
     99%   12.40ms
Requests/sec:  11810.22
== 100 connections
     99%   19.88ms
Requests/sec:  12355.49
== 200 connections
     99%   38.02ms
Requests/sec:  12402.73
== 400 connections
     99%   79.51ms
Requests/sec:  12290.14
== 800 connections
     99%  171.30ms
  Socket errors: connect 0, read 0, write 0, timeout 57
Requests/sec:  11544.83

Throughput plateaus at about 12,400 requests per second and the 99th percentile doubles with every step after 100 connections. The server's practical capacity for this page is around 100-200 concurrent connections. During the test, run htop or vmstat 1 on the server to see which resource is saturated.

To make comparisons meaningful, keep the client, URL, threads, connections and duration identical between runs, repeat each test at least three times, and discard a short warm-up run.

Choosing between ab, wrk and siege

ToolBest forLimitations
abQuick single-URL checks, available almost everywhereSingle threaded, HTTP/1.0, one URL per run
wrkMaximum throughput and accurate latency percentiles, scripted requestsHTTP/1.1 only, no built-in URL list (use Lua)
siegeSimulating users over many URLs with think timeNo percentiles, lower maximum load than wrk

Troubleshooting

  • unable to connect to your_server_ip:http Connection refused from wrk: the server is not listening on that port or a firewall is blocking the client. Check with curl -I http://your_server_ip/ first.
  • Many timeout socket errors in wrk: responses took longer than the default 2-second timeout. The server is overloaded at that connection count. Lower -c, or raise the timeout with --timeout 10s if slow responses are expected.
  • [error] socket: ... Too many open files or siege aborted due to excessive socket failure: raise the client limit with ulimit -n 65535 and check the server's own limits (worker_connections in Nginx).
  • Availability below 100% in siege with an otherwise fast server: check the server's access log for 4xx/5xx responses. Rate limiting or a WAF in front of the server is often the cause.

Conclusion

You installed wrk and siege, measured throughput and latency percentiles, scripted POST requests and custom reports in Lua, replayed a weighted URL list, and found the load at which your server saturates. Next, profile the server at that load with perf or strace to see where time goes, enable compression and caching and measure the difference, or repeat the tests from a second region to include real network latency.