Linux's default TCP settings are a compromise that works for most machines, but a server that moves a lot of data over long distances, or handles thousands of connections per second, can hit limits in the default buffer sizes, the congestion control algorithm or the connection queues. In this tutorial you will measure your current throughput with iperf3, size TCP buffers from the bandwidth-delay product, switch congestion control to BBR, adjust queues and port ranges for busy servers, and verify each change with ss and nstat on Ubuntu 24.04.
Prerequisites
To follow this tutorial, you will need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS, with a non-root user that has
sudoprivileges. - A second Linux machine to test against, ideally at a realistic distance from the server (for example a machine in another region, or your own network).
- Basic familiarity with sysctl and
/etc/sysctl.d/.
NoteTCP tuning mostly helps two cases: high throughput over high-latency paths (large file transfers, backups, streaming to distant users), and servers with very high connection rates. If your traffic is small requests on a low-latency network, the defaults are usually fine, and you should measure before changing anything.
Step 1 - Measuring baseline throughput
Install iperf3 on both machines:
sudo apt update
sudo apt install iperf3
On the server, open the test port temporarily and start iperf3 in server mode:
sudo ufw allow 5201/tcp
iperf3 -s
On the client, measure latency first, then run a 30-second test with one TCP stream. Replace your_server_ip with the server's IP:
ping -c 10 your_server_ip
iperf3 -c your_server_ip -t 30
[ ID] Interval Transfer Bitrate Retr
[ 5] 0.00-30.00 sec 1.02 GBytes 292 Mbits/sec 214 sender
[ 5] 0.00-30.03 sec 1.02 GBytes 291 Mbits/sec receiver
Run it again with -R to measure the opposite direction (server sending to client), which is the direction that matters for a server delivering content. Write down the bitrate, the retransmissions (Retr) and the average round-trip time from ping. A single stream that is far below your link speed on a high-latency path usually points to buffer limits.
Step 2 - Calculating the bandwidth-delay product
TCP can only have as much unacknowledged data in flight as its window allows. To fill a link, the window must be at least the bandwidth-delay product (BDP):
BDP (bytes) = bandwidth (bits/s) x round-trip time (s) / 8
A few examples:
| Link | RTT | BDP |
|---|---|---|
| 1 Gbit/s | 10 ms | 1.25 MB |
| 1 Gbit/s | 100 ms | 12.5 MB |
| 10 Gbit/s | 50 ms | 62.5 MB |
Check the current limits:
sysctl net.ipv4.tcp_rmem net.ipv4.tcp_wmem net.core.rmem_max net.core.wmem_max
net.ipv4.tcp_rmem = 4096 131072 6291456
net.ipv4.tcp_wmem = 4096 16384 4194304
net.core.rmem_max = 212992
net.core.wmem_max = 212992
The three values of tcp_rmem and tcp_wmem are the minimum, default and maximum buffer size per socket, in bytes. The kernel auto-tunes each connection's buffer between those values. With a 6 MB receive maximum and a 4 MB send maximum, a single connection cannot fill 1 Gbit/s beyond about 40 ms of RTT.
Step 3 - Raising the TCP buffer limits
Create a file for your network tuning:
sudo nano /etc/sysctl.d/90-tcp-tuning.conf
Set the maximum to cover your largest expected BDP. 16 MB covers 1 Gbit/s up to roughly 100 ms, which fits most internet-facing servers:
# TCP buffers: min, default, max (bytes)
net.ipv4.tcp_rmem = 4096 131072 16777216
net.ipv4.tcp_wmem = 4096 16384 16777216
# Upper limit for applications that set SO_RCVBUF / SO_SNDBUF themselves
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216
Only raise the maximum values; leave the minimum and default alone. The defaults are what most connections use, and raising them multiplies memory use by the number of connections. Auto-tuning only grows the buffer of connections that actually need it. For 10 Gbit/s links with high latency, use 64 MB (67108864) as the maximum instead.
net.core.rmem_max and wmem_max do not affect auto-tuning; they cap the size that an application can request explicitly with setsockopt(). Raising them keeps such applications from being limited to about 200 KB.
Apply and verify:
sudo sysctl --system
sysctl net.ipv4.tcp_rmem
net.ipv4.tcp_rmem = 4096 131072 16777216
Step 4 - Enabling BBR congestion control
The congestion control algorithm decides how fast TCP sends. The default, CUBIC, reduces its rate sharply when it detects packet loss, so a small amount of random loss on a long path keeps throughput far below the link capacity. BBR, developed by Google, instead models the path's bandwidth and RTT, and keeps throughput high on lossy, high-latency paths. It only changes how this server sends, so it helps downloads from the server, not uploads to it.
Check which algorithms are available:
sysctl net.ipv4.tcp_available_congestion_control
net.ipv4.tcp_available_congestion_control = reno cubic
BBR is built as the tcp_bbr module on Ubuntu kernels. Load it now and at every boot:
sudo modprobe tcp_bbr
echo tcp_bbr | sudo tee /etc/modules-load.d/bbr.conf
Add the congestion control settings to your file:
sudo nano /etc/sysctl.d/90-tcp-tuning.conf
# Congestion control
net.core.default_qdisc = fq
net.ipv4.tcp_congestion_control = bbr
The fq queueing discipline paces packets per flow. BBR works without it on current kernels, but fq is the recommended companion. Apply the settings and verify:
sudo sysctl --system
sysctl net.ipv4.tcp_congestion_control net.ipv4.tcp_available_congestion_control
net.ipv4.tcp_congestion_control = bbr
net.ipv4.tcp_available_congestion_control = reno cubic bbr
default_qdisc only applies when a network interface's queue is created, so existing interfaces keep their current qdisc until the next reboot. After rebooting, find your main interface and check its qdisc:
ip route show default
tc qdisc show dev eth0
Replace eth0 with the interface name shown after dev in the first command. The output should include fq.
New connections use BBR as soon as the setting is applied. Confirm it on a live connection by running an iperf3 test from the client with -R and, on the server, inspecting the connection:
ss -tin 'sport = :5201'
ESTAB 0 1864216 203.0.113.20:5201 198.51.100.7:51844
bbr wscale:9,9 rto:228 rtt:26.1/0.4 mss:1448 cwnd:512 bytes_sent:1307362688 ... delivery_rate 912Mbps ...
The line shows bbr as the algorithm, along with the current RTT, congestion window and delivery rate.
Step 5 - Tuning queues for high connection rates
On servers that accept many new connections per second (load balancers, reverse proxies, API servers), two queues can overflow: the SYN queue (half-open connections) and the accept queue (connections waiting for the application to call accept()). Check the kernel counters first:
nstat -az TcpExtListenOverflows TcpExtListenDrops TcpExtTCPReqQFullDrop
If they are 0 after a period of peak traffic, your queues are large enough. If they increase, look at the listening sockets:
ss -ltn
For sockets in the LISTEN state, Send-Q shows the configured backlog and Recv-Q the number of connections currently waiting in the accept queue. A Recv-Q close to Send-Q means the application is not accepting connections fast enough or its backlog is too small.
Raise the kernel limits in your file:
# Connection queues
net.core.somaxconn = 8192
net.ipv4.tcp_max_syn_backlog = 8192
net.core.netdev_max_backlog = 16384
somaxconn is only a ceiling: the application must request a larger backlog as well. In Nginx, add it to the listen directive (the default is 511):
listen 443 ssl backlog=8192;
Keep net.ipv4.tcp_syncookies = 1, which is the Ubuntu default. It lets the server keep accepting connections when the SYN queue is full, for example during a SYN flood.
Step 6 - Handling TIME_WAIT and port exhaustion
Every TCP connection that a host closes first stays in the TIME_WAIT state for 60 seconds. This duration is fixed in the kernel. On the server side, many TIME_WAIT sockets are harmless. The problem appears on machines that open many outgoing connections to the same destination (a reverse proxy talking to a backend, an application talking to a database), because each connection uses one local port from a limited range.
Count the sockets in TIME_WAIT and check the port range:
ss -Htan state time-wait | wc -l
sysctl net.ipv4.ip_local_port_range net.ipv4.tcp_tw_reuse
18342
net.ipv4.ip_local_port_range = 32768 60999
net.ipv4.tcp_tw_reuse = 2
The default range gives 28,232 ports per destination IP and port. If your application logs errors such as Cannot assign requested address, widen the range and allow reuse of TIME_WAIT sockets for outgoing connections:
# Outgoing connections
net.ipv4.ip_local_port_range = 10240 65535
net.ipv4.tcp_tw_reuse = 1
tcp_tw_reuse = 1 is safe: it only applies to new outgoing connections and relies on TCP timestamps to reject old packets. The default value 2 enables it only for loopback traffic. If a service listens on a port inside the new range, reserve it with net.ipv4.ip_local_reserved_ports.
The best fix, however, is to stop opening a new connection per request. Enable keepalive connections to your upstreams (for example keepalive 32; in an Nginx upstream block) or use a connection pool in your application.
Step 7 - Improving long-lived and idle connections
Two more settings help specific traffic patterns:
# Keep the congestion window after a connection has been idle
net.ipv4.tcp_slow_start_after_idle = 0
# Detect path MTU black holes and lower the packet size automatically
net.ipv4.tcp_mtu_probing = 1
By default, a connection that sits idle for a moment falls back to slow start, which hurts HTTP/2, WebSocket and database connections that send in bursts. tcp_mtu_probing = 1 helps when a firewall on the path drops the ICMP messages that Path MTU Discovery depends on, which otherwise shows up as connections that hang once they send large packets (common with VPNs and tunnels).
Several settings from older guides should be left alone:
| Setting | Why leave it |
|---|---|
net.ipv4.tcp_tw_recycle | Removed from the kernel in 4.12. It broke clients behind NAT. |
net.ipv4.tcp_timestamps, tcp_sack, tcp_window_scaling | Already enabled. Disabling them reduces performance. |
net.ipv4.tcp_fin_timeout | Controls FIN_WAIT_2, not TIME_WAIT, so lowering it does not reduce TIME_WAIT sockets. |
net.ipv4.tcp_mem | Auto-sized from RAM at boot. Wrong values can cause drops under memory pressure. |
Step 8 - Applying and verifying the complete configuration
Your final /etc/sysctl.d/90-tcp-tuning.conf should contain only the sections that apply to your workload. A complete version for a busy internet-facing server looks like this:
# TCP buffers: min, default, max (bytes)
net.ipv4.tcp_rmem = 4096 131072 16777216
net.ipv4.tcp_wmem = 4096 16384 16777216
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216
# Congestion control
net.core.default_qdisc = fq
net.ipv4.tcp_congestion_control = bbr
# Connection queues
net.core.somaxconn = 8192
net.ipv4.tcp_max_syn_backlog = 8192
net.core.netdev_max_backlog = 16384
# Outgoing connections
net.ipv4.ip_local_port_range = 10240 65535
net.ipv4.tcp_tw_reuse = 1
# Long-lived connections and MTU
net.ipv4.tcp_slow_start_after_idle = 0
net.ipv4.tcp_mtu_probing = 1
Apply it, reboot once to confirm the settings and the fq qdisc survive a restart, then repeat the tests from Step 1 in both directions, with one stream and with several (-P 4):
iperf3 -c your_server_ip -t 30 -R
iperf3 -c your_server_ip -t 30 -R -P 4
Compare the bitrate and the number of retransmissions with your baseline. On the server, you can also track retransmissions over time, which should stay low under normal load:
nstat -az TcpRetransSegs TcpOutSegs
When you finish testing, stop iperf3 with Ctrl+C and close the test port:
sudo ufw delete allow 5201/tcp
Troubleshooting
sysctl: setting key "net.ipv4.tcp_congestion_control": No such file or directory at boot. The tcp_bbr module was not loaded when the setting was applied. Check that /etc/modules-load.d/bbr.conf contains tcp_bbr and that lsmod | grep bbr shows the module after boot.
Throughput did not improve. The limit may be elsewhere: the client's receive buffer (tune the client too), CPU usage on either end (watch top during the test), a rate limit on the network path, or disk speed if you test with real file transfers. iperf3 with -P 4 reaching a much higher total than a single stream points to a per-connection limit such as buffers.
Many retransmissions after enabling BBR. BBR tolerates some loss by design and may show more retransmissions than CUBIC for the same or higher throughput. Judge by throughput and application latency, not retransmissions alone.
nf_conntrack: table full, dropping packet in dmesg. On servers with a stateful firewall and many connections, the connection tracking table fills up. Check the usage with sysctl net.netfilter.nf_conntrack_count net.netfilter.nf_conntrack_max and raise net.netfilter.nf_conntrack_max if needed.
Conclusion
You measured your baseline, sized TCP buffers from the bandwidth-delay product, switched to BBR with fq, and adjusted queues, ports and idle behavior only where the counters showed a need. Keep the configuration in a single file so it is easy to review and reproduce. As next steps, apply the same buffer settings to clients that pull large transfers from this server, tune per-service limits such as Nginx worker connections, and monitor TcpRetransSegs and listen overflows alongside your other metrics.
