Linux uses spare RAM as disk cache, so a server that looks "full" in free is often perfectly healthy. The real problems are when available memory runs out, the system starts swapping heavily, or the kernel's OOM killer terminates a process. In this tutorial you will read memory metrics correctly, find which process or service uses the memory, check for leaks and OOM kills, and apply fixes on Ubuntu 24.04.

Prerequisites

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS. The commands also work on Debian 12.
  • A non-root user with sudo privileges.
  • sysstat for pidstat (sudo apt install sysstat) if you want to follow the leak check in Step 5.

Step 1 - Reading free correctly

Start with free in human-readable units:

free -h
               total        used        free      shared  buff/cache   available
Mem:           7.7Gi       5.1Gi       312Mi       214Mi       2.6Gi       2.3Gi
Swap:          2.0Gi       1.4Gi       620Mi

The column that matters is available: an estimate of how much memory new work can use without swapping, including cache the kernel can drop. free being low is normal, because idle RAM is used as buff/cache.

Use these rules of thumb:

  • available above 10-15% of total: memory is fine, even if free is near zero.
  • available low and swap used growing: real memory pressure. Continue with the next steps.
  • shared large: memory used by tmpfs mounts or shared memory segments (common with PostgreSQL). Check tmpfs with df -h -t tmpfs.

Step 2 - Checking whether the system is swapping

Swap being used is not a problem by itself: the kernel moves rarely used pages there. What hurts performance is constant swapping in and out. Watch it with vmstat, once per second:

vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 2  1 1468004 318212  84120 2641800  820  1204  1900  1320 2410 3820 22 9 41 28  0

The si (swap in) and so (swap out) columns are in KiB per second. Values that stay at zero or near it are fine. Sustained values in the hundreds or thousands, as above, mean the working set does not fit in RAM, and the high wa (I/O wait) shows the server is slowing down because of it.

Step 3 - Finding the processes that use memory

List processes by resident memory (RSS, the RAM actually in use), largest first:

ps -eo pid,user,rss,vsz,cmd --sort=-rss | head -10
    PID USER        RSS    VSZ CMD
   1043 mysql   2412880 4103220 /usr/sbin/mysqld
  18342 www-data 612404  903112 php-fpm: pool www
  18343 www-data 598220  903112 php-fpm: pool www

Values are in KiB. VSZ (virtual size) includes memory the process reserved but never touched, so ignore it when judging usage.

RSS double counts memory shared between processes, such as libraries and the shared pages of php-fpm or Apache workers. For a fair total per application, sum RSS per command name:

ps -eo rss,comm --no-headers | awk '{sum[$2]+=$1} END {for (c in sum) printf "%8.0f MiB  %s\n", sum[c]/1024, c}' | sort -rn | head

For the most accurate view, use smem, which reports PSS (proportional set size): shared pages are divided among the processes that share them:

sudo apt install smem
sudo smem -r -s pss -k | head -15

Step 4 - Finding which service uses the memory

Many applications run as several processes. systemd-cgtop aggregates memory per service (control group); -m sorts by memory:

sudo systemd-cgtop -m

For one service, systemctl status shows its current and peak memory:

systemctl status mysql
● mysql.service - MySQL Community Server
     Active: active (running) since Mon 2026-09-15 07:40:12 UTC; 1 week 3 days ago
     Memory: 2.3G (peak: 2.6G)

If the total of all services is much lower than used in free, the memory is being used by the kernel itself. Check the slab caches:

sudo slabtop -o | head -15

A very large dentry or inode_cache is usually reclaimable cache from scanning many files and is released under pressure. Check SReclaimable versus SUnreclaim:

grep -E '^(MemTotal|MemAvailable|Shmem|SReclaimable|SUnreclaim|SwapTotal|SwapFree):' /proc/meminfo

Step 5 - Checking for a memory leak

A leak shows up as a process whose RSS grows steadily and never goes down, even when load is stable. Record its memory once a minute with pidstat:

pidstat -r -p 18342 60
10:00:01 AM   UID       PID  minflt/s  majflt/s     VSZ     RSS   %MEM  Command
10:01:01 AM    33     18342     12.40      0.00  903112  612404   7.60  php-fpm8.3
10:02:01 AM    33     18342     11.90      0.00  911304  620596   7.70  php-fpm8.3
10:03:01 AM    33     18342     13.10      0.00  919496  628788   7.80  php-fpm8.3

Let it run for a while (press Ctrl+C to stop). An RSS that keeps growing by a similar amount every interval is a leak. majflt/s above zero means the process is reading its own pages back from swap or disk, a sign of memory pressure.

For worker-based applications, the usual mitigation while the bug is fixed is to recycle workers after a number of requests: pm.max_requests in PHP-FPM, MaxConnectionsPerChild in Apache, or --max-requests in Gunicorn.

Step 6 - Checking for OOM killer events

When memory and swap run out, the kernel's OOM killer terminates the process with the highest "badness" score to save the system. Search the kernel log:

sudo journalctl -k | grep -iE 'out of memory|oom-kill|killed process'
Sep 25 03:14:07 web01 kernel: oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0,global_oom,task_memcg=/system.slice/mysql.service,task=mysqld,pid=1043,uid=112
Sep 25 03:14:07 web01 kernel: Out of memory: Killed process 1043 (mysqld) total-vm:4103220kB, anon-rss:2398812kB, file-rss:0kB, shmem-rss:0kB, UID:112 pgtables:5420kB oom_score_adj:0

The message names the process that was killed, but not necessarily the one that caused the pressure. The OOM killer picks the largest target, so a database is often killed because another process grew. Compare the time of the kill with the application logs and with your sar -r history if sysstat collection is enabled.

systemd also reports services killed this way:

systemctl list-units --state=failed

Step 7 - Adding swap as a safety net

A small swap file gives the kernel room to move idle pages out of RAM and turns a sudden OOM kill into a gradual slowdown you can react to. Check whether you already have swap:

swapon --show

If the command prints nothing, create a 2 GB swap file:

sudo fallocate -l 2G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile

Make it permanent by adding it to /etc/fstab:

echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

Verify:

swapon --show
NAME      TYPE SIZE USED PRIO
/swapfile file   2G   0B   -2

Step 8 - Limiting a service's memory

To stop one service from taking memory from everything else, cap it with systemd. MemoryHigh throttles the service and makes the kernel reclaim its memory aggressively above the threshold; MemoryMax is a hard limit at which the OOM killer acts only inside that service:

sudo systemctl set-property app-worker.service MemoryHigh=1536M MemoryMax=2G

The change applies immediately and is saved as a drop-in. Confirm it:

systemctl show app-worker.service -p MemoryHigh -p MemoryMax
MemoryHigh=1610612736
MemoryMax=2147483648

Also check the application's own memory settings, which are usually the real fix: innodb_buffer_pool_size for MySQL, shared_buffers for PostgreSQL, pm.max_children for PHP-FPM, or the JVM -Xmx value. Their sum across all services must fit in RAM with room left for the operating system and cache.

Conclusion

You learned to read available instead of free, detect real swapping with vmstat, rank processes and services by memory with ps, smem and systemd-cgtop, confirm leaks with pidstat -r and find OOM kills in the kernel log. With a swap file and per-service limits in place, one misbehaving process can no longer take down the whole server. As next steps, tune the memory settings of your database and application workers, and enable sysstat collection so sar -r keeps a history of memory usage.