When a server stops responding, the cause is almost always one of a handful of things: the network path, the firewall, an exhausted resource (CPU, memory, disk) or a crashed service. This guide gives you a fixed order of checks, from outside the server to inside it, so you can find which one it is quickly instead of guessing. The commands are for Ubuntu 24.04 and Debian 12 and work the same on most other distributions.

Prerequisites

To follow this guide you need:

  • A Linux server that is not responding, and a user with sudo privileges on it.
  • A second machine (your workstation) with ping, mtr, nc, curl and dig to test from outside. On Ubuntu they come from the iputils-ping, mtr-tiny, netcat-openbsd, curl and bind9-dnsutils packages.
  • Access to an out-of-band console in case SSH does not work. On a CubePath VPS, this is the VNC console in the panel.

Throughout the guide, replace your_domain and your_server_ip with your own values.

Step 1 - Finding out what exactly is not responding

"The server is down" can mean very different things. Check from your workstation, from the bottom of the network stack upwards, and stop at the first test that fails.

First, confirm that DNS points to the right place:

dig +short your_domain

If the result is empty or not your_server_ip, the problem is DNS, not the server. Fix the record and wait for its TTL.

Next, test basic reachability:

ping -c 4 your_server_ip
4 packets transmitted, 4 received, 0% packet loss, time 3004ms
rtt min/avg/max/mdev = 11.2/11.6/12.1/0.3 ms

Some servers block ICMP on purpose, so a failed ping alone does not prove the server is down. Test the TCP ports your services use:

nc -zv -w 5 your_server_ip 22
nc -zv -w 5 your_server_ip 443
Connection to your_server_ip 22 port [tcp/ssh] succeeded!
nc: connect to your_server_ip port 443 (tcp) failed: Connection refused

The error message tells you a lot:

ResultMeaning
succeededSomething is listening and the network path is fine.
Connection refusedThe server is up and reachable, but nothing listens on that port: the service has stopped.
timed outPackets are dropped: a firewall, a network problem, or a server too overloaded to answer.

For websites, check the HTTP layer too. A response, even an error, proves the web server is alive:

curl -sS -o /dev/null -w "%{http_code} %{time_total}s\n" https://your_domain
502 0.084s

A 502 or 504 means the web server works but the application behind it does not. Jump to Step 5.

If ping and every port time out, check the network path. mtr combines ping and traceroute and shows where packets are lost:

mtr -rw -c 20 your_server_ip

Loss that starts at one hop and continues to the end points to that hop. Loss on a single intermediate hop that does not carry on to the destination is usually just a router deprioritizing ICMP and can be ignored.

Step 2 - Getting a shell on the server

Try SSH with verbose output. It shows at which stage the connection stops:

ssh -v your_user@your_server_ip
  • If it stops at Connecting to ..., the server is unreachable or port 22 is filtered.
  • If it stops after Connection established or SSH2_MSG_KEXINIT sent and hangs, the server accepts connections but is too overloaded to finish the handshake, which usually means memory exhaustion or extreme load.
  • Permission denied (publickey) is an authentication problem, not a server outage.

If SSH does not work, open the out-of-band console (the VNC console on a CubePath VPS). It works like a physical screen and keyboard, so it is independent of the network configuration, the firewall and SSH. Log in with a local user and password.

The console also shows messages the kernel printed. Text such as Out of memory: Killed process or blocked for more than 120 seconds already points to the cause.

If the console is frozen and does not accept input, the only option is a forced reboot from the panel. After it comes back, use Step 6 to read the logs of the previous boot and find out what happened.

Step 3 - Checking load and CPU

Once you have a shell, start with an overview:

uptime
 10:15:02 up 12 days,  3:41,  1 user,  load average: 14.52, 12.88, 9.10

The three numbers are the average number of processes running or waiting over 1, 5 and 15 minutes. Compare them with the number of CPU cores:

nproc

A load average well above the core count means the server is overloaded. It is rising if the first value is higher than the last one.

Load includes processes waiting for disk, not only for CPU. vmstat tells you which it is:

vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st gu
 9  0      0 182340  51200 902144    0    0     2    10 1850 2400 96  3  1  0  0  0

Read these columns:

  • r high and us + sy near 100: the CPU is saturated by processes.
  • b high and wa high: processes are waiting for the disk (see Step 4).
  • si and so above zero on every line: the server is swapping, so it is short on memory.
  • st high on a virtual machine: the hypervisor is not giving the VM the CPU time it asks for.

Find the processes using the most CPU:

ps -eo pid,user,%cpu,%mem,etime,cmd --sort=-%cpu | head -n 10

For a live view, run top and press P to sort by CPU or M to sort by memory. Press q to exit.

Step 4 - Checking memory, swap and disk

Memory and the OOM killer

free -h
               total        used        free      shared  buff/cache   available
Mem:           3.8Gi       3.6Gi       102Mi        12Mi       190Mi       118Mi
Swap:             0B          0B          0B

Look at the available column, not free: Linux uses spare memory as cache and releases it on demand. When available is close to zero, the kernel starts killing processes. Check whether that already happened:

sudo journalctl -k --since "24 hours ago" | grep -iE "out of memory|oom-kill"
Sep 25 09:58:11 server kernel: Out of memory: Killed process 2211 (php-fpm8.3) total-vm:812345kB, anon-rss:602112kB

A killed process explains why a service stopped. Find the biggest memory users:

ps -eo pid,user,%mem,rss,cmd --sort=-rss | head -n 10

Disk space and inodes

A full disk breaks databases, log writers and package managers, and can prevent SSH logins:

df -h
Filesystem      Size  Used Avail Use% Mounted on
/dev/vda1        38G   38G     0 100% /

A disk can also run out of inodes (file entries) while still showing free space, typically because of millions of small cache or session files:

df -i

Find what is using the space on the full file system. The -x option stays on one file system:

sudo du -xh --max-depth=2 / 2>/dev/null | sort -rh | head -n 15

Disk I/O

If vmstat showed high wa, look at the devices. iostat comes with the sysstat package:

sudo apt install sysstat
iostat -xz 1 3

A %util close to 100 and a high await (milliseconds per request) mean the disk is saturated. Find the processes responsible:

sudo pidstat -d 1 5

Step 5 - Checking services and ports

List services that systemd marks as failed:

systemctl --failed
  UNIT          LOAD   ACTIVE SUB    DESCRIPTION
● mysql.service loaded failed failed MySQL Community Server

Check the state of the service you care about and its latest log lines:

sudo systemctl status nginx --no-pager
sudo journalctl -u nginx -n 50 --no-pager

Confirm which processes are listening on which ports. A service can be running but bound to the wrong address, for example 127.0.0.1 instead of all interfaces:

sudo ss -tlnp
State  Recv-Q Send-Q Local Address:Port  Peer Address:Port Process
LISTEN 0      511          0.0.0.0:80         0.0.0.0:*     users:(("nginx",pid=812,fd=6))
LISTEN 0      4096       127.0.0.1:3306       0.0.0.0:*     users:(("mysqld",pid=901,fd=21))
LISTEN 0      4096               *:22               *:*     users:(("sshd",pid=640,fd=3))

Count connections by state. Thousands of connections in SYN-RECV or from a single address point to a flood or a misbehaving client rather than a local fault:

ss -s
ss -tn state established | awk 'NR>1 {sub(/:[0-9]+$/, "", $4); print $4}' | sort | uniq -c | sort -rn | head

Finally, check the firewall. A rule added by mistake produces exactly the timed out symptom from Step 1:

sudo ufw status verbose

If the server does not use UFW, list the active rules with sudo nft list ruleset.

Step 6 - Reading the logs

The systemd journal collects kernel, service and system messages in one place. Show errors and worse from the current boot:

sudo journalctl -b -p err --no-pager | tail -n 50

If the server rebooted or you had to force a reboot, the interesting messages are in the previous boot. List the recorded boots and read the end of the previous one:

journalctl --list-boots
sudo journalctl -b -1 -n 100 --no-pager

This only works when the journal is stored on disk. Ubuntu does this by default if /var/log/journal exists. If it does not, create it so the next incident leaves a trace:

sudo mkdir -p /var/log/journal
sudo systemctl restart systemd-journald

Also check the kernel messages for hardware and file system errors:

sudo dmesg -T --level=err,warn | tail -n 30

Messages such as I/O error, EXT4-fs error or Remounting filesystem read-only indicate a storage problem. Stop writing to the disk and contact your provider's support.

Step 7 - Fixing the most common causes

Once you know the cause, apply the matching fix.

A process is using all the CPU or memory. Stop it gracefully first, and force it only if it does not exit:

sudo kill PID
sudo kill -9 PID

If the process belongs to a service, restart the service instead: sudo systemctl restart service_name.

A service has stopped. Restart it and confirm that it stays up:

sudo systemctl restart mysql
sudo systemctl status mysql --no-pager

If it crashes again, the log from journalctl -u explains why. To have systemd restart it automatically in the future, create an override:

sudo systemctl edit mysql
[Service]
Restart=on-failure
RestartSec=5

The disk is full. Free space safely before touching application data. Shrink the journal and clean the package cache:

sudo journalctl --vacuum-size=200M
sudo apt clean
sudo apt autoremove --purge

Then remove what du showed in Step 4, typically old backups, rotated logs or application caches. Deleting a file that a process still has open does not free the space; find such files and restart the process that holds them:

sudo lsof +L1

The server keeps running out of memory. Reduce the memory use of the service (fewer PHP-FPM or database workers), move to a plan with more RAM, or add a swap file as a buffer so the kernel does not kill processes immediately:

sudo fallocate -l 2G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

Verify with free -h that the swap is active.

The firewall blocks a service. Allow the port and check the result:

sudo ufw allow 443/tcp
sudo ufw status

Step 8 - Preparing for the next incident

A few changes make the next diagnosis much faster:

  • Keep resource history. Enable sysstat data collection so you can see what CPU, memory and disk looked like before the incident. Answer "Yes" when asked:

    sudo dpkg-reconfigure sysstat
    

    After the next incident, sar -q shows the load history and sar -r the memory history for the day.

  • Keep a persistent journal as described in Step 6.

  • Monitor from outside. An external uptime check on your main ports alerts you before users do.

  • Set Restart=on-failure on your critical services, so a crash becomes a short blip instead of an outage.

Conclusion

Diagnosing an unresponsive server comes down to working through the layers in order: DNS and network from outside, then access, CPU, memory, disk, services and logs from inside. The first check that fails usually points straight to the cause. As next steps, set up external monitoring with alerts, configure sysstat and a persistent journal on all your servers, and write down the commands from this guide in a runbook your team can follow during an incident.