When a Linux server is compromised, the first hour decides whether you end up with a clear picture of what happened or with a wiped disk and a guess. This playbook gives you a fixed order of work for a suspected intrusion on an Ubuntu 24.04 server (VPS or bare metal): preparation you do in advance, triage, containment, evidence collection, eradication by rebuilding, recovery and a post-incident review. Commands are for Ubuntu and Debian; the Rocky Linux equivalents are noted where they differ.

Treat it as a template. Copy it into your own documentation, fill in names, contacts and paths, and rehearse it before you need it.

Prerequisites

To use this playbook you need:

  • Servers running Ubuntu 24.04 LTS (or Debian 12, Rocky Linux 9) and a user with sudo privileges on them.
  • Out-of-band access that does not depend on SSH, such as the web console of your provider's control panel (on CubePath, the VPS console in the panel) or IPMI on bare metal. If you isolate the network, this is how you get back in.
  • A separate machine to receive evidence, with enough disk for the logs of the affected server. It must not trust the affected server (no SSH keys from it, no shared credentials).
  • An agreed communication channel for the incident that does not run on the affected infrastructure.

Phase 0 - Preparing before an incident

Most of what makes a response work is done beforehand. On every server, and in your runbook, have the following in place.

Tools installed in advance. Installing packages on a compromised host means downloading over a network you are about to cut and trusting a package manager the attacker may control. Install the basics now:

sudo apt update
sudo apt install nftables lsof psmisc debsums

On Rocky Linux use sudo dnf install nftables lsof psmisc.

Logs that survive the host. An attacker with root can delete /var/log. Ship logs to another system (a central syslog, journald remote or your logging stack) so the record of what happened does not depend on the compromised machine.

A collection script. Install the script from Phase 3 on every server now, so that during an incident you run one known command instead of typing dozens by hand.

Contacts and decisions. Write down who leads an incident, who can approve taking a production server offline, who handles customer communication and when legal or data protection obligations apply. Keep this list outside the systems it describes.

Phase 1 - Detection and triage

Signals that start this playbook include unknown processes using CPU (typical for cryptocurrency miners), outgoing traffic you cannot explain, new user accounts or SSH keys, modified system binaries, alerts from your provider about abuse from your IP, or a web application serving content you did not publish.

Start an incident log immediately: a plain text file on your own workstation, not on the server. Record every command you run and every observation with a UTC timestamp. It becomes the basis of the timeline and the report.

Connect to the server and record who and what is active right now. Do not reboot and do not kill anything yet: memory, processes and connections are the most volatile evidence and disappear first.

date -u
w
last -F -n 30

Look for processes you do not recognize, especially ones running from /tmp, /var/tmp or /dev/shm, or whose binary was deleted after starting:

ps auxwwf
sudo find /proc -maxdepth 2 -name exe -lname '*(deleted)' 2>/dev/null

A result such as /proc/4121/exe means process 4121 is running from a file that no longer exists on disk, a common trick to hide malware.

Check network connections and the processes behind them:

sudo ss -tunap

Look for established connections to unfamiliar addresses and for listeners on unexpected ports.

Check persistence mechanisms: cron, systemd units and timers, and SSH keys. The -newer /etc/hostname test lists unit files changed after the system was installed:

sudo ls -la /etc/cron.d /var/spool/cron/crontabs
systemctl list-timers --all --no-pager
sudo find /etc/systemd/system /usr/lib/systemd/system -name '*.service' -newer /etc/hostname
sudo find / -xdev -name authorized_keys -exec ls -l {} \;

Finally, compare installed files against the checksums shipped in the packages. Lines for files under /usr/bin, /usr/sbin or /usr/lib marked with 5 (checksum mismatch) are a strong sign of tampering; changes to configuration files (marked c) are usually your own edits:

sudo dpkg --verify
??5??????   /usr/bin/ps
??5?????? c /etc/ssh/sshd_config

On Rocky Linux use sudo rpm -Va.

If triage confirms or strongly suggests a compromise, declare an incident, notify the incident lead and move to containment.

Phase 2 - Containment

The goal is to stop the attacker from doing more damage or exfiltrating more data while keeping the machine running so evidence stays in memory.

Isolating the network with nftables

The most reliable isolation happens outside the server, in a provider-level firewall or at the switch, because an attacker with root can undo anything you configure on the host. If that is not available, add an nftables table that drops all traffic except your own SSH session. A drop verdict in any nftables base chain is final, so this table blocks traffic regardless of the UFW or firewalld rules already loaded.

Find the public IP you are connecting from (for example, from echo $SSH_CLIENT on the server), then write the ruleset to a file. Loading it from a file applies it atomically, which avoids locking yourself out halfway through:

sudo nano /root/ir-contain.nft
table inet ir_contain {
    chain input {
        type filter hook input priority -10; policy drop;
        iif "lo" accept
        ct state established,related accept
        ip saddr 198.51.100.10 tcp dport 22 accept
    }
    chain output {
        type filter hook output priority -10; policy drop;
        oif "lo" accept
        ct state established,related accept
    }
}

Replace 198.51.100.10 with your own IP address. If you connect over IPv6, use ip6 saddr with your IPv6 address instead. Keep the provider console open in another window, then load the table:

sudo nft -f /root/ir-contain.nft
sudo nft list table inet ir_contain

Your current SSH session keeps working because its packets belong to an established connection. New outbound connections, including the attacker's command and control traffic, DNS and package downloads, are now dropped. To lift the isolation later, delete the table:

sudo nft delete table inet ir_contain

Cutting existing connections and sessions

With new connections blocked, close established connections to a hostile address:

sudo ss -K dst 203.0.113.66

Freeze a suspicious process instead of killing it, so its memory can still be examined:

sudo kill -STOP 4121

Lock a compromised account and end its sessions. -L locks the password and -e 1 expires the account, which also blocks logins with SSH keys:

sudo usermod -L -e 1 webapp
sudo pkill -KILL -u webapp

Record each action with its timestamp in the incident log.

Phase 3 - Collecting evidence

Collect evidence from the most volatile to the least volatile: process and network state first, then logs and files. Store it with checksums so you can later show it has not changed.

The collection script

Install this script in advance on every server (Phase 0):

sudo nano /usr/local/sbin/ir-collect
#!/usr/bin/env bash
# ir-collect: capture volatile state and logs from a Linux host for incident response.
# Usage: sudo ir-collect [output_base_dir]
set -uo pipefail

if [[ $EUID -ne 0 ]]; then
  echo "Run as root: sudo $0" >&2
  exit 1
fi

base="${1:-/root}"
name="ir-$(hostname -s)-$(date -u +%Y%m%dT%H%M%SZ)"
out="${base}/${name}"
mkdir -p -m 700 "$out"

# run <file> <command...>: save a command's output, never abort on errors.
run() {
  local file="$1"; shift
  { echo "# $(date -u +%FT%TZ) $*"; "$@"; } > "${out}/${file}" 2>&1
}

run date.txt           date -u
run uptime.txt         uptime
run who.txt            w
run last.txt           last -F -n 100
run lastb.txt          lastb -F -n 100
run ps.txt             ps auxwwf
run deleted-exe.txt    find /proc -maxdepth 2 -name exe -lname '*(deleted)'
run ss.txt             ss -tunap
run lsof-net.txt       lsof -nP -i
run ip-addr.txt        ip addr
run ip-route.txt       ip route
run ip-neigh.txt       ip neigh
run nft-ruleset.txt    nft list ruleset
run lsmod.txt          lsmod
run mounts.txt         findmnt
run units.txt          systemctl list-units --all --no-pager
run timers.txt         systemctl list-timers --all --no-pager
run crontabs.txt       ls -la /etc/cron.d /etc/cron.hourly /etc/cron.daily /var/spool/cron/crontabs
run tmp-files.txt      find /tmp /var/tmp /dev/shm -xdev -type f -printf '%TY-%Tm-%Td %TH:%TM %u %m %s %p\n'
run recent-system.txt  find /etc /usr/bin /usr/sbin /usr/lib/systemd -xdev -type f -mtime -14 -printf '%TY-%Tm-%Td %TH:%TM %u %m %p\n'
run authorized-keys.txt find / -xdev -name authorized_keys -printf '%TY-%Tm-%Td %TH:%TM %u %p\n' -exec cat {} \;
run journal.txt        journalctl --since '-7 days' --no-pager
if command -v dpkg >/dev/null; then
  run pkg-verify.txt   dpkg --verify
else
  run pkg-verify.txt   rpm -Va
fi

# Copy on-disk logs as they are.
cp -a /var/log "${out}/var-log" 2>> "${out}/errors.txt"

# Checksums of everything collected, then a single archive.
(cd "$out" && find . -type f ! -name SHA256SUMS -exec sha256sum {} + > SHA256SUMS)
tar -czf "${out}.tar.gz" -C "$base" "$name"
sha256sum "${out}.tar.gz" | tee "${out}.tar.gz.sha256"
echo "Evidence archive: ${out}.tar.gz"

Make it executable and owned by root:

sudo chown root:root /usr/local/sbin/ir-collect
sudo chmod 700 /usr/local/sbin/ir-collect

The script deliberately does not use set -e: during an incident a missing command or unreadable file must not stop the rest of the collection. Every command's output goes to its own file with a timestamp header.

Running it and moving the evidence off the host

Run the collection:

sudo /usr/local/sbin/ir-collect /root
3f5b0c...e91a  /root/ir-web01-20260925T101502Z.tar.gz
Evidence archive: /root/ir-web01-20260925T101502Z.tar.gz

Write down the SHA-256 value in your incident log. Then pull the archive from your evidence machine instead of pushing it from the compromised host, so the affected server never gets credentials for your evidence store. Your SSH login is still allowed by the containment rules, but /root is not readable by your user, so first place a copy in your home directory:

sudo install -m 600 -o your_user /root/ir-web01-20260925T101502Z.tar.gz /home/your_user/

Then, on the evidence machine:

scp [email protected]:ir-web01-20260925T101502Z.tar.gz .
sha256sum ir-web01-20260925T101502Z.tar.gz

The checksum on the evidence machine must match the one you recorded.

Memory and disk images

For incidents that may involve law enforcement, insurance or a detailed malware analysis, also capture:

  • Memory: a full RAM image with a tool such as Microsoft's AVML or LiME, brought in from trusted media and written to external storage. This is the only way to analyze fileless malware or recover keys.
  • Disk: a snapshot of the disk from your provider or hypervisor, taken while the system is still running or right after powering it off. Analyze copies, never the original.

If you cannot do this properly, do not improvise: an incomplete, undocumented image is of little value. The collection above is enough to answer most questions about how the attacker got in.

Phase 4 - Eradication: rebuild, do not clean

Once an attacker has had root on a server, you cannot prove that you have found every backdoor. Kernel modules, modified libraries, a replaced sshd or a line in a unit file you never look at can all survive a cleanup. The reliable fix is to rebuild:

  1. Provision a new server from a clean, current image.
  2. Apply your configuration with your usual automation (Ansible, cloud-init or your own runbook), not by copying files from the compromised host.
  3. Restore application data from a backup taken before the earliest sign of compromise. Treat data from later backups as untrusted and review it before use.
  4. Deploy the application from your repository, not from the old server's directories.

Before bringing the new server online, fix the entry point you identified, or it will be compromised again the same way. Common ones are:

  • An outdated web application or plugin with a known vulnerability.
  • SSH with password authentication and a weak or reused password.
  • A service exposed to the internet that should only listen on localhost or a private network (databases, Redis, admin panels).
  • Leaked credentials or API keys, for example committed to a public repository.

Rotating credentials

Assume every secret the compromised server could read is known to the attacker. Rotate:

  • SSH keys and passwords of every account on the server, and any key the server used to log in elsewhere.
  • Database passwords, API keys, tokens and cloud credentials stored in configuration files or environment variables.
  • TLS private keys, then reissue the certificates.
  • Passwords of application users if the password hashes were accessible.

Phase 5 - Recovery

Bring the rebuilt service back gradually and watch it closely.

Harden SSH on the new server with a drop-in file instead of editing the main configuration:

sudo nano /etc/ssh/sshd_config.d/10-hardening.conf
PermitRootLogin no
PasswordAuthentication no
KbdInteractiveAuthentication no
MaxAuthTries 3
AllowUsers your_user

Check the syntax before restarting, so a mistake does not lock you out:

sudo sshd -t && sudo systemctl restart ssh

On Rocky Linux the service is called sshd.

Enable the firewall with only the ports the service needs, point DNS or the load balancer at the new server, and keep the old one isolated (not deleted) until the investigation is finished. For the following days, review authentication logs and outbound connections daily:

sudo journalctl -u ssh --since yesterday | grep -E 'Accepted|Failed'
sudo ss -tunp state established

Phase 6 - Post-incident review

Hold a blameless review within a week, while memories are fresh. Base it on the incident log and the collected evidence, and produce a written report with:

  • Timeline: first malicious activity, detection, containment, eradication and recovery, all in UTC. Build it from the auth log, the journal, web server logs and file timestamps in the evidence archive.
  • Entry point and root cause: how the attacker got in and why that was possible.
  • Impact: which systems and data were accessed or modified, and whether customers, partners or authorities must be notified.
  • Detection: how the incident was noticed, how long it went undetected and which alert would have caught it earlier.
  • Response metrics: time to detect, time to contain and time to recover.
  • Actions: concrete follow-ups with an owner and a due date, for example patching policy, removing password authentication, adding log shipping or new alerts.

Update this playbook with anything that slowed you down during the incident.

Troubleshooting

You locked yourself out after loading the containment table. Log in through the provider's web console or IPMI and run sudo nft delete table inet ir_contain. Then correct the IP address in /root/ir-contain.nft and load it again.

ss -K does nothing. The kernel must be built with socket destroy support (CONFIG_INET_DIAG_DESTROY), which Ubuntu kernels include. Run it as root, and check that the address actually matches with sudo ss -tn dst 203.0.113.66 first.

dpkg --verify reports many configuration files. That is expected; files marked c are configuration files that you or your automation changed. Focus on binaries and libraries without the c marker, and confirm individual files with sudo debsums -s package_name.

The evidence archive is too large to copy. Most of the size is /var/log. Transfer the archive with rsync --partial so interrupted copies resume, or split it with split -b 1G and verify each part with sha256sum on both sides.

Conclusion

This playbook turns a security incident on a Linux server into an ordered sequence: prepare, triage without destroying evidence, contain the network, collect and hash evidence, rebuild instead of cleaning, rotate every exposed secret and learn from the review. As next steps, run a tabletop exercise with your team using this document, install ir-collect on all your servers with your configuration management, and ship logs off-host so the evidence survives even if an attacker deletes it.