Configuration drift is the gap between how a server is supposed to be configured and how it actually is. It creeps in through hotfixes made over SSH, packages installed by hand, or resources changed in a web console, and it is the reason a rebuild from code sometimes does not behave like the server it replaces. In this tutorial you will detect drift at three layers on Ubuntu 24.04: infrastructure managed by Terraform, server configuration managed by Ansible, and any file on disk with AIDE. You will also schedule the Terraform check with a systemd timer so drift is reported without anyone having to remember to look.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS, such as a CubePath VPS, with a non-root user that has sudo privileges. This is where the checks run.
  • For Step 1 and Step 2: an existing Terraform project with its state available (local or remote backend) and Terraform 1.5 or newer.
  • For Step 3: one or more servers you manage with Ansible and SSH access to them from this machine.
  • AIDE (Step 4 and Step 5) works on its own, with no Terraform or Ansible needed.

You can follow only the sections that match the tools you use.

How drift detection works

Each tool compares a desired state with the real one, and reports the difference without changing anything:

LayerDesired stateCommandDrift signal
InfrastructureTerraform code and stateterraform plan -detailed-exitcodeExit code 2
Server configurationAnsible playbooksansible-playbook --check --diffTasks reported as changed
Files on diskAIDE database snapshotaide --checkNon-zero exit code with a list of files

Terraform and Ansible can only see what they manage. AIDE covers the rest: a binary replaced by hand, a new cron job, an edited file under /etc that no playbook knows about.

Step 1 - Detecting infrastructure drift with Terraform

terraform plan refreshes the state from the real infrastructure and compares it with your code. The -detailed-exitcode flag turns the result into something a script can act on:

  • 0: no changes, the infrastructure matches the code.
  • 1: the plan failed (credentials, syntax, network).
  • 2: there are changes, meaning either the infrastructure drifted or the code has changes that were never applied.

Run it from your Terraform project directory. The -lock=false flag avoids taking the state lock, which is safe for a read-only plan and keeps a scheduled check from blocking a real deployment:

cd ~/your_terraform_project
terraform plan -detailed-exitcode -input=false -lock=false
echo "exit code: $?"

When nothing has drifted, the output ends like this:

No changes. Your infrastructure matches the configuration.
...
exit code: 0

When something was changed or deleted outside of Terraform, the plan shows what it would do to restore the desired state:

  # docker_container.web will be created
  + resource "docker_container" "web" {
...
Plan: 1 to add, 0 to change, 0 to destroy.
exit code: 2

Step 2 - Scheduling the Terraform check with systemd

A drift check is only useful if it runs regularly. Create a small script that runs the plan and turns the exit code into a clear message:

sudo nano /usr/local/bin/tf-drift-check
#!/usr/bin/env bash
set -euo pipefail

tf_dir="${1:?Usage: tf-drift-check <terraform-directory>}"
cd "$tf_dir"

terraform init -input=false -no-color > /dev/null

rc=0
plan_output=$(terraform plan -detailed-exitcode -input=false -lock=false -no-color 2>&1) || rc=$?

case "$rc" in
  0)
    echo "No drift in $tf_dir"
    ;;
  2)
    echo "Drift detected in $tf_dir"
    echo "$plan_output"
    exit 2
    ;;
  *)
    echo "terraform plan failed in $tf_dir"
    echo "$plan_output"
    exit 1
    ;;
esac

Make it executable and test it by hand:

sudo chmod 755 /usr/local/bin/tf-drift-check
tf-drift-check ~/your_terraform_project
No drift in /home/your_user/your_terraform_project

Next, create a service unit that runs the script as your user. Any variables or credentials the plan needs (for example TF_VAR_db_password or cloud API tokens) go in an environment file readable only by that user:

mkdir -p ~/.config
install -m 600 /dev/null ~/.config/tf-drift.env
nano ~/.config/tf-drift.env
TF_VAR_db_password=your_strong_password

Create /etc/systemd/system/tf-drift-check.service, replacing your_user and the project path:

sudo nano /etc/systemd/system/tf-drift-check.service
[Unit]
Description=Terraform drift check
Wants=network-online.target
After=network-online.target

[Service]
Type=oneshot
User=your_user
EnvironmentFile=/home/your_user/.config/tf-drift.env
ExecStart=/usr/local/bin/tf-drift-check /home/your_user/your_terraform_project

Then a timer that runs it every six hours:

sudo nano /etc/systemd/system/tf-drift-check.timer
[Unit]
Description=Run the Terraform drift check every 6 hours

[Timer]
OnCalendar=00/6:00
RandomizedDelaySec=10m
Persistent=true

[Install]
WantedBy=timers.target

Load the units, start the timer, and trigger one run immediately to test it:

sudo systemctl daemon-reload
sudo systemctl enable --now tf-drift-check.timer
sudo systemctl start tf-drift-check.service

Check the result in the journal:

journalctl -u tf-drift-check.service -n 20 --no-pager
Sep 25 12:00:03 web1 systemd[1]: Starting tf-drift-check.service - Terraform drift check...
Sep 25 12:00:09 web1 tf-drift-check[4121]: No drift in /home/your_user/your_terraform_project
Sep 25 12:00:09 web1 systemd[1]: Finished tf-drift-check.service - Terraform drift check.

When drift is found, the script exits with code 2 and the unit ends up in the failed state, so it shows up in systemctl --failed and in any monitoring that watches failed units. Confirm the next scheduled run with:

systemctl list-timers tf-drift-check.timer

Step 3 - Detecting configuration drift with Ansible

Ansible playbooks are idempotent: running one against a server that already matches reports every task as ok. In check mode, Ansible reports what it would change without changing it, and --diff shows the exact lines. Install Ansible from the Ubuntu archive if you do not have it:

sudo apt update
sudo apt install -y ansible

As an example, this playbook enforces an SSH hardening drop-in and makes sure fail2ban is installed. Create it next to your inventory:

nano ~/ansible/site.yml
- name: Baseline configuration
  hosts: all
  become: true
  tasks:
    - name: Install baseline packages
      ansible.builtin.apt:
        name:
          - fail2ban
          - unattended-upgrades
        state: present

    - name: Enforce SSH hardening
      ansible.builtin.copy:
        dest: /etc/ssh/sshd_config.d/10-hardening.conf
        owner: root
        group: root
        mode: "0644"
        content: |
          PasswordAuthentication no
          PermitRootLogin no
      notify: Reload ssh

  handlers:
    - name: Reload ssh
      ansible.builtin.service:
        name: ssh
        state: reloaded

Apply it once so the servers match the playbook:

cd ~/ansible
ansible-playbook -i inventory.ini site.yml

Now simulate drift, as someone fixing an issue by hand would. Log in to one of the managed servers (web1 in this example) and overwrite the file:

echo 'PasswordAuthentication yes' | sudo tee /etc/ssh/sshd_config.d/10-hardening.conf

Back on the control machine, run the playbook in check mode:

ansible-playbook -i inventory.ini site.yml --check --diff
TASK [Enforce SSH hardening] ***************************************************
--- before: /etc/ssh/sshd_config.d/10-hardening.conf
+++ after: /etc/ssh/sshd_config.d/10-hardening.conf
@@ -1 +1,2 @@
-PasswordAuthentication yes
+PasswordAuthentication no
+PermitRootLogin no

changed: [web1]

RUNNING HANDLER [Reload ssh] ***************************************************
changed: [web1]

PLAY RECAP *********************************************************************
web1   : ok=4    changed=2    unreachable=0    failed=0    skipped=0

Any changed count above zero in the recap is drift. Nothing on the server has been modified yet. To limit the check to one host, add --limit web1.

For a machine-readable summary, for example from a cron job, use the JSON callback shipped in the ansible.posix collection (included in Ubuntu's ansible package) and count the changed tasks with jq:

sudo apt install -y jq
ANSIBLE_STDOUT_CALLBACK=ansible.posix.json ansible-playbook -i inventory.ini site.yml --check | jq '[.stats[].changed] | add'
2

Step 4 - Creating a file integrity baseline with AIDE

AIDE (Advanced Intrusion Detection Environment) records the checksums, permissions and ownership of files in a database, then reports anything that differs on later checks. It catches changes that no configuration tool manages. Install it on each server you want to monitor:

sudo apt install -y aide

The Ubuntu package ships a configuration that covers system binaries, libraries and /etc. Add rules for your own application directories in a separate file under /etc/aide/aide.conf.d/. File names in that directory may contain only letters, digits, _ and -, so do not use a .conf extension:

sudo nano /etc/aide/aide.conf.d/99_local_app
# Application code and configuration: report content, permission and owner changes
/srv/myapp R
# Uploads and cache change constantly: ignore them
!/srv/myapp/uploads
!/srv/myapp/cache

R is AIDE's built-in rule group for read-only files: it checks permissions, ownership, size, timestamps and checksums. A line starting with ! excludes a path. Replace /srv/myapp with your application's directory.

Build the initial database. This reads every monitored file and can take several minutes:

sudo aideinit
Running aide --init...
...
AIDE initialized database at /var/lib/aide/aide.db.new

aideinit writes the new database to /var/lib/aide/aide.db.new and copies it to /var/lib/aide/aide.db, which is the baseline used by checks. Build the baseline on a server you trust, right after provisioning or after a verified deployment.

Step 5 - Checking files against the baseline

Run a check against the baseline:

sudo aide --config /etc/aide/aide.conf --check

On an unchanged system, AIDE reports that everything matches:

AIDE found NO differences between database and filesystem. Looks okay!!

Make a change by hand to see what drift looks like, for example editing a file in /etc:

echo "# test change" | sudo tee -a /etc/hosts
sudo aide --config /etc/aide/aide.conf --check
AIDE found differences between database and filesystem!!
...
Summary:
  Total number of entries:      118223
  Added entries:                0
  Removed entries:              0
  Changed entries:              1
...

Below the summary, AIDE lists each changed file (here /etc/hosts) with the old and new values of every attribute that differs, such as size, modification time and checksums. The exit code tells you what kind of change was found: it is a bit mask where 1 means added files, 2 removed files and 4 changed files, so 5 means files were both added and changed. Revert the test line with sudo nano /etc/hosts.

When a change is legitimate, such as a package upgrade or a deployment, update the baseline so it stops being reported:

sudo aide --config /etc/aide/aide.conf --update
sudo cp /var/lib/aide/aide.db.new /var/lib/aide/aide.db

Review the report printed by --update before copying the database: whatever you accept here becomes the new "known good" state.

The aide-common package also installs a daily check as the dailyaidecheck timer, which mails the report to root. Confirm it is active:

systemctl list-timers 'dailyaidecheck*'

For the report to reach you, configure a mail relay on the server and forward root's mail to a real address in /etc/aliases.

Step 6 - Remediating drift

Detection tells you what differs. Deciding what to do is a human step, and there are two valid answers:

  • The change was wrong: reapply the desired state. For Terraform, run terraform apply after reviewing the plan. For Ansible, run the playbook without --check. For a file only AIDE caught, restore it from your configuration management or from the package (sudo apt install --reinstall package_name), and find out who changed it and why.
  • The change was right: make the code match reality. Update the Terraform code (or use an import block for a resource created by hand) until the plan is clean, add the change to the Ansible playbook, and update the AIDE baseline.

Avoid automatically reverting drift from the scheduled job. An emergency fix applied by hand at 3 AM is also drift, and silently undoing it can bring back the outage it fixed.

Troubleshooting

  • The Terraform timer reports failures but the plan works in your shell: the service does not load your shell environment. Put every variable and credential the plan needs in the EnvironmentFile, and check journalctl -u tf-drift-check.service for the real error.
  • Error acquiring the state lock: an old version of the script ran without -lock=false while someone was applying. Make sure the plan uses -lock=false; never force-unlock while an apply may be running.
  • Ansible check mode fails on a task that depends on an earlier one: in check mode, a package that would be installed is not actually installed, so a later task that uses it fails. Add when: not ansible_check_mode to that task, or accept the failure as a sign of drift.
  • AIDE reports hundreds of changes after every upgrade: this is expected after apt upgrade. Update the baseline after each maintenance window, and exclude paths that change constantly (logs, caches, spool directories) with ! rules.

Conclusion

You now detect drift at three layers: Terraform plan exit codes show when infrastructure no longer matches the code, Ansible check mode shows which configuration a server has lost, and AIDE reports any file changed outside of those tools. The Terraform check runs on a systemd timer and AIDE runs daily, so drift is reported continuously instead of discovered during an incident.

As next steps, forward failed systemd units and AIDE reports to your alerting channel, run the Terraform drift check in CI on a schedule for each environment, and treat every drift report as a prompt to move the manual change into code.