Configuration drift is the gap between how a server is supposed to be configured and how it actually is. It creeps in through hotfixes made over SSH, packages installed by hand, or resources changed in a web console, and it is the reason a rebuild from code sometimes does not behave like the server it replaces. In this tutorial you will detect drift at three layers on Ubuntu 24.04: infrastructure managed by Terraform, server configuration managed by Ansible, and any file on disk with AIDE. You will also schedule the Terraform check with a systemd timer so drift is reported without anyone having to remember to look.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS, such as a CubePath VPS, with a non-root user that has
sudoprivileges. This is where the checks run. - For Step 1 and Step 2: an existing Terraform project with its state available (local or remote backend) and Terraform 1.5 or newer.
- For Step 3: one or more servers you manage with Ansible and SSH access to them from this machine.
- AIDE (Step 4 and Step 5) works on its own, with no Terraform or Ansible needed.
You can follow only the sections that match the tools you use.
How drift detection works
Each tool compares a desired state with the real one, and reports the difference without changing anything:
| Layer | Desired state | Command | Drift signal |
|---|---|---|---|
| Infrastructure | Terraform code and state | terraform plan -detailed-exitcode | Exit code 2 |
| Server configuration | Ansible playbooks | ansible-playbook --check --diff | Tasks reported as changed |
| Files on disk | AIDE database snapshot | aide --check | Non-zero exit code with a list of files |
Terraform and Ansible can only see what they manage. AIDE covers the rest: a binary replaced by hand, a new cron job, an edited file under /etc that no playbook knows about.
Step 1 - Detecting infrastructure drift with Terraform
terraform plan refreshes the state from the real infrastructure and compares it with your code. The -detailed-exitcode flag turns the result into something a script can act on:
0: no changes, the infrastructure matches the code.1: the plan failed (credentials, syntax, network).2: there are changes, meaning either the infrastructure drifted or the code has changes that were never applied.
Run it from your Terraform project directory. The -lock=false flag avoids taking the state lock, which is safe for a read-only plan and keeps a scheduled check from blocking a real deployment:
cd ~/your_terraform_project
terraform plan -detailed-exitcode -input=false -lock=false
echo "exit code: $?"
When nothing has drifted, the output ends like this:
No changes. Your infrastructure matches the configuration.
...
exit code: 0
When something was changed or deleted outside of Terraform, the plan shows what it would do to restore the desired state:
# docker_container.web will be created
+ resource "docker_container" "web" {
...
Plan: 1 to add, 0 to change, 0 to destroy.
exit code: 2
Note
terraform plan -refresh-onlyshows only the changes made outside Terraform, without the code changes. It is useful when investigating, but many providers report cosmetic differences (empty lists, default values) in refresh-only mode, so a regular plan gives fewer false alarms for automated checks.
Step 2 - Scheduling the Terraform check with systemd
A drift check is only useful if it runs regularly. Create a small script that runs the plan and turns the exit code into a clear message:
sudo nano /usr/local/bin/tf-drift-check
#!/usr/bin/env bash
set -euo pipefail
tf_dir="${1:?Usage: tf-drift-check <terraform-directory>}"
cd "$tf_dir"
terraform init -input=false -no-color > /dev/null
rc=0
plan_output=$(terraform plan -detailed-exitcode -input=false -lock=false -no-color 2>&1) || rc=$?
case "$rc" in
0)
echo "No drift in $tf_dir"
;;
2)
echo "Drift detected in $tf_dir"
echo "$plan_output"
exit 2
;;
*)
echo "terraform plan failed in $tf_dir"
echo "$plan_output"
exit 1
;;
esac
Make it executable and test it by hand:
sudo chmod 755 /usr/local/bin/tf-drift-check
tf-drift-check ~/your_terraform_project
No drift in /home/your_user/your_terraform_project
Next, create a service unit that runs the script as your user. Any variables or credentials the plan needs (for example TF_VAR_db_password or cloud API tokens) go in an environment file readable only by that user:
mkdir -p ~/.config
install -m 600 /dev/null ~/.config/tf-drift.env
nano ~/.config/tf-drift.env
TF_VAR_db_password=your_strong_password
Create /etc/systemd/system/tf-drift-check.service, replacing your_user and the project path:
sudo nano /etc/systemd/system/tf-drift-check.service
[Unit]
Description=Terraform drift check
Wants=network-online.target
After=network-online.target
[Service]
Type=oneshot
User=your_user
EnvironmentFile=/home/your_user/.config/tf-drift.env
ExecStart=/usr/local/bin/tf-drift-check /home/your_user/your_terraform_project
Then a timer that runs it every six hours:
sudo nano /etc/systemd/system/tf-drift-check.timer
[Unit]
Description=Run the Terraform drift check every 6 hours
[Timer]
OnCalendar=00/6:00
RandomizedDelaySec=10m
Persistent=true
[Install]
WantedBy=timers.target
Load the units, start the timer, and trigger one run immediately to test it:
sudo systemctl daemon-reload
sudo systemctl enable --now tf-drift-check.timer
sudo systemctl start tf-drift-check.service
Check the result in the journal:
journalctl -u tf-drift-check.service -n 20 --no-pager
Sep 25 12:00:03 web1 systemd[1]: Starting tf-drift-check.service - Terraform drift check...
Sep 25 12:00:09 web1 tf-drift-check[4121]: No drift in /home/your_user/your_terraform_project
Sep 25 12:00:09 web1 systemd[1]: Finished tf-drift-check.service - Terraform drift check.
When drift is found, the script exits with code 2 and the unit ends up in the failed state, so it shows up in systemctl --failed and in any monitoring that watches failed units. Confirm the next scheduled run with:
systemctl list-timers tf-drift-check.timer
Step 3 - Detecting configuration drift with Ansible
Ansible playbooks are idempotent: running one against a server that already matches reports every task as ok. In check mode, Ansible reports what it would change without changing it, and --diff shows the exact lines. Install Ansible from the Ubuntu archive if you do not have it:
sudo apt update
sudo apt install -y ansible
As an example, this playbook enforces an SSH hardening drop-in and makes sure fail2ban is installed. Create it next to your inventory:
nano ~/ansible/site.yml
- name: Baseline configuration
hosts: all
become: true
tasks:
- name: Install baseline packages
ansible.builtin.apt:
name:
- fail2ban
- unattended-upgrades
state: present
- name: Enforce SSH hardening
ansible.builtin.copy:
dest: /etc/ssh/sshd_config.d/10-hardening.conf
owner: root
group: root
mode: "0644"
content: |
PasswordAuthentication no
PermitRootLogin no
notify: Reload ssh
handlers:
- name: Reload ssh
ansible.builtin.service:
name: ssh
state: reloaded
Apply it once so the servers match the playbook:
cd ~/ansible
ansible-playbook -i inventory.ini site.yml
Now simulate drift, as someone fixing an issue by hand would. Log in to one of the managed servers (web1 in this example) and overwrite the file:
echo 'PasswordAuthentication yes' | sudo tee /etc/ssh/sshd_config.d/10-hardening.conf
Back on the control machine, run the playbook in check mode:
ansible-playbook -i inventory.ini site.yml --check --diff
TASK [Enforce SSH hardening] ***************************************************
--- before: /etc/ssh/sshd_config.d/10-hardening.conf
+++ after: /etc/ssh/sshd_config.d/10-hardening.conf
@@ -1 +1,2 @@
-PasswordAuthentication yes
+PasswordAuthentication no
+PermitRootLogin no
changed: [web1]
RUNNING HANDLER [Reload ssh] ***************************************************
changed: [web1]
PLAY RECAP *********************************************************************
web1 : ok=4 changed=2 unreachable=0 failed=0 skipped=0
Any changed count above zero in the recap is drift. Nothing on the server has been modified yet. To limit the check to one host, add --limit web1.
Notecheck mode is only as accurate as your tasks.
commandandshelltasks are skipped in check mode unless you setcheck_mode: falseon them, so configuration applied that way is invisible to this check. Prefer purpose-built modules such ascopy,template,lineinfileandapt.
For a machine-readable summary, for example from a cron job, use the JSON callback shipped in the ansible.posix collection (included in Ubuntu's ansible package) and count the changed tasks with jq:
sudo apt install -y jq
ANSIBLE_STDOUT_CALLBACK=ansible.posix.json ansible-playbook -i inventory.ini site.yml --check | jq '[.stats[].changed] | add'
2
Step 4 - Creating a file integrity baseline with AIDE
AIDE (Advanced Intrusion Detection Environment) records the checksums, permissions and ownership of files in a database, then reports anything that differs on later checks. It catches changes that no configuration tool manages. Install it on each server you want to monitor:
sudo apt install -y aide
The Ubuntu package ships a configuration that covers system binaries, libraries and /etc. Add rules for your own application directories in a separate file under /etc/aide/aide.conf.d/. File names in that directory may contain only letters, digits, _ and -, so do not use a .conf extension:
sudo nano /etc/aide/aide.conf.d/99_local_app
# Application code and configuration: report content, permission and owner changes
/srv/myapp R
# Uploads and cache change constantly: ignore them
!/srv/myapp/uploads
!/srv/myapp/cache
R is AIDE's built-in rule group for read-only files: it checks permissions, ownership, size, timestamps and checksums. A line starting with ! excludes a path. Replace /srv/myapp with your application's directory.
Build the initial database. This reads every monitored file and can take several minutes:
sudo aideinit
Running aide --init...
...
AIDE initialized database at /var/lib/aide/aide.db.new
aideinit writes the new database to /var/lib/aide/aide.db.new and copies it to /var/lib/aide/aide.db, which is the baseline used by checks. Build the baseline on a server you trust, right after provisioning or after a verified deployment.
Step 5 - Checking files against the baseline
Run a check against the baseline:
sudo aide --config /etc/aide/aide.conf --check
On an unchanged system, AIDE reports that everything matches:
AIDE found NO differences between database and filesystem. Looks okay!!
Make a change by hand to see what drift looks like, for example editing a file in /etc:
echo "# test change" | sudo tee -a /etc/hosts
sudo aide --config /etc/aide/aide.conf --check
AIDE found differences between database and filesystem!!
...
Summary:
Total number of entries: 118223
Added entries: 0
Removed entries: 0
Changed entries: 1
...
Below the summary, AIDE lists each changed file (here /etc/hosts) with the old and new values of every attribute that differs, such as size, modification time and checksums. The exit code tells you what kind of change was found: it is a bit mask where 1 means added files, 2 removed files and 4 changed files, so 5 means files were both added and changed. Revert the test line with sudo nano /etc/hosts.
When a change is legitimate, such as a package upgrade or a deployment, update the baseline so it stops being reported:
sudo aide --config /etc/aide/aide.conf --update
sudo cp /var/lib/aide/aide.db.new /var/lib/aide/aide.db
Review the report printed by --update before copying the database: whatever you accept here becomes the new "known good" state.
The aide-common package also installs a daily check as the dailyaidecheck timer, which mails the report to root. Confirm it is active:
systemctl list-timers 'dailyaidecheck*'
For the report to reach you, configure a mail relay on the server and forward root's mail to a real address in /etc/aliases.
Step 6 - Remediating drift
Detection tells you what differs. Deciding what to do is a human step, and there are two valid answers:
- The change was wrong: reapply the desired state. For Terraform, run
terraform applyafter reviewing the plan. For Ansible, run the playbook without--check. For a file only AIDE caught, restore it from your configuration management or from the package (sudo apt install --reinstall package_name), and find out who changed it and why. - The change was right: make the code match reality. Update the Terraform code (or use an
importblock for a resource created by hand) until the plan is clean, add the change to the Ansible playbook, and update the AIDE baseline.
Avoid automatically reverting drift from the scheduled job. An emergency fix applied by hand at 3 AM is also drift, and silently undoing it can bring back the outage it fixed.
Troubleshooting
- The Terraform timer reports failures but the plan works in your shell: the service does not load your shell environment. Put every variable and credential the plan needs in the
EnvironmentFile, and checkjournalctl -u tf-drift-check.servicefor the real error. Error acquiring the state lock: an old version of the script ran without-lock=falsewhile someone was applying. Make sure the plan uses-lock=false; never force-unlock while an apply may be running.- Ansible check mode fails on a task that depends on an earlier one: in check mode, a package that would be installed is not actually installed, so a later task that uses it fails. Add
when: not ansible_check_modeto that task, or accept the failure as a sign of drift. - AIDE reports hundreds of changes after every upgrade: this is expected after
apt upgrade. Update the baseline after each maintenance window, and exclude paths that change constantly (logs, caches, spool directories) with!rules.
Conclusion
You now detect drift at three layers: Terraform plan exit codes show when infrastructure no longer matches the code, Ansible check mode shows which configuration a server has lost, and AIDE reports any file changed outside of those tools. The Terraform check runs on a systemd timer and AIDE runs daily, so drift is reported continuously instead of discovered during an incident.
As next steps, forward failed systemd units and AIDE reports to your alerting channel, run the Terraform drift check in CI on a schedule for each environment, and treat every drift report as a prompt to move the manual change into code.
