Grafana Alerting evaluates queries against your data sources on a schedule and sends notifications when a condition holds for long enough. It has three building blocks: alert rules decide when something is wrong, contact points define where notifications go (email, Slack and many others), and notification policies route each alert to a contact point based on its labels. In this tutorial you will configure all three on Grafana running on Ubuntu 24.04, test a real alert end to end and then store the configuration as provisioning files.
Prerequisites
To follow this tutorial, you will need:
- A server running Ubuntu 24.04 LTS, such as a CubePath VPS, with a non-root user with
sudoprivileges. - Grafana 11 or later installed from the official APT repository (
apt.grafana.com) and reachable in a browser, with an admin account. - A Prometheus data source configured in Grafana that scrapes Node Exporter, for example the Ubuntu
prometheusandprometheus-node-exporterpackages, which scrape the local node under thenodejob. - An SMTP account for sending email (your mail provider or a transactional email service) and, optionally, a Slack incoming webhook URL.
The menu paths below match recent Grafana versions. If a label differs slightly in yours, the same options exist under Alerting in the main menu.
Step 1 - Configuring SMTP in Grafana
Grafana needs an SMTP server to send email notifications. Open the main configuration file:
sudo nano /etc/grafana/grafana.ini
Find the [smtp] section, uncomment the lines below and set your values:
[smtp]
enabled = true
host = smtp.example.com:587
user = your_smtp_user
password = """your_smtp_password"""
from_address = grafana@your_domain
from_name = Grafana
Wrapping the password in triple quotes is required when it contains # or ;, which the INI parser would otherwise treat as the start of a comment. Port 587 uses STARTTLS, which Grafana negotiates automatically.
Restart Grafana and check that it started cleanly:
sudo systemctl restart grafana-server
sudo journalctl -u grafana-server -n 20 --no-pager
The log should not contain errors mentioning smtp. You will test real delivery in the next step.
Step 2 - Creating contact points
A contact point is a named list of integrations. Create one for email and one for Slack so that you can route warnings and critical alerts differently.
- In Grafana, go to Alerting > Contact points and click + Add contact point.
- Set Name to
ops-email, choose the Email integration and enter one or more addresses separated by;in Addresses. - Click Test, then Send test notification. Check the inbox: the message arrives from the
from_addressyou configured. - Click Save contact point.
For Slack, first create an incoming webhook in your Slack workspace (a Slack app with Incoming Webhooks enabled, attached to a channel such as #alerts). Then:
- Click + Add contact point again and name it
ops-slack. - Choose the Slack integration and paste the webhook address into Webhook URL.
- Click Test and confirm that the message appears in the channel, then save.
If a test fails, Grafana shows the error returned by the SMTP server or Slack, for example an authentication failure or an invalid webhook.
Step 3 - Setting up notification policies
Notification policies form a tree. Every alert enters at the default policy and moves to the first nested policy whose label matchers it satisfies. Here, critical alerts go to Slack and everything else to email.
Go to Alerting > Notification policies.
- Edit the Default policy: set the default contact point to
ops-email. Under the timing options, keep Group by asgrafana_folder, alertname, and set Repeat interval to4hso an unresolved alert is not re-sent every hour. - Click + New child policy (the name may be New specific policy in some versions). Add the matcher
severity=criticaland choose the contact pointops-slack. Save.
The timing settings control noise:
| Setting | Meaning | Typical value |
|---|---|---|
| Group wait | How long to wait to batch the first alerts of a new group | 30s |
| Group interval | Minimum time between notifications about new alerts in the same group | 5m |
| Repeat interval | How often to re-send a notification that is still firing | 4h |
Alerts that match the child policy stop there. Alerts without severity=critical stay at the default policy and are emailed.
Step 4 - Creating an alert rule
Create a rule that fires when a Node Exporter target stops responding. Go to Alerting > Alert rules and click + New alert rule.
-
Name:
Instance down. -
Query: select the Prometheus data source and enter the query below in code mode:
up{job="node"} -
Expressions: Grafana adds a Reduce and a Threshold expression by default. For an instant Prometheus query you can remove the Reduce step and set the Threshold to use the query as input with the condition IS BELOW
1. Set the threshold expression as the alert condition. Click Preview to see one row per instance with its current state (Normalif the node is up). -
Folder and labels: create or choose a folder such as
Infrastructureand add the labelseverity=critical. This label is what the notification policy from step 3 matches. -
Evaluation: create an evaluation group named
infrastructurewith an interval of1m, and set the Pending period to2m. The condition must be true for two consecutive minutes before the alert fires, which filters out a single failed scrape. -
Notifications: keep the option to use the notification policy (instead of choosing a contact point directly), so routing stays in one place.
-
Annotations: set Summary to:
{{ $labels.instance }} is downand Description to:
Prometheus has not been able to scrape {{ $labels.instance }} (job {{ $labels.job }}) for 2 minutes. -
In Configure no data and error handling, set the state for No data to
Alertingif a missing series should be treated as a failure, or keep the default otherwise.
Click Save rule and exit. The rule appears in the list with the state Normal.
A second useful rule follows the same steps with a different query. For high CPU usage, use this query with a threshold IS ABOVE 90, a pending period of 10m and the label severity = warning:
100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])))
Because it has severity=warning, it will be routed to email by the default policy.
Step 5 - Testing the alert end to end
Stop Node Exporter on a monitored server to make the Instance down rule fire:
sudo systemctl stop prometheus-node-exporter
Watch the rule in Alerting > Alert rules. Within about a minute the state changes to Pending, and after the two-minute pending period it becomes Firing. The notification policy routes it to ops-slack because of the severity=critical label, and a message appears in the Slack channel after the group wait.
Start the exporter again:
sudo systemctl start prometheus-node-exporter
On the next evaluations the alert returns to Normal and Grafana sends a Resolved notification to the same channel.
Step 6 - Silencing alerts during maintenance
A silence stops notifications for matching alerts for a fixed period without changing any rule, which is what you want during planned maintenance.
- Go to Alerting > Silences and click + Add silence (or New silence).
- Set the start and end time, for example the next two hours.
- Add a matcher, for example
instance=web01:9100. The preview lists the currently firing alerts that the silence will affect. - Add a comment explaining why, then save.
The alert rules keep evaluating and their state is still visible, but no notifications are sent until the silence expires or you remove it. For recurring windows, such as a nightly backup that always spikes CPU, use mute timings in the notification policies instead.
Step 7 - Managing alerting as code with provisioning
Configuration made in the UI lives in the Grafana database. To version it and recreate it on another instance, export it to files that Grafana loads at startup from /etc/grafana/provisioning/alerting/.
Export what you created: in Alerting > Alert rules, use the Export option (on the rule, on its folder or on the whole list) and choose YAML. Contact points and notification policies have their own Export buttons on their pages. The exported YAML is already in provisioning format.
As an example, this file defines the two contact points and the routing tree from steps 2 and 3:
sudo nano /etc/grafana/provisioning/alerting/notifications.yaml
apiVersion: 1
contactPoints:
- orgId: 1
name: ops-email
receivers:
- uid: ops-email
type: email
settings:
addresses: ops@your_domain
- orgId: 1
name: ops-slack
receivers:
- uid: ops-slack
type: slack
settings:
url: $SLACK_WEBHOOK_URL
policies:
- orgId: 1
receiver: ops-email
group_by: ['grafana_folder', 'alertname']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- receiver: ops-slack
object_matchers:
- ['severity', '=', 'critical']
Grafana expands environment variables in provisioning files, so the webhook stays out of the file. Set the variable for the service with a systemd override:
sudo systemctl edit grafana-server
[Service]
Environment="SLACK_WEBHOOK_URL=https://hooks.slack.com/services/your/webhook/path"
Save the exported rules YAML in the same directory, for example as /etc/grafana/provisioning/alerting/rules.yaml. Then remove the objects you created by hand in the UI (provisioning cannot take over an existing contact point with a different UID), and restart Grafana:
sudo chown -R root:grafana /etc/grafana/provisioning/alerting
sudo chmod 0640 /etc/grafana/provisioning/alerting/*.yaml
sudo systemctl restart grafana-server
sudo journalctl -u grafana-server -n 50 --no-pager | grep -i provisioning
Provisioned resources are marked as Provisioned in the UI and cannot be edited there, so all changes go through the files and your version control. Note that the notification policy tree is a single object: provisioning it replaces the whole tree defined in the UI.
Troubleshooting
- Email test fails with
SMTP not configured:enabled = trueis missing or still commented out in[smtp], or Grafana was not restarted after the change. - Alert fires but nobody is notified: open Alerting > Notification policies and check the matchers. The label on the rule must match exactly, including case (
severity=critical). Also check Alerting > Silences for an active silence. - Rule shows
ErrororNoData: open the rule and click Preview. A query error, a wrong data source or a query that returns a range instead of a single value per series are the usual causes. - Provisioning file ignored: check
sudo journalctl -u grafana-serverfor YAML errors, and make sure the file ends in.yamlor.ymland is readable by thegrafanagroup.
Conclusion
Grafana now evaluates alert rules against Prometheus, routes critical alerts to Slack and everything else to email, supports silences for maintenance and keeps its alerting configuration in files that you can version. The same rules, contact points and policies extend to any data source that Grafana supports for alerting, such as Loki or a SQL database.
As next steps, add notification templates to control the message format, create mute timings for known noisy windows, and link each rule to a dashboard panel so that notifications include a direct link to the relevant graph.
