kube-prometheus-stack is a Helm chart that deploys the Prometheus Operator together with Prometheus, Alertmanager, Grafana, node-exporter and kube-state-metrics, preconfigured with dashboards and alerting rules for Kubernetes itself. In this tutorial you will install the chart with a small values file, reach each web UI, scrape a sample application with a ServiceMonitor, and add your own alert with a PrometheusRule. The commands run from any workstation with kubectl access to the cluster; the examples assume an Ubuntu 24.04 workstation.
Prerequisites
To follow this guide you need:
- A working Kubernetes cluster (1.30 or newer) with at least 4 GB of free memory across the worker nodes, for example on CubePath VPS or bare metal servers.
kubectlconfigured against that cluster with cluster-admin rights. Check withkubectl get nodes.- Helm 3 installed on your workstation. Check with
helm version. jqfor reading the Prometheus API output:sudo apt install jq.- A default StorageClass so Prometheus, Alertmanager and Grafana can keep their data. Check with
kubectl get storageclassand look for(default)next to one of them.
Step 1 - Adding the Helm repository
The chart is published by the Prometheus community. Add the repository and refresh the index:
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
Confirm that the chart is available:
helm search repo prometheus-community/kube-prometheus-stack
NAME CHART VERSION APP VERSION DESCRIPTION
prometheus-community/kube-prometheus-stack ... v0.8x.x kube-prometheus-stack collects Kubernetes manif...
Step 2 - Writing a values file
The chart works out of the box, but three things are worth setting from the start: persistent storage (otherwise metrics are lost when the Prometheus pod restarts), a Grafana admin password you choose, and the selector settings.
By default Prometheus only picks up ServiceMonitor, PodMonitor and PrometheusRule objects that carry the label release: <your release name>. That is the most common reason custom monitors are silently ignored. Setting the *SelectorNilUsesHelmValues options to false makes Prometheus select all of them in every namespace.
Create the file:
nano kube-prometheus-values.yaml
Add the following content, replacing your_strong_password with a real password:
grafana:
adminPassword: your_strong_password
persistence:
enabled: true
size: 5Gi
prometheus:
prometheusSpec:
retention: 15d
retentionSize: 18GB
serviceMonitorSelectorNilUsesHelmValues: false
podMonitorSelectorNilUsesHelmValues: false
ruleSelectorNilUsesHelmValues: false
resources:
requests:
cpu: 200m
memory: 1Gi
limits:
memory: 3Gi
storageSpec:
volumeClaimTemplate:
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 20Gi
alertmanager:
alertmanagerSpec:
storage:
volumeClaimTemplate:
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 2Gi
retentionSize is kept a little below the volume size so Prometheus deletes old blocks before the disk fills up. No storageClassName is set, so the cluster's default StorageClass is used; add one under each spec if you want a specific class.
Step 3 - Installing the chart
Install the chart into a dedicated monitoring namespace. The release name kube-prometheus-stack determines the names of every service used later in this guide:
helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
--namespace monitoring --create-namespace \
-f kube-prometheus-values.yaml
helm upgrade --install installs the release the first time and upgrades it on later runs, so you can use the same command whenever you change the values file.
Wait until all pods are ready:
kubectl -n monitoring wait --for=condition=Ready pods --all --timeout=300s
kubectl -n monitoring get pods
NAME READY STATUS RESTARTS AGE
alertmanager-kube-prometheus-stack-alertmanager-0 2/2 Running 0 2m
kube-prometheus-stack-grafana-6c9b8d7f8d-xk2lq 3/3 Running 0 2m
kube-prometheus-stack-kube-state-metrics-5d9c6b8c7f-9hq4w 1/1 Running 0 2m
kube-prometheus-stack-operator-7f6c9d5b8-pl7m2 1/1 Running 0 2m
kube-prometheus-stack-prometheus-node-exporter-4zk8r 1/1 Running 0 2m
kube-prometheus-stack-prometheus-node-exporter-v7c2n 1/1 Running 0 2m
prometheus-kube-prometheus-stack-prometheus-0 2/2 Running 0 2m
There is one node-exporter pod per node. Check that the persistent volumes were bound:
kubectl -n monitoring get pvc
All claims should show STATUS Bound. If they stay Pending, see the troubleshooting section.
Step 4 - Accessing Prometheus, Grafana and Alertmanager
The chart creates ClusterIP services, which are not reachable from outside the cluster. For administration, kubectl port-forward is the simplest and safest option because nothing is exposed publicly.
Open a tunnel to Prometheus:
kubectl -n monitoring port-forward svc/kube-prometheus-stack-prometheus 9090:9090
Browse to http://localhost:9090, open Status > Targets (called Target health in Prometheus 3) and check that the targets are UP. In a second terminal you can also query the API directly:
curl -s 'http://localhost:9090/api/v1/query?query=up' | jq '.data.result | length'
The output is the number of scraped targets, which should be greater than zero.
Open a tunnel to Grafana, whose service listens on port 80:
kubectl -n monitoring port-forward svc/kube-prometheus-stack-grafana 3000:80
Browse to http://localhost:3000 and log in as admin with the password from your values file. If you forget it, read it from the secret the chart created:
kubectl -n monitoring get secret kube-prometheus-stack-grafana -o jsonpath='{.data.admin-password}' | base64 -d; echo
Under Dashboards you will find the bundled dashboards, for example Kubernetes / Compute Resources / Cluster and Node Exporter / Nodes. They already show data because Prometheus is configured as the default data source.
Alertmanager works the same way:
kubectl -n monitoring port-forward svc/kube-prometheus-stack-alertmanager 9093:9093
At http://localhost:9093 you should see at least the Watchdog alert. It fires permanently on purpose, to prove that the alerting pipeline works end to end.
NoteIf you want to publish Grafana on a domain instead, put it behind your Ingress controller with TLS and keep Prometheus and Alertmanager private, since they have no authentication.
Step 5 - Scraping your own application with a ServiceMonitor
A ServiceMonitor tells the operator which Services to scrape and on which port. To test it, deploy the small example application used by the Prometheus Operator project, which exposes metrics on port 8080.
Create a manifest:
nano example-app.yaml
apiVersion: v1
kind: Namespace
metadata:
name: demo
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: example-app
namespace: demo
spec:
replicas: 2
selector:
matchLabels:
app: example-app
template:
metadata:
labels:
app: example-app
spec:
containers:
- name: example-app
image: quay.io/brancz/prometheus-example-app:v0.5.0
ports:
- name: web
containerPort: 8080
---
apiVersion: v1
kind: Service
metadata:
name: example-app
namespace: demo
labels:
app: example-app
spec:
selector:
app: example-app
ports:
- name: web
port: 8080
targetPort: web
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: example-app
namespace: demo
spec:
selector:
matchLabels:
app: example-app
endpoints:
- port: web
path: /metrics
interval: 30s
Two details matter: the ServiceMonitor selects the Service by its labels (not the pods), and endpoints[].port refers to the name of the Service port, not its number.
Apply it:
kubectl apply -f example-app.yaml
After about a minute, query Prometheus (with the port-forward from Step 4 still running):
curl -s 'http://localhost:9090/api/v1/query' --data-urlencode 'query=up{namespace="demo"}' | jq -r '.data.result[] | "\(.metric.pod) \(.value[1])"'
example-app-7d8c9f6b5d-4jxkq 1
example-app-7d8c9f6b5d-q9w2m 1
A value of 1 means the pod was scraped successfully. The job label defaults to the Service name, so these targets appear as serviceMonitor/demo/example-app/0 on the targets page.
Step 6 - Adding an alert with a PrometheusRule
Alerting rules are also Kubernetes objects. The following rule fires when no example-app target has been reachable for two minutes. Using absent(... == 1) covers both cases: pods that fail to be scraped and pods that no longer exist.
nano example-app-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: example-app
namespace: demo
spec:
groups:
- name: example-app.rules
rules:
- alert: ExampleAppDown
expr: absent(up{job="example-app", namespace="demo"} == 1)
for: 2m
labels:
severity: critical
annotations:
summary: "example-app has no healthy targets"
description: "No example-app pod in namespace demo has been scraped successfully for 2 minutes."
Apply it and confirm that Prometheus loaded the rule:
kubectl apply -f example-app-rules.yaml
curl -s http://localhost:9090/api/v1/rules | jq -r '.data.groups[].name' | grep example-app
example-app.rules
Test the alert by scaling the application to zero:
kubectl -n demo scale deployment example-app --replicas=0
After two to three minutes, the alert changes from pending to firing on the Alerts page in Prometheus and then appears in Alertmanager. Scale back up to resolve it:
kubectl -n demo scale deployment example-app --replicas=2
Step 7 - Sending alerts to Slack
Alertmanager receives the alerts but does not notify anyone until you define a receiver. Add an alertmanager.config block to kube-prometheus-values.yaml. This replaces the chart's default configuration, so it includes a route that keeps the Watchdog alert out of your channel.
First create an incoming webhook in Slack, then add this block at the same level as the existing alertmanager.alertmanagerSpec key, replacing your_slack_webhook_url and #alerts:
alertmanager:
config:
global:
resolve_timeout: 5m
route:
receiver: slack
group_by: ["alertname", "namespace"]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- receiver: "null"
matchers:
- alertname = "Watchdog"
receivers:
- name: "null"
- name: slack
slack_configs:
- api_url: your_slack_webhook_url
channel: "#alerts"
send_resolved: true
Make sure there is only one top-level alertmanager: key in the file, with both config and alertmanagerSpec under it. Apply the change:
helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
--namespace monitoring -f kube-prometheus-values.yaml
Open Alertmanager again and check Status: the loaded configuration should show your slack receiver. Repeat the scale-to-zero test from Step 6 and the alert should arrive in Slack, followed by a resolved message when you scale back up.
Troubleshooting
The ServiceMonitor target never appears. Check that the Service labels match spec.selector.matchLabels, that the port name in endpoints matches a named Service port, and that the Service actually has endpoints (kubectl -n demo get endpoints example-app). If you did not set the *SelectorNilUsesHelmValues: false options, add the label release: kube-prometheus-stack to the ServiceMonitor.
PVCs stay Pending. There is no default StorageClass. Either mark one as default or set storageClassName in each volumeClaimTemplate. Inspect the reason with kubectl -n monitoring describe pvc <name>.
kube-controller-manager, kube-scheduler, etcd or kube-proxy targets are down. On kubeadm clusters these components listen on 127.0.0.1 only, so Prometheus cannot reach them and the related ...Down alerts fire. Either change their bind address to an address reachable from the pods, or disable those components in the chart (for example kubeControllerManager.enabled: false) if you do not need their metrics.
Prometheus is OOMKilled. Memory grows with the number of time series. Raise resources.limits.memory, reduce retention, or drop high cardinality metrics from your applications.
Conclusion
You now have Prometheus, Alertmanager and Grafana running with persistent storage, your own application scraped through a ServiceMonitor, and a tested alert delivered to Slack. From here you can add ServiceMonitors for your databases and ingress controller, build Grafana dashboards for your services, and route critical alerts to additional receivers such as email or PagerDuty.
