kube-prometheus-stack is a Helm chart that deploys the Prometheus Operator together with Prometheus, Alertmanager, Grafana, node-exporter and kube-state-metrics, preconfigured with dashboards and alerting rules for Kubernetes itself. In this tutorial you will install the chart with a small values file, reach each web UI, scrape a sample application with a ServiceMonitor, and add your own alert with a PrometheusRule. The commands run from any workstation with kubectl access to the cluster; the examples assume an Ubuntu 24.04 workstation.

Prerequisites

To follow this guide you need:

  • A working Kubernetes cluster (1.30 or newer) with at least 4 GB of free memory across the worker nodes, for example on CubePath VPS or bare metal servers.
  • kubectl configured against that cluster with cluster-admin rights. Check with kubectl get nodes.
  • Helm 3 installed on your workstation. Check with helm version.
  • jq for reading the Prometheus API output: sudo apt install jq.
  • A default StorageClass so Prometheus, Alertmanager and Grafana can keep their data. Check with kubectl get storageclass and look for (default) next to one of them.

Step 1 - Adding the Helm repository

The chart is published by the Prometheus community. Add the repository and refresh the index:

helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update

Confirm that the chart is available:

helm search repo prometheus-community/kube-prometheus-stack
NAME                                        CHART VERSION  APP VERSION  DESCRIPTION
prometheus-community/kube-prometheus-stack  ...            v0.8x.x      kube-prometheus-stack collects Kubernetes manif...

Step 2 - Writing a values file

The chart works out of the box, but three things are worth setting from the start: persistent storage (otherwise metrics are lost when the Prometheus pod restarts), a Grafana admin password you choose, and the selector settings.

By default Prometheus only picks up ServiceMonitor, PodMonitor and PrometheusRule objects that carry the label release: <your release name>. That is the most common reason custom monitors are silently ignored. Setting the *SelectorNilUsesHelmValues options to false makes Prometheus select all of them in every namespace.

Create the file:

nano kube-prometheus-values.yaml

Add the following content, replacing your_strong_password with a real password:

grafana:
  adminPassword: your_strong_password
  persistence:
    enabled: true
    size: 5Gi

prometheus:
  prometheusSpec:
    retention: 15d
    retentionSize: 18GB
    serviceMonitorSelectorNilUsesHelmValues: false
    podMonitorSelectorNilUsesHelmValues: false
    ruleSelectorNilUsesHelmValues: false
    resources:
      requests:
        cpu: 200m
        memory: 1Gi
      limits:
        memory: 3Gi
    storageSpec:
      volumeClaimTemplate:
        spec:
          accessModes: ["ReadWriteOnce"]
          resources:
            requests:
              storage: 20Gi

alertmanager:
  alertmanagerSpec:
    storage:
      volumeClaimTemplate:
        spec:
          accessModes: ["ReadWriteOnce"]
          resources:
            requests:
              storage: 2Gi

retentionSize is kept a little below the volume size so Prometheus deletes old blocks before the disk fills up. No storageClassName is set, so the cluster's default StorageClass is used; add one under each spec if you want a specific class.

Step 3 - Installing the chart

Install the chart into a dedicated monitoring namespace. The release name kube-prometheus-stack determines the names of every service used later in this guide:

helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
  --namespace monitoring --create-namespace \
  -f kube-prometheus-values.yaml

helm upgrade --install installs the release the first time and upgrades it on later runs, so you can use the same command whenever you change the values file.

Wait until all pods are ready:

kubectl -n monitoring wait --for=condition=Ready pods --all --timeout=300s
kubectl -n monitoring get pods
NAME                                                        READY   STATUS    RESTARTS   AGE
alertmanager-kube-prometheus-stack-alertmanager-0           2/2     Running   0          2m
kube-prometheus-stack-grafana-6c9b8d7f8d-xk2lq              3/3     Running   0          2m
kube-prometheus-stack-kube-state-metrics-5d9c6b8c7f-9hq4w   1/1     Running   0          2m
kube-prometheus-stack-operator-7f6c9d5b8-pl7m2              1/1     Running   0          2m
kube-prometheus-stack-prometheus-node-exporter-4zk8r        1/1     Running   0          2m
kube-prometheus-stack-prometheus-node-exporter-v7c2n        1/1     Running   0          2m
prometheus-kube-prometheus-stack-prometheus-0               2/2     Running   0          2m

There is one node-exporter pod per node. Check that the persistent volumes were bound:

kubectl -n monitoring get pvc

All claims should show STATUS Bound. If they stay Pending, see the troubleshooting section.

Step 4 - Accessing Prometheus, Grafana and Alertmanager

The chart creates ClusterIP services, which are not reachable from outside the cluster. For administration, kubectl port-forward is the simplest and safest option because nothing is exposed publicly.

Open a tunnel to Prometheus:

kubectl -n monitoring port-forward svc/kube-prometheus-stack-prometheus 9090:9090

Browse to http://localhost:9090, open Status > Targets (called Target health in Prometheus 3) and check that the targets are UP. In a second terminal you can also query the API directly:

curl -s 'http://localhost:9090/api/v1/query?query=up' | jq '.data.result | length'

The output is the number of scraped targets, which should be greater than zero.

Open a tunnel to Grafana, whose service listens on port 80:

kubectl -n monitoring port-forward svc/kube-prometheus-stack-grafana 3000:80

Browse to http://localhost:3000 and log in as admin with the password from your values file. If you forget it, read it from the secret the chart created:

kubectl -n monitoring get secret kube-prometheus-stack-grafana -o jsonpath='{.data.admin-password}' | base64 -d; echo

Under Dashboards you will find the bundled dashboards, for example Kubernetes / Compute Resources / Cluster and Node Exporter / Nodes. They already show data because Prometheus is configured as the default data source.

Alertmanager works the same way:

kubectl -n monitoring port-forward svc/kube-prometheus-stack-alertmanager 9093:9093

At http://localhost:9093 you should see at least the Watchdog alert. It fires permanently on purpose, to prove that the alerting pipeline works end to end.

Step 5 - Scraping your own application with a ServiceMonitor

A ServiceMonitor tells the operator which Services to scrape and on which port. To test it, deploy the small example application used by the Prometheus Operator project, which exposes metrics on port 8080.

Create a manifest:

nano example-app.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: demo
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: example-app
  namespace: demo
spec:
  replicas: 2
  selector:
    matchLabels:
      app: example-app
  template:
    metadata:
      labels:
        app: example-app
    spec:
      containers:
        - name: example-app
          image: quay.io/brancz/prometheus-example-app:v0.5.0
          ports:
            - name: web
              containerPort: 8080
---
apiVersion: v1
kind: Service
metadata:
  name: example-app
  namespace: demo
  labels:
    app: example-app
spec:
  selector:
    app: example-app
  ports:
    - name: web
      port: 8080
      targetPort: web
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: example-app
  namespace: demo
spec:
  selector:
    matchLabels:
      app: example-app
  endpoints:
    - port: web
      path: /metrics
      interval: 30s

Two details matter: the ServiceMonitor selects the Service by its labels (not the pods), and endpoints[].port refers to the name of the Service port, not its number.

Apply it:

kubectl apply -f example-app.yaml

After about a minute, query Prometheus (with the port-forward from Step 4 still running):

curl -s 'http://localhost:9090/api/v1/query' --data-urlencode 'query=up{namespace="demo"}' | jq -r '.data.result[] | "\(.metric.pod) \(.value[1])"'
example-app-7d8c9f6b5d-4jxkq 1
example-app-7d8c9f6b5d-q9w2m 1

A value of 1 means the pod was scraped successfully. The job label defaults to the Service name, so these targets appear as serviceMonitor/demo/example-app/0 on the targets page.

Step 6 - Adding an alert with a PrometheusRule

Alerting rules are also Kubernetes objects. The following rule fires when no example-app target has been reachable for two minutes. Using absent(... == 1) covers both cases: pods that fail to be scraped and pods that no longer exist.

nano example-app-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: example-app
  namespace: demo
spec:
  groups:
    - name: example-app.rules
      rules:
        - alert: ExampleAppDown
          expr: absent(up{job="example-app", namespace="demo"} == 1)
          for: 2m
          labels:
            severity: critical
          annotations:
            summary: "example-app has no healthy targets"
            description: "No example-app pod in namespace demo has been scraped successfully for 2 minutes."

Apply it and confirm that Prometheus loaded the rule:

kubectl apply -f example-app-rules.yaml
curl -s http://localhost:9090/api/v1/rules | jq -r '.data.groups[].name' | grep example-app
example-app.rules

Test the alert by scaling the application to zero:

kubectl -n demo scale deployment example-app --replicas=0

After two to three minutes, the alert changes from pending to firing on the Alerts page in Prometheus and then appears in Alertmanager. Scale back up to resolve it:

kubectl -n demo scale deployment example-app --replicas=2

Step 7 - Sending alerts to Slack

Alertmanager receives the alerts but does not notify anyone until you define a receiver. Add an alertmanager.config block to kube-prometheus-values.yaml. This replaces the chart's default configuration, so it includes a route that keeps the Watchdog alert out of your channel.

First create an incoming webhook in Slack, then add this block at the same level as the existing alertmanager.alertmanagerSpec key, replacing your_slack_webhook_url and #alerts:

alertmanager:
  config:
    global:
      resolve_timeout: 5m
    route:
      receiver: slack
      group_by: ["alertname", "namespace"]
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 4h
      routes:
        - receiver: "null"
          matchers:
            - alertname = "Watchdog"
    receivers:
      - name: "null"
      - name: slack
        slack_configs:
          - api_url: your_slack_webhook_url
            channel: "#alerts"
            send_resolved: true

Make sure there is only one top-level alertmanager: key in the file, with both config and alertmanagerSpec under it. Apply the change:

helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
  --namespace monitoring -f kube-prometheus-values.yaml

Open Alertmanager again and check Status: the loaded configuration should show your slack receiver. Repeat the scale-to-zero test from Step 6 and the alert should arrive in Slack, followed by a resolved message when you scale back up.

Troubleshooting

The ServiceMonitor target never appears. Check that the Service labels match spec.selector.matchLabels, that the port name in endpoints matches a named Service port, and that the Service actually has endpoints (kubectl -n demo get endpoints example-app). If you did not set the *SelectorNilUsesHelmValues: false options, add the label release: kube-prometheus-stack to the ServiceMonitor.

PVCs stay Pending. There is no default StorageClass. Either mark one as default or set storageClassName in each volumeClaimTemplate. Inspect the reason with kubectl -n monitoring describe pvc <name>.

kube-controller-manager, kube-scheduler, etcd or kube-proxy targets are down. On kubeadm clusters these components listen on 127.0.0.1 only, so Prometheus cannot reach them and the related ...Down alerts fire. Either change their bind address to an address reachable from the pods, or disable those components in the chart (for example kubeControllerManager.enabled: false) if you do not need their metrics.

Prometheus is OOMKilled. Memory grows with the number of time series. Raise resources.limits.memory, reduce retention, or drop high cardinality metrics from your applications.

Conclusion

You now have Prometheus, Alertmanager and Grafana running with persistent storage, your own application scraped through a ServiceMonitor, and a tested alert delivered to Slack. From here you can add ServiceMonitors for your databases and ingress controller, build Grafana dashboards for your services, and route critical alerts to additional receivers such as email or PagerDuty.