The kube-prometheus-stack Helm chart installs the Prometheus Operator together with Prometheus, Alertmanager, Grafana, node-exporter and kube-state-metrics, already wired up with Kubernetes dashboards and alert rules. In this tutorial you will deploy it with persistent storage, open Grafana and Prometheus, scrape metrics from your own application with a ServiceMonitor, add a custom alert rule and route alerts to Slack.

Prerequisites

To follow this guide you need:

  • A Kubernetes cluster running a currently supported version (1.30 or newer), for example a managed cluster or a self-hosted one on CubePath VPS nodes.
  • At least 4 GB of free memory and 2 vCPUs across the worker nodes for the monitoring stack.
  • A default StorageClass so Prometheus, Alertmanager and Grafana can get persistent volumes.
  • A workstation running Ubuntu 24.04 with kubectl configured for the cluster (kubectl get nodes works) and Helm 3.

If you still need the command line tools on Ubuntu 24.04, both are available as snaps:

sudo snap install kubectl --classic
sudo snap install helm --classic

Confirm the cluster has a default storage class. One of the entries must show (default):

kubectl get storageclass
NAME                   PROVISIONER             RECLAIMPOLICY   VOLUMEBINDINGMODE      ALLOWVOLUMEEXPANSION   AGE
local-path (default)   rancher.io/local-path   Delete          WaitForFirstConsumer   false                  12d

Step 1 - Adding the Helm repository

The chart is published by the prometheus-community project. Add the repository and refresh the chart index:

helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update

Check that the chart is visible:

helm search repo prometheus-community/kube-prometheus-stack

The output lists the chart with its current version and the Prometheus Operator app version it ships.

Step 2 - Writing a values file

The chart works with its defaults, but three changes make it production ready: persistent storage, a retention period, and letting Prometheus pick up ServiceMonitor and PrometheusRule objects from any namespace. By default the operator only selects objects labeled release: <your release name>, which is the most common reason custom monitors are silently ignored.

Create the values file on your workstation:

nano kps-values.yaml

Add the following content:

prometheus:
  prometheusSpec:
    retention: 15d
    retentionSize: 45GB
    # Select ServiceMonitors, PodMonitors and rules from all namespaces,
    # not only those labeled with this Helm release
    serviceMonitorSelectorNilUsesHelmValues: false
    podMonitorSelectorNilUsesHelmValues: false
    ruleSelectorNilUsesHelmValues: false
    resources:
      requests:
        cpu: 500m
        memory: 2Gi
      limits:
        memory: 4Gi
    storageSpec:
      volumeClaimTemplate:
        spec:
          accessModes: ["ReadWriteOnce"]
          resources:
            requests:
              storage: 50Gi

alertmanager:
  alertmanagerSpec:
    storage:
      volumeClaimTemplate:
        spec:
          accessModes: ["ReadWriteOnce"]
          resources:
            requests:
              storage: 2Gi

grafana:
  persistence:
    enabled: true
    size: 10Gi

retentionSize is kept below the volume size so Prometheus deletes old blocks before the disk fills up. No Grafana password is set here on purpose: the chart stores the admin password in a Kubernetes secret, which you will read in Step 4.

Step 3 - Installing the chart

Install the stack into a dedicated monitoring namespace. upgrade --install works both for the first install and for every later change to the values file:

helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
  --namespace monitoring --create-namespace \
  -f kps-values.yaml

The operator creates the Prometheus and Alertmanager StatefulSets after the chart is applied, so give it a minute and then check the pods:

kubectl get pods -n monitoring
NAME                                                        READY   STATUS    RESTARTS   AGE
alertmanager-kube-prometheus-stack-alertmanager-0           2/2     Running   0          90s
kube-prometheus-stack-grafana-6d9f7c8b7-x2lqz               3/3     Running   0          2m
kube-prometheus-stack-kube-state-metrics-5c7b9d8f4-r7m2n    1/1     Running   0          2m
kube-prometheus-stack-operator-7f8c9b6d5-jk4tp              1/1     Running   0          2m
kube-prometheus-stack-prometheus-node-exporter-8hx2c        1/1     Running   0          2m
kube-prometheus-stack-prometheus-node-exporter-qv7wd        1/1     Running   0          2m
prometheus-kube-prometheus-stack-prometheus-0               2/2     Running   0          90s

There is one node-exporter pod per node. Also confirm the persistent volume claims are Bound:

kubectl get pvc -n monitoring

Step 4 - Accessing Grafana and Prometheus

The services are only reachable inside the cluster, which is the safe default. Use kubectl port-forward from your workstation to open them.

Read the Grafana admin password from the secret the chart created:

kubectl get secret -n monitoring kube-prometheus-stack-grafana \
  -o jsonpath="{.data.admin-password}" | base64 -d; echo

Forward Grafana to local port 3000 and leave the command running:

kubectl port-forward -n monitoring svc/kube-prometheus-stack-grafana 3000:80

Open http://localhost:3000 and log in as admin with that password. Under Dashboards you will find the bundled Kubernetes dashboards, such as Kubernetes / Compute Resources / Cluster, Kubernetes / Compute Resources / Namespace (Pods) and Node Exporter / Nodes. They are already populated with data.

In a second terminal, forward Prometheus:

kubectl port-forward -n monitoring svc/kube-prometheus-stack-prometheus 9090:9090

Open http://localhost:9090/targets to see every scrape target. You can also query from the command line. The following returns the number of healthy targets per job:

curl -sG http://localhost:9090/api/v1/query \
  --data-urlencode 'query=sum by (job) (up)'

Every job should report a value of at least 1.

Step 5 - Scraping your own application with a ServiceMonitor

A ServiceMonitor tells the operator which Services to scrape and on which port. To demonstrate, deploy podinfo, a small demo app that exposes Prometheus metrics on /metrics.

Create the manifest:

nano podinfo.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: demo
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: podinfo
  namespace: demo
spec:
  replicas: 2
  selector:
    matchLabels:
      app: podinfo
  template:
    metadata:
      labels:
        app: podinfo
    spec:
      containers:
        - name: podinfo
          image: ghcr.io/stefanprodan/podinfo:6.7.1
          ports:
            - name: http
              containerPort: 9898
          resources:
            requests:
              cpu: 10m
              memory: 32Mi
            limits:
              memory: 128Mi
---
apiVersion: v1
kind: Service
metadata:
  name: podinfo
  namespace: demo
  labels:
    app: podinfo
spec:
  selector:
    app: podinfo
  ports:
    - name: http
      port: 9898
      targetPort: http
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: podinfo
  namespace: demo
spec:
  selector:
    matchLabels:
      app: podinfo
  endpoints:
    - port: http
      path: /metrics
      interval: 30s

The ServiceMonitor selects Services by label (app: podinfo) and refers to the port by its name (http), not its number. Apply it:

kubectl apply -f podinfo.yaml

After about a minute, query Prometheus for the new targets (keep the port-forward from Step 4 running):

curl -sG http://localhost:9090/api/v1/query \
  --data-urlencode 'query=up{namespace="demo"}'

The result contains one series per podinfo pod with the value 1. The same targets appear under serviceMonitor/demo/podinfo/0 on the Targets page.

Step 6 - Adding a custom alert rule

The chart already ships a large set of alerts for common failures, such as KubePodCrashLooping, KubeNodeNotReady and KubePersistentVolumeFillingUp, so you do not need to write those yourself. Add rules for what is specific to your workloads. This example fires when a container uses more than 90% of its memory limit for 10 minutes, a good early warning before an OOM kill.

nano memory-rule.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: container-memory
  namespace: demo
spec:
  groups:
    - name: container-memory
      rules:
        - alert: ContainerMemoryNearLimit
          expr: |
            sum by (namespace, pod, container) (container_memory_working_set_bytes{container!=""})
              /
            sum by (namespace, pod, container) (kube_pod_container_resource_limits{resource="memory"})
              > 0.9
          for: 10m
          labels:
            severity: warning
          annotations:
            summary: "{{ $labels.namespace }}/{{ $labels.pod }} ({{ $labels.container }}) is above 90% of its memory limit"

Apply the rule:

kubectl apply -f memory-rule.yaml

Check that Prometheus loaded it:

curl -s http://localhost:9090/api/v1/rules | grep -o '"name":"ContainerMemoryNearLimit"'
"name":"ContainerMemoryNearLimit"

The rule is also listed on http://localhost:9090/alerts, in the container-memory group.

Step 7 - Sending alerts to Slack

Alertmanager receives firing alerts from Prometheus and routes them to receivers. With the chart, the cleanest way to configure it is through the alertmanager.config value. First create an incoming webhook in your Slack workspace and copy its URL.

Open the values file again:

nano kps-values.yaml

Add a config block under the existing alertmanager key, at the same level as alertmanagerSpec. The complete alertmanager section then looks like this:

alertmanager:
  config:
    global:
      resolve_timeout: 5m
    route:
      receiver: "null"
      group_by: ["namespace", "alertname"]
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 12h
      routes:
        - receiver: "null"
          matchers:
            - alertname = "Watchdog"
        - receiver: slack
          matchers:
            - severity =~ "warning|critical"
    receivers:
      - name: "null"
      - name: slack
        slack_configs:
          - api_url: "https://hooks.slack.com/services/your/webhook/url"
            channel: "#alerts"
            send_resolved: true
  alertmanagerSpec:
    storage:
      volumeClaimTemplate:
        spec:
          accessModes: ["ReadWriteOnce"]
          resources:
            requests:
              storage: 2Gi

Replace the api_url with your webhook. The Watchdog alert always fires by design (it proves the alerting pipeline works), so it is routed to the null receiver to keep it out of Slack.

Apply the change:

helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
  --namespace monitoring -f kps-values.yaml

Forward Alertmanager and open http://localhost:9093/#/status to confirm the loaded configuration contains the slack receiver:

kubectl port-forward -n monitoring svc/kube-prometheus-stack-alertmanager 9093:9093

Any active warning or critical alert is now posted to the Slack channel, followed by a resolved message when it clears.

Troubleshooting

A ServiceMonitor is ignored. Check that serviceMonitorSelectorNilUsesHelmValues: false is in the values file and applied, that the matchLabels in the ServiceMonitor match the labels on the Service (not the pods), and that the port field uses the port name defined in the Service.

Control plane targets are down. On managed clusters, disable the components you cannot scrape by adding these keys at the top level of the values file and running the helm upgrade command again:

kubeEtcd:
  enabled: false
kubeScheduler:
  enabled: false
kubeControllerManager:
  enabled: false

Prometheus pod stays Pending. Run kubectl describe pvc -n monitoring and look at the events. Usually there is no default StorageClass or the requested size is not available.

Prometheus is OOM killed. Memory grows with the number of active series. Raise the memory limit in prometheusSpec.resources, or reduce cardinality by dropping unused metrics with metricRelabelings in the relevant ServiceMonitor.

Conclusion

You now have Prometheus, Alertmanager and Grafana monitoring your Kubernetes nodes and workloads, a pattern for scraping any application with a ServiceMonitor, a custom PrometheusRule, and alerts delivered to Slack. As next steps, expose Grafana through an Ingress with TLS and authentication instead of port-forwarding, add ServiceMonitor objects for your databases and ingress controller, and consider remote storage if you need retention beyond a few weeks.