The Horizontal Pod Autoscaler (HPA) adds or removes replicas of a Deployment or StatefulSet based on observed metrics, most often CPU or memory usage relative to what each pod requested. In this tutorial you will install metrics-server, deploy a CPU-bound test application, create an HPA with the autoscaling/v2 API, generate load to watch it scale out and back in, and tune how fast it reacts. The same steps work on kubeadm, k3s and managed Kubernetes clusters.

Prerequisites

To follow this tutorial you need:

  • A Kubernetes cluster (1.30 or newer) with at least two worker nodes or enough spare CPU for about 5 extra small pods.
  • kubectl configured with cluster-admin access, for example from an Ubuntu 24.04 workstation or a CubePath VPS.

How the HPA decides the replica count

The HPA controller runs inside kube-controller-manager and, every 15 seconds by default, compares the current metric with your target:

desiredReplicas = ceil( currentReplicas * currentMetricValue / targetMetricValue )

For example, 4 replicas averaging 90% CPU with a 60% target gives ceil(4 * 90 / 60) = 6 replicas. Two consequences follow:

  • Utilization is relative to the pod's resource requests. A pod without a CPU request has no utilization, and the HPA cannot scale it.
  • When several metrics are configured, the HPA computes a replica count for each one and uses the highest.

Step 1 - Installing metrics-server

Resource-based autoscaling reads CPU and memory from the Metrics API, which metrics-server provides. k3s and most managed services install it by default. Check first:

kubectl get deployment metrics-server -n kube-system

If the command returns NotFound, install the latest release from the official manifest:

kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

Wait for it to become available:

kubectl rollout status deployment/metrics-server -n kube-system
deployment "metrics-server" successfully rolled out

Confirm that metrics flow. The first values appear after about a minute:

kubectl top nodes
NAME     CPU(cores)   CPU(%)   MEMORY(bytes)   MEMORY(%)
node-1   182m         4%       2451Mi          31%
node-2   96m          2%       1804Mi          23%

Step 2 - Deploying a test application with resource requests

Use the hpa-example image from the Kubernetes project: a PHP page that burns CPU on every request, which makes scaling easy to trigger.

nano php-apache.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: php-apache
  namespace: hpa-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: php-apache
  template:
    metadata:
      labels:
        app: php-apache
    spec:
      containers:
        - name: php-apache
          image: registry.k8s.io/hpa-example
          ports:
            - containerPort: 80
          resources:
            requests:
              cpu: 200m
              memory: 64Mi
            limits:
              cpu: 500m
              memory: 128Mi
---
apiVersion: v1
kind: Service
metadata:
  name: php-apache
  namespace: hpa-demo
spec:
  selector:
    app: php-apache
  ports:
    - port: 80

The cpu: 200m request is what the HPA's percentages are measured against: 50% utilization means 100m of CPU per pod on average.

Create the namespace and apply the manifest:

kubectl create namespace hpa-demo
kubectl apply -f php-apache.yaml
kubectl get pods -n hpa-demo
NAME                          READY   STATUS    RESTARTS   AGE
php-apache-7b6c9d5f8b-m4v9s   1/1     Running   0          15s

Step 3 - Creating the HorizontalPodAutoscaler

Define an HPA that keeps average CPU utilization around 50%, with between 1 and 10 replicas:

nano php-apache-hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: php-apache
  namespace: hpa-demo
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: php-apache
  minReplicas: 1
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 50
kubectl apply -f php-apache-hpa.yaml

Check its status. It can show <unknown> for up to a minute until the first metrics arrive:

kubectl get hpa -n hpa-demo
NAME         REFERENCE               TARGETS       MINPODS   MAXPODS   REPLICAS   AGE
php-apache   Deployment/php-apache   cpu: 1%/50%   1         10        1          45s

Step 4 - Generating load and watching it scale

Open a second terminal and start a load generator that requests the page in a tight loop:

kubectl run load-generator -n hpa-demo --rm -i --tty --restart=Never \
  --image=busybox:1.36 \
  -- /bin/sh -c "while sleep 0.01; do wget -q -O- http://php-apache; done"

In the first terminal, watch the HPA:

kubectl get hpa php-apache -n hpa-demo --watch
NAME         REFERENCE               TARGETS         MINPODS   MAXPODS   REPLICAS   AGE
php-apache   Deployment/php-apache   cpu: 1%/50%     1         10        1          2m
php-apache   Deployment/php-apache   cpu: 248%/50%   1         10        1          2m30s
php-apache   Deployment/php-apache   cpu: 248%/50%   1         10        4          2m45s
php-apache   Deployment/php-apache   cpu: 102%/50%   1         10        5          3m
php-apache   Deployment/php-apache   cpu: 49%/50%    1         10        5          3m30s

CPU jumps well above the target, the HPA adds replicas, and utilization settles around 50%. The exact numbers depend on your nodes. Press CTRL+C to stop watching and look at the decisions the controller made:

kubectl describe hpa php-apache -n hpa-demo
...
Events:
  Type    Reason             Age   From                       Message
  ----    ------             ----  ----                       -------
  Normal  SuccessfulRescale  90s   horizontal-pod-autoscaler  New size: 4; reason: cpu resource utilization (percentage of request) above target
  Normal  SuccessfulRescale  75s   horizontal-pod-autoscaler  New size: 5; reason: cpu resource utilization (percentage of request) above target

Now stop the load generator with CTRL+C in the second terminal and watch again. CPU drops to near 0% within a minute, but the replica count stays at 5 for about five minutes before it returns to 1. That delay is the default scale-down stabilization window, which prevents the HPA from removing pods during a short lull and adding them back seconds later.

Step 5 - Tuning scaling behavior

The behavior field controls how fast the HPA may change the replica count in each direction. Without it, the defaults are:

  • Scale up: no stabilization window; every 15 seconds it may add 100% of the current replicas or 4 pods, whichever is more.
  • Scale down: a 300 second stabilization window (the HPA uses the highest recommendation of the last five minutes); it may remove up to 100% of the pods per 15 seconds after that.

A common production profile scales up quickly but scales down gently. Add this block under spec in php-apache-hpa.yaml, at the same level as metrics:

  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0
      policies:
        - type: Percent
          value: 100
          periodSeconds: 15
        - type: Pods
          value: 4
          periodSeconds: 15
      selectPolicy: Max
    scaleDown:
      stabilizationWindowSeconds: 300
      policies:
        - type: Pods
          value: 1
          periodSeconds: 60
      selectPolicy: Min

How to read it:

  • Each policy limits the change allowed within periodSeconds. Percent is relative to the current replica count; Pods is an absolute number.
  • selectPolicy: Max picks the policy that allows the biggest change, Min the smallest, and Disabled turns scaling in that direction off.
  • With this scale-down policy, the HPA removes at most one pod per minute after the five minute window, so a traffic drop from 10 to 1 replicas takes about 14 minutes.

Apply the change and confirm it was accepted:

kubectl apply -f php-apache-hpa.yaml
kubectl get hpa php-apache -n hpa-demo -o jsonpath='{.spec.behavior.scaleDown}{"\n"}'
{"policies":[{"periodSeconds":60,"type":"Pods","value":1}],"selectPolicy":"Min","stabilizationWindowSeconds":300}

Run the load generator from Step 4 again and stop it to see the slower, step-by-step scale down.

Step 6 - Scaling on memory as well

You can combine several metrics in one HPA. Add memory to the metrics list so the HPA also scales out when pods use more than 70% of their memory request:

  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 50
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 70

Apply it and check that both targets appear:

kubectl apply -f php-apache-hpa.yaml
kubectl get hpa -n hpa-demo
NAME         REFERENCE               TARGETS                        MINPODS   MAXPODS   REPLICAS   AGE
php-apache   Deployment/php-apache   cpu: 1%/50%, memory: 18%/70%   1         10        1          20m

Use memory scaling with care. Many runtimes (JVM, Node.js, most caches) do not release memory when load drops, so the HPA scales out but never scales back in. Memory targets work best for applications whose memory use tracks the number of requests in flight.

You can also target an absolute amount instead of a percentage with type: AverageValue, for example averageValue: 300m for CPU. That form does not depend on requests, but CPU requests are still needed for sensible scheduling.

Scaling on application metrics

CPU is only a proxy for load. For queue workers or APIs, scaling on requests per second or queue length is often more accurate. The HPA supports this through two additional metric types:

  • type: Pods and type: Object read the custom metrics API, which an adapter such as prometheus-adapter serves from Prometheus data.
  • type: External reads the external metrics API, for metrics that do not belong to a Kubernetes object, such as the length of a message queue.

KEDA is a popular way to get both without writing adapter rules: it ships scalers for Prometheus, RabbitMQ, Kafka, Redis, cron schedules and more, and creates and manages the HPA for you.

HPA and the Vertical Pod Autoscaler

The Vertical Pod Autoscaler (VPA) is a separate project that adjusts the CPU and memory requests of pods instead of the number of replicas.

HPAVPA
ChangesNumber of replicasRequests (and limits) per pod
Best forStateless services behind a load balancerRight-sizing workloads that cannot scale out
Installed by defaultYes (controller-manager)No

Do not let both act on the same CPU or memory metric for one workload: the VPA raises requests, which lowers utilization, which makes the HPA scale in, and the two loop. A safe combination is the VPA in recommendation-only mode (updateMode: "Off") to choose good requests, and the HPA for scaling.

Cleaning up

Delete the namespace to remove the Deployment, Service and HPA:

kubectl delete namespace hpa-demo

Troubleshooting

  • TARGETS stays at <unknown>: the HPA cannot read metrics. Check kubectl top pods -n hpa-demo. If that fails, look at kubectl logs -n kube-system deployment/metrics-server. If it works, the target pods are probably missing a CPU or memory request; kubectl describe hpa then shows FailedGetResourceMetric with missing request for cpu.
  • kubectl top returns Metrics API not available: metrics-server is not running or its APIService is not available. Run kubectl get apiservice v1beta1.metrics.k8s.io and check that AVAILABLE is True.
  • Replicas stay at maxReplicas: the workload needs more capacity than allowed, or new pods are Pending because the nodes are full. Check kubectl get pods and add nodes or raise maxReplicas.
  • Replica count flaps up and down: increase scaleDown.stabilizationWindowSeconds or add a gentler scale-down policy as in Step 5, and make sure the Deployment manifest does not set replicas on every apply.

Conclusion

You installed metrics-server, created a CPU-based HorizontalPodAutoscaler, watched it scale a real workload out under load and back in after the stabilization window, tuned its behavior, and added memory as a second metric. Next, set realistic CPU and memory requests for your own workloads (the VPA in recommendation mode helps), combine the HPA with the cluster autoscaler when your nodes can grow, and consider KEDA for scaling on queue length or request rate.