The Horizontal Pod Autoscaler (HPA) adds or removes replicas of a Deployment or StatefulSet based on observed metrics, most often CPU or memory usage relative to what each pod requested. In this tutorial you will install metrics-server, deploy a CPU-bound test application, create an HPA with the autoscaling/v2 API, generate load to watch it scale out and back in, and tune how fast it reacts. The same steps work on kubeadm, k3s and managed Kubernetes clusters.
Prerequisites
To follow this tutorial you need:
- A Kubernetes cluster (1.30 or newer) with at least two worker nodes or enough spare CPU for about 5 extra small pods.
kubectlconfigured with cluster-admin access, for example from an Ubuntu 24.04 workstation or a CubePath VPS.
How the HPA decides the replica count
The HPA controller runs inside kube-controller-manager and, every 15 seconds by default, compares the current metric with your target:
desiredReplicas = ceil( currentReplicas * currentMetricValue / targetMetricValue )
For example, 4 replicas averaging 90% CPU with a 60% target gives ceil(4 * 90 / 60) = 6 replicas. Two consequences follow:
- Utilization is relative to the pod's resource requests. A pod without a CPU request has no utilization, and the HPA cannot scale it.
- When several metrics are configured, the HPA computes a replica count for each one and uses the highest.
Step 1 - Installing metrics-server
Resource-based autoscaling reads CPU and memory from the Metrics API, which metrics-server provides. k3s and most managed services install it by default. Check first:
kubectl get deployment metrics-server -n kube-system
If the command returns NotFound, install the latest release from the official manifest:
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
Wait for it to become available:
kubectl rollout status deployment/metrics-server -n kube-system
deployment "metrics-server" successfully rolled out
Confirm that metrics flow. The first values appear after about a minute:
kubectl top nodes
NAME CPU(cores) CPU(%) MEMORY(bytes) MEMORY(%)
node-1 182m 4% 2451Mi 31%
node-2 96m 2% 1804Mi 23%
NoteOn kubeadm clusters, kubelets often use self-signed serving certificates, and metrics-server logs
x509: cannot validate certificate. The proper fix is to enableserverTLSBootstrap: truein the kubelet configuration and approve the kubelet CSRs. For a lab cluster you can instead add--kubelet-insecure-tlsto the metrics-server container arguments withkubectl edit deployment metrics-server -n kube-system, which skips verification of kubelet certificates.
Step 2 - Deploying a test application with resource requests
Use the hpa-example image from the Kubernetes project: a PHP page that burns CPU on every request, which makes scaling easy to trigger.
nano php-apache.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: php-apache
namespace: hpa-demo
spec:
replicas: 1
selector:
matchLabels:
app: php-apache
template:
metadata:
labels:
app: php-apache
spec:
containers:
- name: php-apache
image: registry.k8s.io/hpa-example
ports:
- containerPort: 80
resources:
requests:
cpu: 200m
memory: 64Mi
limits:
cpu: 500m
memory: 128Mi
---
apiVersion: v1
kind: Service
metadata:
name: php-apache
namespace: hpa-demo
spec:
selector:
app: php-apache
ports:
- port: 80
The cpu: 200m request is what the HPA's percentages are measured against: 50% utilization means 100m of CPU per pod on average.
Create the namespace and apply the manifest:
kubectl create namespace hpa-demo
kubectl apply -f php-apache.yaml
kubectl get pods -n hpa-demo
NAME READY STATUS RESTARTS AGE
php-apache-7b6c9d5f8b-m4v9s 1/1 Running 0 15s
Step 3 - Creating the HorizontalPodAutoscaler
Define an HPA that keeps average CPU utilization around 50%, with between 1 and 10 replicas:
nano php-apache-hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: php-apache
namespace: hpa-demo
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: php-apache
minReplicas: 1
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
kubectl apply -f php-apache-hpa.yaml
Check its status. It can show <unknown> for up to a minute until the first metrics arrive:
kubectl get hpa -n hpa-demo
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
php-apache Deployment/php-apache cpu: 1%/50% 1 10 1 45s
ImportantWhen an HPA manages a Deployment, remove
replicasfrom the Deployment manifest you keep in Git (or leave it out on later applies). Otherwise everykubectl applyor GitOps sync resets the replica count and fights the autoscaler.
Step 4 - Generating load and watching it scale
Open a second terminal and start a load generator that requests the page in a tight loop:
kubectl run load-generator -n hpa-demo --rm -i --tty --restart=Never \
--image=busybox:1.36 \
-- /bin/sh -c "while sleep 0.01; do wget -q -O- http://php-apache; done"
In the first terminal, watch the HPA:
kubectl get hpa php-apache -n hpa-demo --watch
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
php-apache Deployment/php-apache cpu: 1%/50% 1 10 1 2m
php-apache Deployment/php-apache cpu: 248%/50% 1 10 1 2m30s
php-apache Deployment/php-apache cpu: 248%/50% 1 10 4 2m45s
php-apache Deployment/php-apache cpu: 102%/50% 1 10 5 3m
php-apache Deployment/php-apache cpu: 49%/50% 1 10 5 3m30s
CPU jumps well above the target, the HPA adds replicas, and utilization settles around 50%. The exact numbers depend on your nodes. Press CTRL+C to stop watching and look at the decisions the controller made:
kubectl describe hpa php-apache -n hpa-demo
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal SuccessfulRescale 90s horizontal-pod-autoscaler New size: 4; reason: cpu resource utilization (percentage of request) above target
Normal SuccessfulRescale 75s horizontal-pod-autoscaler New size: 5; reason: cpu resource utilization (percentage of request) above target
Now stop the load generator with CTRL+C in the second terminal and watch again. CPU drops to near 0% within a minute, but the replica count stays at 5 for about five minutes before it returns to 1. That delay is the default scale-down stabilization window, which prevents the HPA from removing pods during a short lull and adding them back seconds later.
Step 5 - Tuning scaling behavior
The behavior field controls how fast the HPA may change the replica count in each direction. Without it, the defaults are:
- Scale up: no stabilization window; every 15 seconds it may add 100% of the current replicas or 4 pods, whichever is more.
- Scale down: a 300 second stabilization window (the HPA uses the highest recommendation of the last five minutes); it may remove up to 100% of the pods per 15 seconds after that.
A common production profile scales up quickly but scales down gently. Add this block under spec in php-apache-hpa.yaml, at the same level as metrics:
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
- type: Pods
value: 4
periodSeconds: 15
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Pods
value: 1
periodSeconds: 60
selectPolicy: Min
How to read it:
- Each policy limits the change allowed within
periodSeconds.Percentis relative to the current replica count;Podsis an absolute number. selectPolicy: Maxpicks the policy that allows the biggest change,Minthe smallest, andDisabledturns scaling in that direction off.- With this scale-down policy, the HPA removes at most one pod per minute after the five minute window, so a traffic drop from 10 to 1 replicas takes about 14 minutes.
Apply the change and confirm it was accepted:
kubectl apply -f php-apache-hpa.yaml
kubectl get hpa php-apache -n hpa-demo -o jsonpath='{.spec.behavior.scaleDown}{"\n"}'
{"policies":[{"periodSeconds":60,"type":"Pods","value":1}],"selectPolicy":"Min","stabilizationWindowSeconds":300}
Run the load generator from Step 4 again and stop it to see the slower, step-by-step scale down.
Step 6 - Scaling on memory as well
You can combine several metrics in one HPA. Add memory to the metrics list so the HPA also scales out when pods use more than 70% of their memory request:
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 70
Apply it and check that both targets appear:
kubectl apply -f php-apache-hpa.yaml
kubectl get hpa -n hpa-demo
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
php-apache Deployment/php-apache cpu: 1%/50%, memory: 18%/70% 1 10 1 20m
Use memory scaling with care. Many runtimes (JVM, Node.js, most caches) do not release memory when load drops, so the HPA scales out but never scales back in. Memory targets work best for applications whose memory use tracks the number of requests in flight.
You can also target an absolute amount instead of a percentage with type: AverageValue, for example averageValue: 300m for CPU. That form does not depend on requests, but CPU requests are still needed for sensible scheduling.
Scaling on application metrics
CPU is only a proxy for load. For queue workers or APIs, scaling on requests per second or queue length is often more accurate. The HPA supports this through two additional metric types:
type: Podsandtype: Objectread the custom metrics API, which an adapter such as prometheus-adapter serves from Prometheus data.type: Externalreads the external metrics API, for metrics that do not belong to a Kubernetes object, such as the length of a message queue.
KEDA is a popular way to get both without writing adapter rules: it ships scalers for Prometheus, RabbitMQ, Kafka, Redis, cron schedules and more, and creates and manages the HPA for you.
HPA and the Vertical Pod Autoscaler
The Vertical Pod Autoscaler (VPA) is a separate project that adjusts the CPU and memory requests of pods instead of the number of replicas.
| HPA | VPA | |
|---|---|---|
| Changes | Number of replicas | Requests (and limits) per pod |
| Best for | Stateless services behind a load balancer | Right-sizing workloads that cannot scale out |
| Installed by default | Yes (controller-manager) | No |
Do not let both act on the same CPU or memory metric for one workload: the VPA raises requests, which lowers utilization, which makes the HPA scale in, and the two loop. A safe combination is the VPA in recommendation-only mode (updateMode: "Off") to choose good requests, and the HPA for scaling.
Cleaning up
Delete the namespace to remove the Deployment, Service and HPA:
kubectl delete namespace hpa-demo
Troubleshooting
TARGETSstays at<unknown>: the HPA cannot read metrics. Checkkubectl top pods -n hpa-demo. If that fails, look atkubectl logs -n kube-system deployment/metrics-server. If it works, the target pods are probably missing a CPU or memory request;kubectl describe hpathen showsFailedGetResourceMetricwithmissing request for cpu.kubectl topreturnsMetrics API not available: metrics-server is not running or its APIService is not available. Runkubectl get apiservice v1beta1.metrics.k8s.ioand check thatAVAILABLEisTrue.- Replicas stay at
maxReplicas: the workload needs more capacity than allowed, or new pods arePendingbecause the nodes are full. Checkkubectl get podsand add nodes or raisemaxReplicas. - Replica count flaps up and down: increase
scaleDown.stabilizationWindowSecondsor add a gentler scale-down policy as in Step 5, and make sure the Deployment manifest does not setreplicason every apply.
Conclusion
You installed metrics-server, created a CPU-based HorizontalPodAutoscaler, watched it scale a real workload out under load and back in after the stabilization window, tuned its behavior, and added memory as a second metric. Next, set realistic CPU and memory requests for your own workloads (the VPA in recommendation mode helps), combine the HPA with the cluster autoscaler when your nodes can grow, and consider KEDA for scaling on queue length or request rate.
