KEDA (Kubernetes Event-Driven Autoscaling) lets a Deployment scale on external signals such as queue length, consumer lag, a Prometheus query or a schedule, instead of only CPU and memory. It does this by feeding external metrics to a regular Horizontal Pod Autoscaler (HPA) and by scaling workloads between zero and one replica itself. In this tutorial you will install KEDA with Helm, scale a worker Deployment based on the length of a RabbitMQ queue (including scale to zero), and add a cron trigger for predictable traffic.

Prerequisites

To follow this guide you need:

  • A working Kubernetes cluster running a version supported by the current KEDA release (check the compatibility table at keda.sh). A cluster built on CubePath VPS instances works fine.
  • kubectl configured with cluster-admin access, since KEDA installs CRDs and an APIService.
  • Helm 3 installed on your workstation.
  • About 200 MB of free memory in the cluster for the KEDA components and a small RabbitMQ pod used in the demo.

KEDA does not need Metrics Server for event-driven triggers. You only need Metrics Server if you also use the cpu or memory scalers.

Step 1 - Installing KEDA with Helm

Add the official KEDA chart repository and install the chart into its own namespace:

helm repo add kedacore https://kedacore.github.io/charts
helm repo update
helm install keda kedacore/keda --namespace keda --create-namespace

Wait until the three KEDA components are running:

kubectl get pods -n keda
NAME                                               READY   STATUS    RESTARTS   AGE
keda-admission-webhooks-6b4d8f9c7d-x2kqp           1/1     Running   0          45s
keda-operator-7c9b5d6f8-jq4tn                      1/1     Running   0          45s
keda-operator-metrics-apiserver-5f7d9c8b6d-lm8zr   1/1     Running   0          45s
  • keda-operator watches ScaledObject and ScaledJob resources and handles scaling to and from zero.
  • keda-operator-metrics-apiserver serves external metrics to the HPA.
  • keda-admission-webhooks validates KEDA resources when you apply them.

Confirm that the external metrics API is registered and available:

kubectl get apiservice v1beta1.external.metrics.k8s.io
NAME                              SERVICE                                AVAILABLE   AGE
v1beta1.external.metrics.k8s.io   keda/keda-operator-metrics-apiserver   True        60s

Step 2 - Deploying a demo queue and worker

To see KEDA in action you need an event source and a workload to scale. Create a namespace and a Secret holding the RabbitMQ credentials. Replace your_strong_password with a real password:

kubectl create namespace keda-demo
kubectl create secret generic rabbitmq-auth -n keda-demo \
  --from-literal=username=keda \
  --from-literal=password=your_strong_password \
  --from-literal=host='amqp://keda:[email protected]:5672'

The host key is the AMQP connection string KEDA will use. It has no trailing slash, which makes the client use the default virtual host /.

Create the manifest for a single RabbitMQ instance and a placeholder worker:

nano demo.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: rabbitmq
  namespace: keda-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: rabbitmq
  template:
    metadata:
      labels:
        app: rabbitmq
    spec:
      containers:
        - name: rabbitmq
          image: rabbitmq:4-management
          env:
            - name: RABBITMQ_DEFAULT_USER
              valueFrom:
                secretKeyRef:
                  name: rabbitmq-auth
                  key: username
            - name: RABBITMQ_DEFAULT_PASS
              valueFrom:
                secretKeyRef:
                  name: rabbitmq-auth
                  key: password
          ports:
            - containerPort: 5672
            - containerPort: 15672
---
apiVersion: v1
kind: Service
metadata:
  name: rabbitmq
  namespace: keda-demo
spec:
  selector:
    app: rabbitmq
  ports:
    - name: amqp
      port: 5672
    - name: management
      port: 15672
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: worker
  namespace: keda-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: worker
  template:
    metadata:
      labels:
        app: worker
    spec:
      containers:
        - name: worker
          image: registry.k8s.io/pause:3.10
          resources:
            requests:
              cpu: 10m
              memory: 16Mi

The worker Deployment runs a tiny pause container. In a real setup this would be your consumer application; for the demo it only needs to exist so KEDA can change its replica count. Apply the manifest and wait for RabbitMQ:

kubectl apply -f demo.yaml
kubectl rollout status deployment/rabbitmq -n keda-demo
deployment "rabbitmq" successfully rolled out

KEDA's AMQP check fails if the queue does not exist, so create a durable queue named tasks through the RabbitMQ management API. This runs a temporary curl pod inside the cluster:

kubectl run mq-setup -n keda-demo --rm -i --restart=Never --image=curlimages/curl -- \
  curl -s -o /dev/null -w '%{http_code}\n' -u keda:your_strong_password \
  -X PUT -H 'content-type: application/json' -d '{"durable":true}' \
  http://rabbitmq:15672/api/queues/%2F/tasks
201
pod "mq-setup" deleted

A 201 means the queue was created (204 if it already existed).

Step 3 - Scaling on queue length with a ScaledObject

A ScaledObject connects a workload to one or more triggers. A TriggerAuthentication tells KEDA where to read connection secrets, so credentials never appear in the ScaledObject itself.

nano worker-scaler.yaml
apiVersion: keda.sh/v1alpha1
kind: TriggerAuthentication
metadata:
  name: rabbitmq-trigger-auth
  namespace: keda-demo
spec:
  secretTargetRef:
    - parameter: host
      name: rabbitmq-auth
      key: host
---
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: worker-scaler
  namespace: keda-demo
spec:
  scaleTargetRef:
    name: worker
  pollingInterval: 15
  cooldownPeriod: 60
  minReplicaCount: 0
  maxReplicaCount: 10
  triggers:
    - type: rabbitmq
      metadata:
        protocol: amqp
        queueName: tasks
        mode: QueueLength
        value: "5"
        activationValue: "0"
      authenticationRef:
        name: rabbitmq-trigger-auth

What the important fields do:

FieldMeaning
pollingIntervalHow often, in seconds, KEDA checks the trigger.
cooldownPeriodSeconds to wait after the last active trigger before scaling to zero. It only applies to the 1 to 0 step; scaling between 1 and N is handled by the HPA.
minReplicaCount: 0Allows scale to zero when the queue is empty.
value: "5"Target of 5 messages per replica. 50 messages means 10 replicas.
activationValue: "0"The workload is activated (scaled from 0 to 1) when the queue has more than 0 messages.

Apply it:

kubectl apply -f worker-scaler.yaml
kubectl get scaledobject worker-scaler -n keda-demo
NAME            SCALETARGETKIND      SCALETARGETNAME   MIN   MAX   READY   ACTIVE   ...
worker-scaler   apps/v1.Deployment   worker            0     10    True    False    ...

READY True means KEDA can reach RabbitMQ and read the queue. ACTIVE False means the queue is empty, so after cooldownPeriod (60 seconds) the worker is scaled to zero:

kubectl get deployment worker -n keda-demo
NAME     READY   UP-TO-DATE   AVAILABLE   AGE
worker   0/0     0            0           4m

KEDA also created an HPA named keda-hpa-<scaledobject-name>:

kubectl get hpa -n keda-demo
NAME                     REFERENCE           TARGETS     MINPODS   MAXPODS   REPLICAS   AGE
keda-hpa-worker-scaler   Deployment/worker   <unknown>/5 (avg)   1         10        0          2m

Step 4 - Testing scale up and scale to zero

Publish 50 messages to the tasks queue through the management API:

kubectl run mq-publish -n keda-demo --rm -i --restart=Never --image=curlimages/curl -- sh -c '
for i in $(seq 1 50); do
  curl -s -o /dev/null -u keda:your_strong_password \
    -H "content-type: application/json" \
    -d "{\"properties\":{},\"routing_key\":\"tasks\",\"payload\":\"job-$i\",\"payload_encoding\":\"string\"}" \
    http://rabbitmq:15672/api/exchanges/%2F/amq.default/publish
done
echo published'

Watch the worker Deployment. Within one polling interval KEDA activates it, and the HPA then scales it toward 10 replicas (50 messages / 5 per replica):

kubectl get deployment worker -n keda-demo -w
NAME     READY   UP-TO-DATE   AVAILABLE   AGE
worker   0/0     0            0           6m
worker   0/1     1            0           6m
worker   1/1     1            1           6m
worker   4/4     4            4           6m
worker   10/10   10           10          7m

Press Ctrl+C to stop watching. Because the pause container never consumes messages, the queue stays full. Purge it to simulate the workers finishing the backlog:

kubectl run mq-purge -n keda-demo --rm -i --restart=Never --image=curlimages/curl -- \
  curl -s -o /dev/null -w '%{http_code}\n' -u keda:your_strong_password \
  -X DELETE http://rabbitmq:15672/api/queues/%2F/tasks/contents

The HPA scales down gradually (its default scale-down stabilization window is 5 minutes), and once the trigger has been inactive for cooldownPeriod, KEDA takes the Deployment back to zero. Check the events KEDA recorded:

kubectl describe scaledobject worker-scaler -n keda-demo | tail -n 5
Events:
  Type    Reason                      Age   From           Message
  ----    ------                      ----  ----           -------
  Normal  KEDAScaleTargetActivated    8m    keda-operator  Scaled apps/v1.Deployment keda-demo/worker from 0 to 1, triggered by rabbitMQScaler
  Normal  KEDAScaleTargetDeactivated  1m    keda-operator  Deactivated apps/v1.Deployment keda-demo/worker from 1 to 0

Step 5 - Adding a cron trigger for predictable traffic

A ScaledObject can have several triggers; the HPA uses whichever asks for the most replicas. A cron trigger keeps a minimum number of replicas during a time window, which is useful to pre-warm workers before a known daily peak.

Edit worker-scaler.yaml and add a second entry under triggers:

    - type: cron
      metadata:
        timezone: Europe/Madrid
        start: "0 8 * * 1-5"
        end: "0 20 * * 1-5"
        desiredReplicas: "3"

timezone takes an IANA name, and start/end use standard five-field cron syntax. With this trigger the worker runs at least 3 replicas from 08:00 to 20:00 on weekdays, and can still scale up to 10 if the queue grows. Outside that window it falls back to the queue trigger alone and can scale to zero.

Apply the change and check that the ScaledObject is still ready:

kubectl apply -f worker-scaler.yaml
kubectl get scaledobject worker-scaler -n keda-demo

Other common scalers

The same pattern (ScaledObject plus TriggerAuthentication) applies to every scaler. The key metadata for three frequently used ones:

ScalerKey metadataScales on
kafkabootstrapServers, consumerGroup, topic, lagThresholdConsumer group lag per partition
prometheusserverAddress, query, threshold, activationThresholdValue returned by a PromQL query
crontimezone, start, end, desiredReplicasTime windows

For example, a Prometheus trigger that targets 100 requests per second per replica:

    - type: prometheus
      metadata:
        serverAddress: http://prometheus.monitoring.svc:9090
        query: sum(rate(http_requests_total{job="api"}[2m]))
        threshold: "100"
        activationThreshold: "10"

Note that with Kafka, KEDA never runs more replicas than there are partitions in the topic by default, because extra consumers in the group would sit idle.

For batch work where each message should be handled by a separate pod that exits when done, use a ScaledJob instead of a ScaledObject. It creates Kubernetes Jobs instead of changing a Deployment's replica count.

Troubleshooting

READY is False or Unknown. KEDA cannot query the event source. Describe the ScaledObject and read the operator logs:

kubectl describe scaledobject worker-scaler -n keda-demo
kubectl logs -n keda deploy/keda-operator --tail=50

Typical causes are a wrong host or password in the Secret, a queue that does not exist, or a network policy blocking traffic from the keda namespace to the event source.

The HPA shows <unknown> targets while replicas are above zero. The metrics API server cannot serve the metric. Check kubectl get apiservice v1beta1.external.metrics.k8s.io and the logs of keda-operator-metrics-apiserver.

The workload never scales to zero. Confirm minReplicaCount is 0, that no other trigger (such as an active cron window) is keeping it up, and wait for the full cooldownPeriod after the queue is empty.

Cleaning up

Remove the demo resources when you are done:

kubectl delete namespace keda-demo

Conclusion

You installed KEDA, scaled a Deployment from zero to ten replicas based on RabbitMQ queue length, and added a cron trigger for business hours. Next, replace the demo worker with your real consumer, tune value so each replica handles a realistic share of the backlog, and consider ScaledJob for long-running, one-message-per-pod batch jobs. The KEDA scalers catalog at keda.sh lists the metadata for every supported event source.