A Kubernetes Job runs one or more pods until a task completes successfully, and a CronJob creates Jobs on a schedule, like cron on a Linux server. In this tutorial you will run a simple Job, control its retries and time limits, split work across parallel pods with an Indexed Job, and schedule a nightly PostgreSQL backup with a CronJob. All commands run with kubectl from an Ubuntu 24.04 workstation or any machine with access to the cluster.

Prerequisites

To follow this guide you need:

  • A Kubernetes cluster (1.30 or newer), for example on CubePath VPS, and kubectl configured to use it. Check with kubectl get nodes.
  • Permission to create namespaces, Jobs, CronJobs, Secrets and PVCs.
  • For Step 6 only: a PostgreSQL server reachable from the cluster and a StorageClass that can provision volumes.

Create a namespace for the examples:

kubectl create namespace batch

Step 1 - Running a simple Job

A Job wraps a pod template. Unlike a Deployment, the pod is expected to exit, and the Job is complete once the required number of pods has exited with code 0.

Create the manifest:

nano hello-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: hello
  namespace: batch
spec:
  backoffLimit: 3
  ttlSecondsAfterFinished: 3600
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: hello
          image: busybox:1.36
          command: ["sh", "-c", "echo Processing started; sleep 5; echo Done"]

The pod template of a Job must use restartPolicy: Never or OnFailure; the default Always is rejected. ttlSecondsAfterFinished deletes the Job and its pods one hour after it finishes, so completed Jobs do not pile up.

Run it and wait for completion:

kubectl apply -f hello-job.yaml
kubectl -n batch wait --for=condition=complete job/hello --timeout=120s
job.batch/hello created
job.batch/hello condition met

Check the status and the output. Pods created by a Job carry the label job-name:

kubectl -n batch get job hello
kubectl -n batch logs job/hello
NAME    STATUS     COMPLETIONS   DURATION   AGE
hello   Complete   1/1           9s         15s
Processing started
Done

For quick one-off tasks you can also create a Job without a file: kubectl -n batch create job hello-cli --image=busybox:1.36 -- echo hi.

Step 2 - Controlling retries and time limits

Batch tasks fail, and a Job needs clear limits so that a broken task does not retry forever. The fields that control this are:

FieldEffect
backoffLimitNumber of retries before the Job is marked failed (default 6). Retries wait 10 s, 20 s, 40 s and so on, up to 6 minutes.
activeDeadlineSecondsMaximum total runtime of the Job. When reached, all its pods are terminated and the Job fails.
restartPolicyNever creates a new pod for each retry, keeping the failed pods for inspection. OnFailure restarts the container in the same pod.

To see the behavior, run a Job that always fails:

nano failing-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: failing
  namespace: batch
spec:
  backoffLimit: 2
  activeDeadlineSeconds: 300
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: task
          image: busybox:1.36
          command: ["sh", "-c", "echo Connecting to API; exit 1"]
kubectl apply -f failing-job.yaml
kubectl -n batch get pods -l job-name=failing --watch

After three pods (the first attempt plus two retries) end in Error, press Ctrl+C and inspect the Job:

kubectl -n batch describe job failing | tail -n 5
  Type     Reason                Age   From            Message
  ----     ------                ----  ----            -------
  Normal   SuccessfulCreate      60s   job-controller  Created pod: failing-2xk9p
  Normal   SuccessfulCreate      40s   job-controller  Created pod: failing-7bq4d
  Warning  BackoffLimitExceeded  5s    job-controller  Job has reached the specified backoff limit

Because the pods were kept, you can read the logs of any attempt with kubectl -n batch logs <pod_name>.

Sometimes retrying is pointless, for example when the input is invalid. A podFailurePolicy can fail the Job immediately on a specific exit code, and ignore pod evictions caused by node drains so they do not count against backoffLimit. It requires restartPolicy: Never:

spec:
  backoffLimit: 3
  podFailurePolicy:
    rules:
      - action: FailJob
        onExitCodes:
          containerName: task
          operator: In
          values: [42]
      - action: Ignore
        onPodConditions:
          - type: DisruptionTarget

Delete the failing Job before continuing:

kubectl -n batch delete job failing

Step 3 - Processing work in parallel with an Indexed Job

When a task can be split into independent pieces, such as processing 10 files or 10 date ranges, run several pods at once. Two fields control this:

  • completions: how many pods must succeed in total.
  • parallelism: how many pods run at the same time.

With completionMode: Indexed, every pod receives a unique index from 0 to completions - 1 in the JOB_COMPLETION_INDEX environment variable, so each pod knows which piece to handle without an external queue.

nano indexed-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: process-chunks
  namespace: batch
spec:
  completions: 10
  parallelism: 3
  completionMode: Indexed
  backoffLimit: 4
  ttlSecondsAfterFinished: 3600
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: worker
          image: busybox:1.36
          command: ["sh", "-c", "echo Processing chunk $JOB_COMPLETION_INDEX of 10; sleep 5"]

Apply it and watch the progress:

kubectl apply -f indexed-job.yaml
kubectl -n batch get job process-chunks --watch
NAME             STATUS    COMPLETIONS   DURATION   AGE
process-chunks   Running   0/10          3s         3s
process-chunks   Running   3/10          10s        10s
process-chunks   Running   6/10          17s        17s
process-chunks   Complete  10/10         26s        26s

Each batch of three pods runs together. Check that every index was processed exactly once:

kubectl -n batch logs -l job-name=process-chunks --tail=1 | sort -V
Processing chunk 0 of 10
Processing chunk 1 of 10
...
Processing chunk 9 of 10

In a real worker, use the index to select the input, for example input-$JOB_COMPLETION_INDEX.csv or rows index * 1000 to index * 1000 + 999.

Step 4 - Scheduling work with a CronJob

A CronJob contains a Job template and a schedule in standard five-field cron syntax:

FieldAllowed values
Minute0-59
Hour0-23
Day of month1-31
Month1-12
Day of week0-6 (0 is Sunday)

Some common schedules:

ScheduleMeaning
*/15 * * * *Every 15 minutes
0 */4 * * *Every 4 hours, on the hour
30 2 * * *Every day at 02:30
0 9 * * 1-5Weekdays at 09:00
0 3 1 * *The first day of every month at 03:00

Create a CronJob that runs every two minutes, so you can observe it:

nano heartbeat-cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: heartbeat
  namespace: batch
spec:
  schedule: "*/2 * * * *"
  timeZone: "Etc/UTC"
  concurrencyPolicy: Forbid
  startingDeadlineSeconds: 120
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 1
  jobTemplate:
    spec:
      backoffLimit: 1
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: heartbeat
              image: busybox:1.36
              command: ["sh", "-c", "date; echo heartbeat ok"]

The CronJob-specific fields:

  • timeZone: an IANA time zone name such as Europe/Madrid or America/New_York. Without it, the schedule uses the time zone of the kube-controller-manager, which is usually UTC.
  • concurrencyPolicy: Allow (default) lets runs overlap, Forbid skips a run while the previous one is still active, and Replace stops the running Job and starts the new one. Use Forbid for anything that must not run twice at once, like backups.
  • startingDeadlineSeconds: if a run could not start on time (for example, because the controller was down), it is skipped when it is later than this many seconds.
  • successfulJobsHistoryLimit and failedJobsHistoryLimit: how many finished Jobs to keep for inspection.

Apply it and test it right away instead of waiting for the schedule. kubectl create job --from creates a Job from the CronJob's template:

kubectl apply -f heartbeat-cronjob.yaml
kubectl -n batch create job heartbeat-manual --from=cronjob/heartbeat
kubectl -n batch wait --for=condition=complete job/heartbeat-manual --timeout=60s
kubectl -n batch logs job/heartbeat-manual
Thu Sep 25 11:02:14 UTC 2026
heartbeat ok

After a few minutes, check that scheduled runs are happening:

kubectl -n batch get cronjob heartbeat
kubectl -n batch get jobs
NAME        SCHEDULE      TIMEZONE   SUSPEND   ACTIVE   LAST SCHEDULE   AGE
heartbeat   */2 * * * *   Etc/UTC    False     0        40s             5m

The Jobs list shows entries named heartbeat-<number>, never more than three successful ones because of the history limit.

Step 5 - Suspending and resuming a CronJob

To pause a CronJob during maintenance without deleting it, set suspend to true. Jobs that are already running are not affected:

kubectl -n batch patch cronjob heartbeat -p '{"spec":{"suspend":true}}'
kubectl -n batch get cronjob heartbeat -o jsonpath='{.spec.suspend}{"\n"}'
true

Resume it with "suspend":false. Runs missed while suspended are not executed afterwards unless they are still within startingDeadlineSeconds. When you are done with the example, delete it:

kubectl -n batch delete cronjob heartbeat

Deleting a CronJob also deletes the Jobs and pods it created.

Step 6 - Scheduling a nightly database backup

A realistic CronJob: dump a PostgreSQL database every night at 02:30 Madrid time to a persistent volume and keep the last seven dumps. Store the database password in a Secret instead of in the manifest, replacing your_db_password:

kubectl -n batch create secret generic pg-backup --from-literal=PGPASSWORD='your_db_password'

Create the manifest, replacing your_db_host, your_db_user and your_db_name:

nano pg-backup-cronjob.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: pg-backups
  namespace: batch
spec:
  accessModes: ["ReadWriteOnce"]
  resources:
    requests:
      storage: 10Gi
---
apiVersion: batch/v1
kind: CronJob
metadata:
  name: pg-backup
  namespace: batch
spec:
  schedule: "30 2 * * *"
  timeZone: "Europe/Madrid"
  concurrencyPolicy: Forbid
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 3
  jobTemplate:
    spec:
      backoffLimit: 2
      activeDeadlineSeconds: 3600
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: pg-dump
              image: postgres:16
              env:
                - name: PGHOST
                  value: your_db_host
                - name: PGUSER
                  value: your_db_user
                - name: PGDATABASE
                  value: your_db_name
                - name: PGPASSWORD
                  valueFrom:
                    secretKeyRef:
                      name: pg-backup
                      key: PGPASSWORD
              command:
                - sh
                - -c
                - |
                  set -eu
                  file="/backups/${PGDATABASE}-$(date +%Y%m%d-%H%M%S).dump"
                  pg_dump --format=custom --file="$file"
                  echo "Backup written to $file"
                  find /backups -name '*.dump' -mtime +7 -delete
              volumeMounts:
                - name: backups
                  mountPath: /backups
          volumes:
            - name: backups
              persistentVolumeClaim:
                claimName: pg-backups

pg_dump reads the connection settings from the standard PG* environment variables. set -eu makes the script stop on the first error, so a failed dump fails the Job and old dumps are only cleaned up after a successful one. Use a postgres image with the same or a newer major version than your server.

Apply it and trigger a first run to validate the configuration:

kubectl apply -f pg-backup-cronjob.yaml
kubectl -n batch create job pg-backup-test --from=cronjob/pg-backup
kubectl -n batch wait --for=condition=complete job/pg-backup-test --timeout=600s
kubectl -n batch logs job/pg-backup-test
Backup written to /backups/your_db_name-20260925-110815.dump

Keep in mind that this volume lives in the same cluster as the job. Copy the dumps off the cluster as well, for example to object storage, so a cluster failure does not take the backups with it.

Troubleshooting

The Job pod stays Pending. The scheduler cannot place it, usually because of insufficient CPU or memory, or an unbound PVC. kubectl -n batch describe pod <pod_name> shows the reason in the events.

The CronJob never runs. Check kubectl -n batch describe cronjob <name> for events. Common causes are suspend: true, an invalid timeZone name, or concurrencyPolicy: Forbid combined with a previous Job that is still running or stuck.

The Job is marked failed with DeadlineExceeded. The total runtime reached activeDeadlineSeconds. Increase the limit if the task is legitimately slow, or check the pod logs to find where it hangs.

Too many finished Jobs and pods. Set ttlSecondsAfterFinished on standalone Jobs and history limits on CronJobs. To remove existing completed Jobs in one go, run kubectl -n batch delete jobs --field-selector status.successful=1.

Conclusion

You ran one-off and parallel Jobs, controlled their retries and deadlines, and scheduled recurring work with CronJobs, including a nightly database backup you can test on demand. Next, add CPU and memory requests to your job templates so they do not starve other workloads, alert on failed Jobs with your monitoring stack, and ship backup files off the cluster.