Argo Workflows is a Kubernetes-native workflow engine: each step of a workflow runs as a pod, and the dependencies between steps are described in YAML as a sequence or a directed acyclic graph (DAG). It is widely used for batch jobs, data pipelines, machine learning and CI tasks. In this tutorial you will install Argo Workflows and its CLI, run a first workflow, build a DAG with parallel tasks, reuse logic with WorkflowTemplates, pass artifacts through S3-compatible storage and schedule a nightly job with a CronWorkflow.

Prerequisites

To follow this tutorial you need:

  • A Kubernetes cluster running version 1.30 or later, for example a managed Kubernetes cluster on CubePath, with at least 2 vCPU and 4 GB of RAM free for the controller and workflow pods.
  • kubectl configured with cluster-admin access to that cluster.
  • A Linux workstation (Ubuntu 24.04 in the examples) with curl.
  • For the artifacts step: an S3-compatible bucket and an access key/secret key pair for it.

All commands run from your workstation. Every workflow in this guide uses public images (alpine, busybox), so you can run them as they are.

Step 1 - Installing Argo Workflows

Argo Workflows publishes an install.yaml manifest with each release that installs the workflow controller, the Argo Server (API and web UI) and the Custom Resource Definitions into the argo namespace.

Set the version you want to install. Check the latest release on the releases page; this guide was tested with v4.1.4:

ARGO_WORKFLOWS_VERSION="v4.1.4"

Create the namespace and apply the manifest. The --server-side flag is required because the full CRDs are too large for a client-side apply:

kubectl create namespace argo
kubectl apply --server-side -n argo -f "https://github.com/argoproj/argo-workflows/releases/download/${ARGO_WORKFLOWS_VERSION}/install.yaml"

Wait until both deployments are available:

kubectl -n argo rollout status deployment/workflow-controller
kubectl -n argo rollout status deployment/argo-server
deployment "workflow-controller" successfully rolled out
deployment "argo-server" successfully rolled out

Step 2 - Installing the argo CLI

The argo CLI submits and inspects workflows. By default it talks directly to the Kubernetes API using your kubeconfig, so it works without exposing the Argo Server.

Download the binary for your architecture (use argo-linux-arm64.gz on ARM), make it executable and move it into your PATH:

curl -fsSLO "https://github.com/argoproj/argo-workflows/releases/download/${ARGO_WORKFLOWS_VERSION}/argo-linux-amd64.gz"
gunzip argo-linux-amd64.gz
chmod +x argo-linux-amd64
sudo mv argo-linux-amd64 /usr/local/bin/argo

Verify the installation:

argo version --short
argo: v4.1.4

Step 3 - Creating a service account for workflows

Every pod in a workflow runs with a Kubernetes ServiceAccount. The executor inside those pods needs permission to report task results back to the controller, and the project recommends a dedicated account instead of default.

Create a file with the ServiceAccount, the minimal Role and the RoleBinding:

nano argo-workflow-sa.yaml
apiVersion: v1
kind: ServiceAccount
metadata:
  name: argo-workflow
  namespace: argo
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: argo-executor
  namespace: argo
rules:
  - apiGroups: ["argoproj.io"]
    resources: ["workflowtaskresults"]
    verbs: ["create", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: argo-workflow-executor
  namespace: argo
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: argo-executor
subjects:
  - kind: ServiceAccount
    name: argo-workflow
    namespace: argo

Apply it:

kubectl apply -f argo-workflow-sa.yaml

If a workflow later needs to create Kubernetes resources itself (for example to deploy something), add those permissions to this Role, not to default.

Step 4 - Running your first workflow

A Workflow has an entrypoint and a list of templates. The simplest template runs one container.

Create hello-workflow.yaml:

nano hello-workflow.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
  generateName: hello-
  namespace: argo
spec:
  serviceAccountName: argo-workflow
  entrypoint: hello
  arguments:
    parameters:
      - name: message
        value: "Hello, Argo Workflows!"
  templates:
    - name: hello
      inputs:
        parameters:
          - name: message
      container:
        image: busybox:1.36
        command: [echo]
        args: ["{{inputs.parameters.message}}"]
        resources:
          requests:
            cpu: 50m
            memory: 32Mi

generateName makes Kubernetes append a random suffix, so you can submit the same file many times. Submit it and watch it run:

argo submit -n argo --watch hello-workflow.yaml

When it finishes, the status block ends with:

Name:                hello-x7k2p
Namespace:           argo
ServiceAccount:      argo-workflow
Status:              Succeeded
...
STEP            TEMPLATE  PODNAME          DURATION  MESSAGE
 ✔ hello-x7k2p  hello     hello-x7k2p      4s

Print the logs of the latest workflow:

argo logs -n argo @latest
hello-x7k2p: Hello, Argo Workflows!

Parameters can be overridden at submit time without editing the file:

argo submit -n argo --watch hello-workflow.yaml -p message="Hello from the CLI"

Step 5 - Building a DAG with parallel tasks

A DAG template lists tasks and their dependencies. Tasks without pending dependencies run in parallel. In this pipeline, extract runs first, transform and validate run at the same time, and load waits for both. extract also produces an output parameter that transform consumes.

Create dag-pipeline.yaml:

nano dag-pipeline.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
  generateName: etl-
  namespace: argo
spec:
  serviceAccountName: argo-workflow
  entrypoint: pipeline
  templates:
    - name: pipeline
      dag:
        tasks:
          - name: extract
            template: extract
          - name: transform
            template: step
            dependencies: [extract]
            arguments:
              parameters:
                - name: message
                  value: "transforming {{tasks.extract.outputs.parameters.rows}} rows"
          - name: validate
            template: step
            dependencies: [extract]
            arguments:
              parameters:
                - name: message
                  value: "validating schema"
          - name: load
            template: step
            dependencies: [transform, validate]
            arguments:
              parameters:
                - name: message
                  value: "loading into the warehouse"

    - name: extract
      script:
        image: alpine:3.20
        command: [sh]
        source: |
          echo 42 > /tmp/rows
          echo "extracted 42 rows"
      outputs:
        parameters:
          - name: rows
            valueFrom:
              path: /tmp/rows

    - name: step
      inputs:
        parameters:
          - name: message
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["echo {{inputs.parameters.message}}; sleep 5"]

Submit it:

argo submit -n argo --watch dag-pipeline.yaml

In the final tree you can see that transform and validate started at the same time:

STEP            TEMPLATE   PODNAME                        DURATION
 ✔ etl-9vq4d    pipeline
 ├─✔ extract    extract    etl-9vq4d-extract-1520483145   4s
 ├─✔ transform  step       etl-9vq4d-step-2718296042      9s
 ├─✔ validate   step       etl-9vq4d-step-1016387190      9s
 └─✔ load       step       etl-9vq4d-step-3920284717      8s

Fanning out over a list

To run the same task for each element of a list, use withItems. Argo creates one pod per item and runs them in parallel. Add this task to a DAG to process three files:

          - name: process-file
            template: step
            arguments:
              parameters:
                - name: message
                  value: "processing {{item}}"
            withItems:
              - report-jan.csv
              - report-feb.csv
              - report-mar.csv

A task that depends on process-file only starts when all three pods have succeeded, which gives you fan-out and fan-in with no extra code.

Step 6 - Reusing logic with WorkflowTemplates

A WorkflowTemplate is a workflow definition stored in the cluster that other workflows reference by name. Use it for steps shared by many pipelines, so a fix is made once.

Create notify-template.yaml with a reusable template:

nano notify-template.yaml
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata:
  name: common-steps
  namespace: argo
spec:
  templates:
    - name: notify
      inputs:
        parameters:
          - name: text
      container:
        image: busybox:1.36
        command: [echo]
        args: ["NOTIFY: {{inputs.parameters.text}}"]

Register it in the cluster:

argo template create -n argo notify-template.yaml
argo template list -n argo
NAME
common-steps

Reference it from any workflow with templateRef:

nano use-template.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
  generateName: use-template-
  namespace: argo
spec:
  serviceAccountName: argo-workflow
  entrypoint: main
  templates:
    - name: main
      steps:
        - - name: announce
            templateRef:
              name: common-steps
              template: notify
            arguments:
              parameters:
                - name: text
                  value: "nightly build finished"
argo submit -n argo --watch use-template.yaml
argo logs -n argo @latest
use-template-5r2wn-notify-...: NOTIFY: nightly build finished

Step 7 - Passing artifacts through S3-compatible storage

Parameters are for short strings. Files (reports, datasets, build outputs) are passed as artifacts, which Argo uploads to an artifact repository after a step and downloads before the next one. Any S3-compatible object storage works.

First store the bucket credentials in a Secret. Replace the placeholders with your real keys:

kubectl -n argo create secret generic s3-credentials \
  --from-literal=accessKey='your_access_key' \
  --from-literal=secretKey='your_secret_key'

Then declare the default artifact repository for the namespace with a ConfigMap named artifact-repositories. Replace your_s3_endpoint (host name only, without https://) and your_bucket:

nano artifact-repositories.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: artifact-repositories
  namespace: argo
  annotations:
    workflows.argoproj.io/default-artifact-repository: default-v1
data:
  default-v1: |
    s3:
      endpoint: your_s3_endpoint
      bucket: your_bucket
      insecure: false
      accessKeySecret:
        name: s3-credentials
        key: accessKey
      secretKeySecret:
        name: s3-credentials
        key: secretKey
kubectl apply -f artifact-repositories.yaml

Now create a workflow where one step writes a JSON file and the next reads it:

nano artifacts.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
  generateName: artifacts-
  namespace: argo
spec:
  serviceAccountName: argo-workflow
  entrypoint: main
  templates:
    - name: main
      steps:
        - - name: generate
            template: generate-report
        - - name: consume
            template: print-report
            arguments:
              artifacts:
                - name: report
                  from: "{{steps.generate.outputs.artifacts.report}}"

    - name: generate-report
      container:
        image: alpine:3.20
        command: [sh, -c]
        args: ['echo "{\"status\":\"ok\",\"count\":42}" > /tmp/report.json']
      outputs:
        artifacts:
          - name: report
            path: /tmp/report.json

    - name: print-report
      inputs:
        artifacts:
          - name: report
            path: /tmp/input/report.json
      container:
        image: alpine:3.20
        command: [cat, /tmp/input/report.json]
argo submit -n argo --watch artifacts.yaml
argo logs -n argo @latest
artifacts-p8m3c-print-report-...: {"status":"ok","count":42}

The object is also visible in your bucket under a path that starts with the workflow name, compressed as report.tgz.

Step 8 - Scheduling jobs with CronWorkflows

A CronWorkflow creates a Workflow on a schedule, like a Kubernetes CronJob but with the full workflow model. Since v4.0 the schedule is a list in the schedules field (the old single schedule field has been removed).

nano nightly-report.yaml
apiVersion: argoproj.io/v1alpha1
kind: CronWorkflow
metadata:
  name: nightly-report
  namespace: argo
spec:
  schedules:
    - "0 2 * * *"
  timezone: "Europe/Madrid"
  concurrencyPolicy: Forbid
  startingDeadlineSeconds: 300
  successfulJobsHistoryLimit: 5
  failedJobsHistoryLimit: 3
  workflowSpec:
    serviceAccountName: argo-workflow
    entrypoint: report
    templates:
      - name: report
        container:
          image: busybox:1.36
          command: [sh, -c]
          args: ["echo generating report for {{workflow.scheduledTime}}"]
  • concurrencyPolicy: Forbid skips a run if the previous one is still running.
  • startingDeadlineSeconds: 300 still runs a missed schedule if the controller was down for less than 5 minutes.
  • The history limits keep the last 5 successful and 3 failed workflows.

Create it and check its next run:

argo cron create -n argo nightly-report.yaml
argo cron list -n argo
NAME             AGE   LAST RUN   NEXT RUN   SCHEDULES   TIMEZONE        SUSPENDED
nightly-report   5s    N/A        14h        0 2 * * *   Europe/Madrid   false

To test it without waiting for 2:00, submit a workflow from the CronWorkflow right now:

argo submit -n argo --from cronwf/nightly-report --watch

You can pause and resume the schedule with argo cron suspend nightly-report -n argo and argo cron resume nightly-report -n argo.

Step 9 - Accessing the web UI

The Argo Server provides a web UI to browse workflows, logs and artifacts. Its default authentication mode is client: you log in with a Kubernetes bearer token, and you see exactly what that token is allowed to see.

Create a ServiceAccount for UI access with the read-only role that the installation ships:

kubectl -n argo create serviceaccount argo-ui
kubectl -n argo create rolebinding argo-ui-view \
  --clusterrole=argo-aggregate-to-view \
  --serviceaccount=argo:argo-ui

Generate a short-lived token:

echo "Bearer $(kubectl -n argo create token argo-ui --duration=8h)"

Forward the Argo Server port to your workstation. The server is not exposed outside the cluster:

kubectl -n argo port-forward service/argo-server 2746:2746

Open https://localhost:2746 (note https; the certificate is self-signed, so accept the browser warning), paste the whole Bearer ... string in the token box of the login page and select the argo namespace. You should see the workflows you ran in the previous steps.

Use argo-aggregate-to-edit instead of argo-aggregate-to-view if the account must also submit or delete workflows from the UI.

Troubleshooting

The workflow stays in Pending. The pods cannot be scheduled. Look at the pod events:

kubectl -n argo get pods -l workflows.argoproj.io/workflow=WORKFLOW_NAME
kubectl -n argo describe pod POD_NAME

Insufficient cpu or Insufficient memory means the resource requests do not fit on any node: lower them or add capacity.

Error mentioning workflowtaskresults.argoproj.io is forbidden. The workflow ran with a ServiceAccount that lacks the executor Role from Step 3. Check that spec.serviceAccountName is set to argo-workflow in the workflow (or in workflowSpec for CronWorkflows).

Artifacts fail to upload. Read the logs of the wait container of the step, which is the one that uploads artifacts:

kubectl -n argo logs POD_NAME -c wait

Errors like Access Denied or NoSuchBucket point to the Secret keys or the bucket name. Confirm the ConfigMap has the workflows.argoproj.io/default-artifact-repository annotation.

A DAG task never starts. A task only runs when all its dependencies succeeded. Find the failed node and read its logs:

argo get -n argo WORKFLOW_NAME
argo logs -n argo WORKFLOW_NAME --follow

Conclusion

You installed Argo Workflows and its CLI, created a least-privilege ServiceAccount for workflow pods, and ran a sequential workflow, a parallel DAG, a shared WorkflowTemplate, an artifact-passing pipeline backed by S3 storage and a scheduled CronWorkflow. The same building blocks scale from nightly maintenance jobs to multi-stage data and CI pipelines.

Next steps:

  • Add retries with retryStrategy on steps that call external services.
  • Trigger workflows from Git pushes or webhooks with Argo Events.
  • Enable workflow archiving with PostgreSQL or MySQL to keep history after workflows are garbage collected.