Argo Workflows is a Kubernetes-native workflow engine: each step of a workflow runs as a pod, and the dependencies between steps are described in YAML as a sequence or a directed acyclic graph (DAG). It is widely used for batch jobs, data pipelines, machine learning and CI tasks. In this tutorial you will install Argo Workflows and its CLI, run a first workflow, build a DAG with parallel tasks, reuse logic with WorkflowTemplates, pass artifacts through S3-compatible storage and schedule a nightly job with a CronWorkflow.
Prerequisites
To follow this tutorial you need:
- A Kubernetes cluster running version 1.30 or later, for example a managed Kubernetes cluster on CubePath, with at least 2 vCPU and 4 GB of RAM free for the controller and workflow pods.
kubectlconfigured with cluster-admin access to that cluster.- A Linux workstation (Ubuntu 24.04 in the examples) with
curl. - For the artifacts step: an S3-compatible bucket and an access key/secret key pair for it.
All commands run from your workstation. Every workflow in this guide uses public images (alpine, busybox), so you can run them as they are.
Step 1 - Installing Argo Workflows
Argo Workflows publishes an install.yaml manifest with each release that installs the workflow controller, the Argo Server (API and web UI) and the Custom Resource Definitions into the argo namespace.
Set the version you want to install. Check the latest release on the releases page; this guide was tested with v4.1.4:
ARGO_WORKFLOWS_VERSION="v4.1.4"
Create the namespace and apply the manifest. The --server-side flag is required because the full CRDs are too large for a client-side apply:
kubectl create namespace argo
kubectl apply --server-side -n argo -f "https://github.com/argoproj/argo-workflows/releases/download/${ARGO_WORKFLOWS_VERSION}/install.yaml"
Wait until both deployments are available:
kubectl -n argo rollout status deployment/workflow-controller
kubectl -n argo rollout status deployment/argo-server
deployment "workflow-controller" successfully rolled out
deployment "argo-server" successfully rolled out
Step 2 - Installing the argo CLI
The argo CLI submits and inspects workflows. By default it talks directly to the Kubernetes API using your kubeconfig, so it works without exposing the Argo Server.
Download the binary for your architecture (use argo-linux-arm64.gz on ARM), make it executable and move it into your PATH:
curl -fsSLO "https://github.com/argoproj/argo-workflows/releases/download/${ARGO_WORKFLOWS_VERSION}/argo-linux-amd64.gz"
gunzip argo-linux-amd64.gz
chmod +x argo-linux-amd64
sudo mv argo-linux-amd64 /usr/local/bin/argo
Verify the installation:
argo version --short
argo: v4.1.4
Step 3 - Creating a service account for workflows
Every pod in a workflow runs with a Kubernetes ServiceAccount. The executor inside those pods needs permission to report task results back to the controller, and the project recommends a dedicated account instead of default.
Create a file with the ServiceAccount, the minimal Role and the RoleBinding:
nano argo-workflow-sa.yaml
apiVersion: v1
kind: ServiceAccount
metadata:
name: argo-workflow
namespace: argo
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: argo-executor
namespace: argo
rules:
- apiGroups: ["argoproj.io"]
resources: ["workflowtaskresults"]
verbs: ["create", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: argo-workflow-executor
namespace: argo
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: argo-executor
subjects:
- kind: ServiceAccount
name: argo-workflow
namespace: argo
Apply it:
kubectl apply -f argo-workflow-sa.yaml
If a workflow later needs to create Kubernetes resources itself (for example to deploy something), add those permissions to this Role, not to default.
Step 4 - Running your first workflow
A Workflow has an entrypoint and a list of templates. The simplest template runs one container.
Create hello-workflow.yaml:
nano hello-workflow.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
generateName: hello-
namespace: argo
spec:
serviceAccountName: argo-workflow
entrypoint: hello
arguments:
parameters:
- name: message
value: "Hello, Argo Workflows!"
templates:
- name: hello
inputs:
parameters:
- name: message
container:
image: busybox:1.36
command: [echo]
args: ["{{inputs.parameters.message}}"]
resources:
requests:
cpu: 50m
memory: 32Mi
generateName makes Kubernetes append a random suffix, so you can submit the same file many times. Submit it and watch it run:
argo submit -n argo --watch hello-workflow.yaml
When it finishes, the status block ends with:
Name: hello-x7k2p
Namespace: argo
ServiceAccount: argo-workflow
Status: Succeeded
...
STEP TEMPLATE PODNAME DURATION MESSAGE
✔ hello-x7k2p hello hello-x7k2p 4s
Print the logs of the latest workflow:
argo logs -n argo @latest
hello-x7k2p: Hello, Argo Workflows!
Parameters can be overridden at submit time without editing the file:
argo submit -n argo --watch hello-workflow.yaml -p message="Hello from the CLI"
Step 5 - Building a DAG with parallel tasks
A DAG template lists tasks and their dependencies. Tasks without pending dependencies run in parallel. In this pipeline, extract runs first, transform and validate run at the same time, and load waits for both. extract also produces an output parameter that transform consumes.
Create dag-pipeline.yaml:
nano dag-pipeline.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
generateName: etl-
namespace: argo
spec:
serviceAccountName: argo-workflow
entrypoint: pipeline
templates:
- name: pipeline
dag:
tasks:
- name: extract
template: extract
- name: transform
template: step
dependencies: [extract]
arguments:
parameters:
- name: message
value: "transforming {{tasks.extract.outputs.parameters.rows}} rows"
- name: validate
template: step
dependencies: [extract]
arguments:
parameters:
- name: message
value: "validating schema"
- name: load
template: step
dependencies: [transform, validate]
arguments:
parameters:
- name: message
value: "loading into the warehouse"
- name: extract
script:
image: alpine:3.20
command: [sh]
source: |
echo 42 > /tmp/rows
echo "extracted 42 rows"
outputs:
parameters:
- name: rows
valueFrom:
path: /tmp/rows
- name: step
inputs:
parameters:
- name: message
container:
image: busybox:1.36
command: [sh, -c]
args: ["echo {{inputs.parameters.message}}; sleep 5"]
Submit it:
argo submit -n argo --watch dag-pipeline.yaml
In the final tree you can see that transform and validate started at the same time:
STEP TEMPLATE PODNAME DURATION
✔ etl-9vq4d pipeline
├─✔ extract extract etl-9vq4d-extract-1520483145 4s
├─✔ transform step etl-9vq4d-step-2718296042 9s
├─✔ validate step etl-9vq4d-step-1016387190 9s
└─✔ load step etl-9vq4d-step-3920284717 8s
Fanning out over a list
To run the same task for each element of a list, use withItems. Argo creates one pod per item and runs them in parallel. Add this task to a DAG to process three files:
- name: process-file
template: step
arguments:
parameters:
- name: message
value: "processing {{item}}"
withItems:
- report-jan.csv
- report-feb.csv
- report-mar.csv
A task that depends on process-file only starts when all three pods have succeeded, which gives you fan-out and fan-in with no extra code.
Step 6 - Reusing logic with WorkflowTemplates
A WorkflowTemplate is a workflow definition stored in the cluster that other workflows reference by name. Use it for steps shared by many pipelines, so a fix is made once.
Create notify-template.yaml with a reusable template:
nano notify-template.yaml
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata:
name: common-steps
namespace: argo
spec:
templates:
- name: notify
inputs:
parameters:
- name: text
container:
image: busybox:1.36
command: [echo]
args: ["NOTIFY: {{inputs.parameters.text}}"]
Register it in the cluster:
argo template create -n argo notify-template.yaml
argo template list -n argo
NAME
common-steps
Reference it from any workflow with templateRef:
nano use-template.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
generateName: use-template-
namespace: argo
spec:
serviceAccountName: argo-workflow
entrypoint: main
templates:
- name: main
steps:
- - name: announce
templateRef:
name: common-steps
template: notify
arguments:
parameters:
- name: text
value: "nightly build finished"
argo submit -n argo --watch use-template.yaml
argo logs -n argo @latest
use-template-5r2wn-notify-...: NOTIFY: nightly build finished
Step 7 - Passing artifacts through S3-compatible storage
Parameters are for short strings. Files (reports, datasets, build outputs) are passed as artifacts, which Argo uploads to an artifact repository after a step and downloads before the next one. Any S3-compatible object storage works.
First store the bucket credentials in a Secret. Replace the placeholders with your real keys:
kubectl -n argo create secret generic s3-credentials \
--from-literal=accessKey='your_access_key' \
--from-literal=secretKey='your_secret_key'
Then declare the default artifact repository for the namespace with a ConfigMap named artifact-repositories. Replace your_s3_endpoint (host name only, without https://) and your_bucket:
nano artifact-repositories.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: artifact-repositories
namespace: argo
annotations:
workflows.argoproj.io/default-artifact-repository: default-v1
data:
default-v1: |
s3:
endpoint: your_s3_endpoint
bucket: your_bucket
insecure: false
accessKeySecret:
name: s3-credentials
key: accessKey
secretKeySecret:
name: s3-credentials
key: secretKey
kubectl apply -f artifact-repositories.yaml
Now create a workflow where one step writes a JSON file and the next reads it:
nano artifacts.yaml
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
generateName: artifacts-
namespace: argo
spec:
serviceAccountName: argo-workflow
entrypoint: main
templates:
- name: main
steps:
- - name: generate
template: generate-report
- - name: consume
template: print-report
arguments:
artifacts:
- name: report
from: "{{steps.generate.outputs.artifacts.report}}"
- name: generate-report
container:
image: alpine:3.20
command: [sh, -c]
args: ['echo "{\"status\":\"ok\",\"count\":42}" > /tmp/report.json']
outputs:
artifacts:
- name: report
path: /tmp/report.json
- name: print-report
inputs:
artifacts:
- name: report
path: /tmp/input/report.json
container:
image: alpine:3.20
command: [cat, /tmp/input/report.json]
argo submit -n argo --watch artifacts.yaml
argo logs -n argo @latest
artifacts-p8m3c-print-report-...: {"status":"ok","count":42}
The object is also visible in your bucket under a path that starts with the workflow name, compressed as report.tgz.
Step 8 - Scheduling jobs with CronWorkflows
A CronWorkflow creates a Workflow on a schedule, like a Kubernetes CronJob but with the full workflow model. Since v4.0 the schedule is a list in the schedules field (the old single schedule field has been removed).
nano nightly-report.yaml
apiVersion: argoproj.io/v1alpha1
kind: CronWorkflow
metadata:
name: nightly-report
namespace: argo
spec:
schedules:
- "0 2 * * *"
timezone: "Europe/Madrid"
concurrencyPolicy: Forbid
startingDeadlineSeconds: 300
successfulJobsHistoryLimit: 5
failedJobsHistoryLimit: 3
workflowSpec:
serviceAccountName: argo-workflow
entrypoint: report
templates:
- name: report
container:
image: busybox:1.36
command: [sh, -c]
args: ["echo generating report for {{workflow.scheduledTime}}"]
concurrencyPolicy: Forbidskips a run if the previous one is still running.startingDeadlineSeconds: 300still runs a missed schedule if the controller was down for less than 5 minutes.- The history limits keep the last 5 successful and 3 failed workflows.
Create it and check its next run:
argo cron create -n argo nightly-report.yaml
argo cron list -n argo
NAME AGE LAST RUN NEXT RUN SCHEDULES TIMEZONE SUSPENDED
nightly-report 5s N/A 14h 0 2 * * * Europe/Madrid false
To test it without waiting for 2:00, submit a workflow from the CronWorkflow right now:
argo submit -n argo --from cronwf/nightly-report --watch
You can pause and resume the schedule with argo cron suspend nightly-report -n argo and argo cron resume nightly-report -n argo.
Step 9 - Accessing the web UI
The Argo Server provides a web UI to browse workflows, logs and artifacts. Its default authentication mode is client: you log in with a Kubernetes bearer token, and you see exactly what that token is allowed to see.
Create a ServiceAccount for UI access with the read-only role that the installation ships:
kubectl -n argo create serviceaccount argo-ui
kubectl -n argo create rolebinding argo-ui-view \
--clusterrole=argo-aggregate-to-view \
--serviceaccount=argo:argo-ui
Generate a short-lived token:
echo "Bearer $(kubectl -n argo create token argo-ui --duration=8h)"
Forward the Argo Server port to your workstation. The server is not exposed outside the cluster:
kubectl -n argo port-forward service/argo-server 2746:2746
Open https://localhost:2746 (note https; the certificate is self-signed, so accept the browser warning), paste the whole Bearer ... string in the token box of the login page and select the argo namespace. You should see the workflows you ran in the previous steps.
Use argo-aggregate-to-edit instead of argo-aggregate-to-view if the account must also submit or delete workflows from the UI.
Troubleshooting
The workflow stays in Pending. The pods cannot be scheduled. Look at the pod events:
kubectl -n argo get pods -l workflows.argoproj.io/workflow=WORKFLOW_NAME
kubectl -n argo describe pod POD_NAME
Insufficient cpu or Insufficient memory means the resource requests do not fit on any node: lower them or add capacity.
Error mentioning workflowtaskresults.argoproj.io is forbidden. The workflow ran with a ServiceAccount that lacks the executor Role from Step 3. Check that spec.serviceAccountName is set to argo-workflow in the workflow (or in workflowSpec for CronWorkflows).
Artifacts fail to upload. Read the logs of the wait container of the step, which is the one that uploads artifacts:
kubectl -n argo logs POD_NAME -c wait
Errors like Access Denied or NoSuchBucket point to the Secret keys or the bucket name. Confirm the ConfigMap has the workflows.argoproj.io/default-artifact-repository annotation.
A DAG task never starts. A task only runs when all its dependencies succeeded. Find the failed node and read its logs:
argo get -n argo WORKFLOW_NAME
argo logs -n argo WORKFLOW_NAME --follow
Conclusion
You installed Argo Workflows and its CLI, created a least-privilege ServiceAccount for workflow pods, and ran a sequential workflow, a parallel DAG, a shared WorkflowTemplate, an artifact-passing pipeline backed by S3 storage and a scheduled CronWorkflow. The same building blocks scale from nightly maintenance jobs to multi-stage data and CI pipelines.
Next steps:
- Add retries with
retryStrategyon steps that call external services. - Trigger workflows from Git pushes or webhooks with Argo Events.
- Enable workflow archiving with PostgreSQL or MySQL to keep history after workflows are garbage collected.
