An Internal Developer Platform (IDP) is the set of tools, templates and APIs a platform team offers so that application developers can create services, request infrastructure and deploy without opening tickets. This guide explains the design patterns that make an IDP work on Kubernetes: a software catalog, golden path templates, a declarative self-service API, built-in guardrails and delivery metrics. Each pattern comes with a concrete, working example (Backstage, Crossplane, Argo CD and native Kubernetes admission policies) that you can adapt instead of a generic checklist.
Prerequisites
This is a design guide, not a single installation, but the examples assume:
- A Kubernetes cluster (1.30 or later) where you have cluster-admin access, for example a managed Kubernetes cluster on CubePath.
kubectlconfigured against that cluster.- A Git hosting service (GitHub in the examples) and a CI system.
- Optionally, a Backstage instance, Crossplane v2 and Argo CD if you want to run the examples as written.
- Familiarity with GitOps (the desired state lives in Git and a controller applies it).
The layers of an IDP
An IDP is not one product. It is a thin layer of opinionated interfaces on top of tools you probably already run. Thinking in layers helps you decide what to build and what to reuse:
| Layer | What developers get | Common tools |
|---|---|---|
| Portal and catalog | One place to find services, owners, docs and templates | Backstage, Port |
| Templates (golden paths) | A new repository that already builds, deploys and is monitored | Backstage Software Templates, cookiecutter |
| Infrastructure API | Databases, caches and buckets requested as YAML | Crossplane, Terraform with a controller |
| Delivery | Git push triggers build and deployment | Argo CD, Flux, Argo Workflows |
| Guardrails | Policies enforced automatically, not in code review | ValidatingAdmissionPolicy, Kyverno, OPA Gatekeeper |
| Observability | Dashboards and alerts wired in from day one | Prometheus, Grafana, Loki |
The platform team owns the integration between these layers and the defaults, not the individual applications. A useful rule: if a developer has to open a ticket for something that happens more than once a week, it is a candidate for the platform.
Pattern 1 - Treat the software catalog as the source of truth
Every other pattern depends on knowing which services exist, who owns them and what they depend on. In Backstage this is a catalog-info.yaml file at the root of each repository, which the catalog ingests.
A minimal but useful entry for a service looks like this:
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
name: order-service
description: Manages customer orders and fulfillment
annotations:
github.com/project-slug: my-org/order-service
backstage.io/techdocs-ref: dir:.
argocd/app-name: order-service-production
tags:
- java
- orders
links:
- url: https://grafana.example.com/d/order-service
title: Grafana dashboard
spec:
type: service
lifecycle: production
owner: group:order-team
system: ecommerce
dependsOn:
- component:default/payment-service
- resource:default/orders-database
providesApis:
- order-api
The fields that matter most are owner (so alerts and incidents reach a team, not a person), lifecycle (so you can find experimental services that should not receive production traffic) and dependsOn (so you can answer "what breaks if this database goes down"). Annotations connect the entry to plugins such as Argo CD or TechDocs.
Keep the file in the service repository rather than in a central list. When the owner changes, the change goes through the same pull request process as the code.
Pattern 2 - Golden paths as software templates
A golden path is the supported, pre-configured way to create a new service. It encodes your decisions (base image, CI pipeline, Helm chart, dashboards, catalog-info.yaml) so that a new service is production ready on its first commit. Developers can leave the path, but then they own the extra work.
A typical layout for templates in a platform repository:
platform/
templates/
go-service/
template.yaml # Backstage template definition
skeleton/ # Files copied into the new repository
catalog-info.yaml
Dockerfile
chart/
.github/workflows/ci.yml
static-frontend/
The Backstage Software Template below asks for a name, owner and description, renders the skeleton, creates the GitHub repository and registers it in the catalog:
apiVersion: scaffolder.backstage.io/v1beta3
kind: Template
metadata:
name: go-service
title: Go service
description: Go HTTP service with CI, Helm chart and catalog entry
tags:
- go
- recommended
spec:
owner: group:platform-team
type: service
parameters:
- title: Service information
required:
- name
- owner
properties:
name:
title: Name
type: string
pattern: '^[a-z][a-z0-9-]*$'
description: Lowercase letters, digits and hyphens
owner:
title: Owner
type: string
ui:field: OwnerPicker
ui:options:
catalogFilter:
kind: Group
description:
title: Description
type: string
steps:
- id: fetch
name: Render skeleton
action: fetch:template
input:
url: ./skeleton
values:
name: ${{ parameters.name }}
owner: ${{ parameters.owner }}
description: ${{ parameters.description }}
- id: publish
name: Create GitHub repository
action: publish:github
input:
repoUrl: github.com?owner=my-org&repo=${{ parameters.name }}
description: ${{ parameters.description }}
defaultBranch: main
repoVisibility: private
- id: register
name: Register in catalog
action: catalog:register
input:
repoContentsUrl: ${{ steps['publish'].output.repoContentsUrl }}
catalogInfoPath: /catalog-info.yaml
output:
links:
- title: Repository
url: ${{ steps['publish'].output.remoteUrl }}
- title: Open in catalog
icon: catalog
entityRef: ${{ steps['register'].output.entityRef }}
publish:github creates the repository and pushes the rendered files in one step, so you do not need a separate "create repository" action. Inside skeleton/, files use the same ${{ values.name }} syntax to insert the parameters.
Design rules that keep golden paths healthy:
- Start with one template for your most common service type. A template nobody uses is maintenance with no return.
- Version the skeleton like code and review changes to it. Every change affects all future services.
- Put the CI pipeline, Dockerfile and chart in the skeleton, but keep reusable logic (shared CI workflows, base Helm charts) in central repositories that the skeleton references. That way a fix reaches existing services too, not only new ones.
Pattern 3 - A declarative self-service infrastructure API
Developers should request a database the same way they request a Deployment: by committing a small manifest. The platform translates that request into the real resources, with sizes, backups and network rules chosen by the platform team.
With Crossplane v2 you define the API with a CompositeResourceDefinition (XRD). The following XRD creates a namespaced Database kind with a single size field:
apiVersion: apiextensions.crossplane.io/v2
kind: CompositeResourceDefinition
metadata:
name: databases.platform.example.com
spec:
group: platform.example.com
scope: Namespaced
names:
kind: Database
plural: databases
versions:
- name: v1alpha1
served: true
referenceable: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
size:
type: string
enum: [small, medium]
required:
- size
A Composition then maps that request to concrete resources. Crossplane v2 can compose any Kubernetes resource, so this example uses a CloudNativePG Cluster running in the same namespace. It requires the CloudNativePG operator and the function-patch-and-transform Crossplane function to be installed:
apiVersion: apiextensions.crossplane.io/v1
kind: Composition
metadata:
name: database-cnpg
spec:
compositeTypeRef:
apiVersion: platform.example.com/v1alpha1
kind: Database
mode: Pipeline
pipeline:
- step: render-postgres
functionRef:
name: function-patch-and-transform
input:
apiVersion: pt.fn.crossplane.io/v1beta1
kind: Resources
resources:
- name: postgres
base:
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
spec:
instances: 2
storage:
size: 10Gi
patches:
- type: FromCompositeFieldPath
fromFieldPath: spec.size
toFieldPath: spec.storage.size
transforms:
- type: map
map:
small: 10Gi
medium: 50Gi
The developer only writes this, in their own namespace:
apiVersion: platform.example.com/v1alpha1
kind: Database
metadata:
name: orders-db
namespace: order-team
spec:
size: small
Once applied, check that the request was reconciled:
kubectl get databases -n order-team
NAME SYNCED READY COMPOSITION AGE
orders-db True True database-cnpg 3m
CloudNativePG stores the application credentials in a Secret named after the cluster with an -app suffix, which the application mounts through secretKeyRef. The key design choice is that the developer chooses what (small), and the platform decides how (storage class, instance count, backups). You can later switch the Composition to a managed database on another provider without changing a single developer manifest.
Pattern 4 - Guardrails enforced by the platform
Guardrails replace review comments like "please set memory limits" with automatic checks. Enforce them at admission time so they apply to every path into the cluster: Argo CD, CI pipelines and manual kubectl.
Kubernetes has a built-in, generally available mechanism for this since 1.30: ValidatingAdmissionPolicy, which uses CEL expressions and needs no extra controller. This policy rejects Deployments whose containers do not set a memory limit:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: require-memory-limits
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: ["apps"]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["deployments"]
validations:
- expression: >-
object.spec.template.spec.containers.all(c,
has(c.resources) && has(c.resources.limits) &&
'memory' in c.resources.limits)
message: "Every container must set resources.limits.memory"
A policy does nothing until it is bound. Bind it only to namespaces labeled as tenant namespaces, so system components are not affected:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: require-memory-limits
spec:
policyName: require-memory-limits
validationActions: ["Deny"]
matchResources:
namespaceSelector:
matchLabels:
platform.example.com/tenant: "true"
Test it with a Deployment that has no limits:
kubectl label namespace order-team platform.example.com/tenant=true
kubectl create deployment test-nolimits --image=nginx -n order-team
error: failed to create deployment: deployments.apps "test-nolimits" is forbidden: ValidatingAdmissionPolicy 'require-memory-limits' with binding 'require-memory-limits' denied request: Every container must set resources.limits.memory
Tip: roll out a new policy with validationActions: ["Warn", "Audit"] first. Developers see a warning, you see which workloads would fail, and only then do you switch to Deny.
Combine admission policies with guardrails that live earlier in the path: the golden path template already ships a chart with limits, so most services pass the policy without anyone thinking about it.
Pattern 5 - Secrets through identity, not copied values
Developers should never copy credentials into CI variables or manifests. The platform gives each workload an identity (its Kubernetes ServiceAccount) and maps that identity to the secrets it may read. With HashiCorp Vault and its Kubernetes auth method enabled, a per-team policy and role look like this:
vault policy write order-team - <<'EOF'
path "secret/data/order-team/*" {
capabilities = ["read", "list"]
}
EOF
vault write auth/kubernetes/role/order-service \
bound_service_account_names=order-service \
bound_service_account_namespaces=order-team \
policies=order-team \
ttl=1h
Only pods running as the order-service ServiceAccount in the order-team namespace can log in with that role, and the token they receive expires after an hour. The platform automates creating these per-team policies (from the catalog owner field, for example) so developers only ask for "access to my team's secrets".
Pattern 6 - GitOps delivery with drift correction
With GitOps, deployments are pull requests and the cluster continuously converges to what is in Git. That makes the platform self-service by default: developers merge, Argo CD applies.
Enable automated sync with self-healing so manual changes in the cluster are reverted to the Git state:
argocd app set order-service-production --sync-policy automated --self-heal
To review what differs between Git and the live cluster before syncing:
argocd app diff order-service-production
The command prints nothing and exits with code 0 when there is no drift. Keep emergency changes possible, but make the path explicit: a hotfix is a revert or a small pull request, not a kubectl edit that the next sync will undo.
Pattern 7 - Measure the platform like a product
A platform that nobody adopts has no value, however good its architecture. Track two kinds of signals:
- Delivery metrics (DORA): deployment frequency, lead time for changes, change failure rate and time to restore service.
- Adoption and satisfaction: percentage of services created from a golden path, number of tickets to the platform team, and a short periodic developer survey.
Argo CD exposes a Prometheus counter of sync operations, which gives you a first approximation of deployment frequency per application over the last 7 days:
sum by (name) (increase(argocd_app_sync_total{phase="Succeeded"}[7d]))
Lead time needs timestamps from two systems (commit or merge time from Git, sync time from Argo CD), so it is usually computed from webhook events stored in a database rather than from a single metric.
Common pitfalls
The platform is too complex to adopt. Start with one golden path and one self-service resource. If less than half of new services use the path, find out why before adding features.
The platform team becomes a bottleneck. Anything that requires a platform engineer to act on a ticket should become a template, an API or a pull request that CI validates.
Templates drift from reality. Services created a year ago do not get fixes made to the skeleton later. Move shared logic into reusable CI workflows and base charts that the services reference, so updates propagate.
Guardrails block teams without explanation. Every policy needs a clear message and a link to documentation, and should start in Warn mode.
Conclusion
An effective internal developer platform is a small set of well-chosen interfaces: a catalog that knows who owns what, golden path templates, a declarative infrastructure API, automatic guardrails and GitOps delivery, all measured by how much they speed up developers. Build one pattern at a time and treat the platform as a product with its own users.
Next steps:
- Orchestrate CI and batch jobs inside the cluster with Argo Workflows.
- Give developers a fast inner loop against the shared cluster with Telepresence.
- Add Backstage TechDocs so every template ships with its own documentation.
