Every container in Kubernetes can declare how much CPU and memory it needs (requests) and the most it may use (limits). Requests drive scheduling, limits are enforced by the kernel, and together they decide which pods are evicted first when a node runs short of memory. In this tutorial you will set requests and limits on a workload, check its Quality of Service (QoS) class, reproduce an out-of-memory kill, apply namespace defaults with a LimitRange, cap a namespace with a ResourceQuota and use real usage data to right-size your pods.

Prerequisites

To follow this tutorial, you will need:

  • A running Kubernetes cluster, such as a CubePath managed Kubernetes cluster or a self-managed cluster on VPS or bare metal.
  • kubectl configured with permissions to create namespaces.
  • The Kubernetes Metrics Server installed, which provides kubectl top. Check with kubectl top nodes; if it returns error: Metrics API not available, install it with kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml.

Create a namespace for the examples:

kubectl create namespace resources-demo

Step 1 - Understanding requests and limits

The two settings do different jobs:

SettingUsed byEffect
requests.cpuScheduler, kernel CPU sharesThe pod is only placed on a node with this much unreserved CPU. Under contention, CPU time is shared in proportion to requests.
requests.memoryScheduler, evictionThe pod is only placed on a node with this much unreserved memory. Pods using more than their request are evicted first under node memory pressure.
limits.cpuKernel (CFS quota)The container is throttled when it uses its CPU quota in a 100 ms period. It is never killed for CPU.
limits.memoryKernel (cgroup)If the container exceeds it, the kernel kills the process and the container shows OOMKilled.

CPU is measured in cores or millicores (500m is half a core). Memory uses bytes with binary suffixes: Mi and Gi. Write 128Mi, not 128m: a lowercase m means millibytes.

Check how much of each node is already reserved by requests:

kubectl describe node your_node_name | grep -A 8 "Allocated resources"
Allocated resources:
  (Total limits may be over 100 percent, i.e., overcommitted.)
  Resource           Requests      Limits
  --------           --------      ------
  cpu                1250m (31%)   3 (75%)
  memory             1920Mi (24%)  4Gi (52%)

The scheduler compares the Requests column, not actual usage, against the node's allocatable capacity.

Step 2 - Setting requests and limits on a Deployment

Requests and limits are set per container under resources. Create a Deployment manifest:

nano web.yaml

This example gives NGINX a quarter core and 64 MiB guaranteed, lets it burst to one core, and caps memory at 128 MiB:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
  namespace: resources-demo
spec:
  replicas: 2
  selector:
    matchLabels:
      app: web
  template:
    metadata:
      labels:
        app: web
    spec:
      containers:
      - name: nginx
        image: nginx:stable
        resources:
          requests:
            cpu: 250m
            memory: 64Mi
          limits:
            cpu: "1"
            memory: 128Mi

Apply it and confirm the pods are running:

kubectl apply -f web.yaml
kubectl get pods -n resources-demo
NAME                   READY   STATUS    RESTARTS   AGE
web-6d9f8c7b5d-4kq2z   1/1     Running   0          15s
web-6d9f8c7b5d-x7m9p   1/1     Running   0          15s

If a pod stays Pending instead, its requests do not fit on any node. Step 5 shows how to confirm that.

Step 3 - Checking the QoS class

Kubernetes assigns every pod a QoS class from its resources. The class decides the eviction order when a node runs out of memory:

QoS classConditionEviction order
GuaranteedEvery container has CPU and memory limits equal to its requestsLast
BurstableAt least one container has a request or limit, but not GuaranteedMiddle, those furthest above their memory request first
BestEffortNo container has any request or limitFirst

Check the class of the web pods:

kubectl get pods -n resources-demo -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass
NAME                   QOS
web-6d9f8c7b5d-4kq2z   Burstable
web-6d9f8c7b5d-x7m9p   Burstable

They are Burstable because limits are higher than requests. Critical stateful workloads such as databases are good candidates for Guaranteed: set limits equal to requests for both CPU and memory in every container.

Step 4 - Reproducing an OOMKill

To see what happens when a container exceeds its memory limit, run a pod that tries to allocate 250 MiB with a 100 MiB limit. The polinux/stress image runs the stress tool:

nano memory-hog.yaml
apiVersion: v1
kind: Pod
metadata:
  name: memory-hog
  namespace: resources-demo
spec:
  restartPolicy: Never
  containers:
  - name: stress
    image: polinux/stress
    command: ["stress"]
    args: ["--vm", "1", "--vm-bytes", "250M", "--vm-hang", "1"]
    resources:
      requests:
        memory: 50Mi
      limits:
        memory: 100Mi

Apply it and check its status after a few seconds:

kubectl apply -f memory-hog.yaml
kubectl get pod memory-hog -n resources-demo
NAME         READY   STATUS      RESTARTS   AGE
memory-hog   0/1     OOMKilled   0          12s

The termination reason and exit code are also recorded in the pod status:

kubectl get pod memory-hog -n resources-demo -o jsonpath='{.status.containerStatuses[0].state.terminated.reason} {.status.containerStatuses[0].state.terminated.exitCode}{"\n"}'
OOMKilled 137

Exit code 137 means the process received SIGKILL. In a Deployment the container would be restarted and, if it keeps exceeding the limit, end up in CrashLoopBackOff. The fix is either a higher memory limit or less memory use in the application, never removing the limit on a shared cluster. Delete the test pod:

kubectl delete pod memory-hog -n resources-demo

Step 5 - Setting namespace defaults with LimitRange

Containers without resources are BestEffort and invisible to the scheduler's capacity math. A LimitRange fills in default requests and limits for containers that do not set them, and rejects values outside a range.

nano limitrange.yaml
apiVersion: v1
kind: LimitRange
metadata:
  name: container-defaults
  namespace: resources-demo
spec:
  limits:
  - type: Container
    defaultRequest:
      cpu: 100m
      memory: 128Mi
    default:
      cpu: 500m
      memory: 256Mi
    max:
      cpu: "2"
      memory: 2Gi

defaultRequest and default are applied to containers that omit requests or limits; max rejects any container asking for more. Apply it and test with a pod that has no resources block:

kubectl apply -f limitrange.yaml
kubectl run defaults-test -n resources-demo --image=nginx:stable
kubectl get pod defaults-test -n resources-demo -o jsonpath='{.spec.containers[0].resources}{"\n"}'
{"limits":{"cpu":"500m","memory":"256Mi"},"requests":{"cpu":"100m","memory":"128Mi"}}

A LimitRange only affects pods created after it exists. Existing pods keep their values until they are recreated.

Step 6 - Capping a namespace with ResourceQuota

A ResourceQuota limits the total resources that all pods in a namespace may request. It is how you keep one team or application from reserving an entire cluster.

nano quota.yaml
apiVersion: v1
kind: ResourceQuota
metadata:
  name: compute-quota
  namespace: resources-demo
spec:
  hard:
    requests.cpu: "2"
    requests.memory: 2Gi
    limits.cpu: "4"
    limits.memory: 4Gi
    pods: "10"

Apply it and check current usage:

kubectl apply -f quota.yaml
kubectl describe resourcequota compute-quota -n resources-demo
Name:            compute-quota
Namespace:       resources-demo
Resource         Used   Hard
--------         ----   ----
limits.cpu       2500m  4
limits.memory    512Mi  4Gi
pods             3      10
requests.cpu     600m   2
requests.memory  256Mi  2Gi

Once a quota covers CPU or memory, every new pod in the namespace must declare those values. The LimitRange from Step 5 supplies them automatically. Try to exceed the quota by scaling the web Deployment:

kubectl scale deployment web -n resources-demo --replicas=8
kubectl get deployment web -n resources-demo
NAME   READY   UP-TO-DATE   AVAILABLE   AGE
web    3/8     3            3           6m

The Deployment cannot create the extra pods. The reason is recorded on its ReplicaSet:

kubectl describe replicaset -n resources-demo -l app=web | grep -m 1 "exceeded quota"
  Warning  FailedCreate  10s  replicaset-controller  Error creating: pods "web-6d9f8c7b5d-p2k8w" is forbidden: exceeded quota: compute-quota, requested: limits.cpu=1, used: limits.cpu=3500m, limited: limits.cpu=4

Scale back down:

kubectl scale deployment web -n resources-demo --replicas=2

Step 7 - Right-sizing with real usage

Good values come from measurement, not guesses. Once a workload has run under normal traffic, compare its usage with its requests:

kubectl top pods -n resources-demo --containers
POD                    NAME    CPU(cores)   MEMORY(bytes)
web-6d9f8c7b5d-4kq2z   nginx   2m           5Mi
web-6d9f8c7b5d-x7m9p   nginx   1m           5Mi

kubectl top shows current usage only. For decisions, look at usage over days in your monitoring system (for example, the p95 of container_cpu_usage_seconds_total and the peak of container_memory_working_set_bytes in Prometheus). A practical approach:

  • Memory request: around the normal peak working set, plus 10 to 20% headroom.
  • Memory limit: equal to the request, or modestly above it. Memory cannot be reclaimed from a process without killing it, so a large gap between request and limit makes node-level OOM more likely.
  • CPU request: around the p95 of usage, so the scheduler reserves what the pod really needs.
  • CPU limit: optional. Many teams leave CPU limits off for latency-sensitive services to avoid throttling, and rely on requests for fair sharing. Keep limits where you need a hard cap, such as on multi-tenant clusters.

To check whether a container is being throttled by its CPU limit, read its cgroup statistics. On nodes with cgroup v2 (Ubuntu 22.04 and later):

kubectl exec -n resources-demo deploy/web -- cat /sys/fs/cgroup/cpu.stat
usage_usec 1843211
user_usec 1203344
system_usec 639867
nr_periods 5213
nr_throttled 0
throttled_usec 0

If nr_throttled grows steadily relative to nr_periods, the limit is too low for the workload's bursts. Raise it or remove it.

When you are done, remove the demo namespace:

kubectl delete namespace resources-demo

Troubleshooting

  • Pod Pending with Insufficient cpu or Insufficient memory. No node has enough unrequested capacity. Lower the requests if they are oversized, or add nodes. Check with kubectl describe pod your_pod.
  • Container restarts with OOMKilled. The memory limit is below the application's real peak. Check kubectl describe pod for Last State: Terminated, Reason: OOMKilled and raise the limit, or tune the application's heap (for example, JVM -XX:MaxRAMPercentage).
  • Pods Evicted with The node was low on resource: memory. The node ran out of memory because the sum of usage exceeded capacity. Pods using far more than their memory request are evicted first; set realistic memory requests.
  • must specify limits.cpu when creating a pod. A ResourceQuota covers that resource and the pod has no value for it. Add resources to the pod or create a LimitRange with defaults.

Conclusion

You set requests and limits on a Deployment, checked its QoS class, reproduced an OOMKill, added namespace defaults with a LimitRange, capped a namespace with a ResourceQuota and used real usage data to size containers. Next, you can scale replicas on CPU usage with a Horizontal Pod Autoscaler, which depends on correct CPU requests, protect critical pods with PriorityClass resources, or add alerts for containers that are regularly throttled or OOMKilled.