Every container in Kubernetes can declare how much CPU and memory it needs (requests) and the most it may use (limits). Requests drive scheduling, limits are enforced by the kernel, and together they decide which pods are evicted first when a node runs short of memory. In this tutorial you will set requests and limits on a workload, check its Quality of Service (QoS) class, reproduce an out-of-memory kill, apply namespace defaults with a LimitRange, cap a namespace with a ResourceQuota and use real usage data to right-size your pods.
Prerequisites
To follow this tutorial, you will need:
- A running Kubernetes cluster, such as a CubePath managed Kubernetes cluster or a self-managed cluster on VPS or bare metal.
kubectlconfigured with permissions to create namespaces.- The Kubernetes Metrics Server installed, which provides
kubectl top. Check withkubectl top nodes; if it returnserror: Metrics API not available, install it withkubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml.
Create a namespace for the examples:
kubectl create namespace resources-demo
Step 1 - Understanding requests and limits
The two settings do different jobs:
| Setting | Used by | Effect |
|---|---|---|
requests.cpu | Scheduler, kernel CPU shares | The pod is only placed on a node with this much unreserved CPU. Under contention, CPU time is shared in proportion to requests. |
requests.memory | Scheduler, eviction | The pod is only placed on a node with this much unreserved memory. Pods using more than their request are evicted first under node memory pressure. |
limits.cpu | Kernel (CFS quota) | The container is throttled when it uses its CPU quota in a 100 ms period. It is never killed for CPU. |
limits.memory | Kernel (cgroup) | If the container exceeds it, the kernel kills the process and the container shows OOMKilled. |
CPU is measured in cores or millicores (500m is half a core). Memory uses bytes with binary suffixes: Mi and Gi. Write 128Mi, not 128m: a lowercase m means millibytes.
Check how much of each node is already reserved by requests:
kubectl describe node your_node_name | grep -A 8 "Allocated resources"
Allocated resources:
(Total limits may be over 100 percent, i.e., overcommitted.)
Resource Requests Limits
-------- -------- ------
cpu 1250m (31%) 3 (75%)
memory 1920Mi (24%) 4Gi (52%)
The scheduler compares the Requests column, not actual usage, against the node's allocatable capacity.
Step 2 - Setting requests and limits on a Deployment
Requests and limits are set per container under resources. Create a Deployment manifest:
nano web.yaml
This example gives NGINX a quarter core and 64 MiB guaranteed, lets it burst to one core, and caps memory at 128 MiB:
apiVersion: apps/v1
kind: Deployment
metadata:
name: web
namespace: resources-demo
spec:
replicas: 2
selector:
matchLabels:
app: web
template:
metadata:
labels:
app: web
spec:
containers:
- name: nginx
image: nginx:stable
resources:
requests:
cpu: 250m
memory: 64Mi
limits:
cpu: "1"
memory: 128Mi
Apply it and confirm the pods are running:
kubectl apply -f web.yaml
kubectl get pods -n resources-demo
NAME READY STATUS RESTARTS AGE
web-6d9f8c7b5d-4kq2z 1/1 Running 0 15s
web-6d9f8c7b5d-x7m9p 1/1 Running 0 15s
If a pod stays Pending instead, its requests do not fit on any node. Step 5 shows how to confirm that.
Step 3 - Checking the QoS class
Kubernetes assigns every pod a QoS class from its resources. The class decides the eviction order when a node runs out of memory:
| QoS class | Condition | Eviction order |
|---|---|---|
Guaranteed | Every container has CPU and memory limits equal to its requests | Last |
Burstable | At least one container has a request or limit, but not Guaranteed | Middle, those furthest above their memory request first |
BestEffort | No container has any request or limit | First |
Check the class of the web pods:
kubectl get pods -n resources-demo -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass
NAME QOS
web-6d9f8c7b5d-4kq2z Burstable
web-6d9f8c7b5d-x7m9p Burstable
They are Burstable because limits are higher than requests. Critical stateful workloads such as databases are good candidates for Guaranteed: set limits equal to requests for both CPU and memory in every container.
Step 4 - Reproducing an OOMKill
To see what happens when a container exceeds its memory limit, run a pod that tries to allocate 250 MiB with a 100 MiB limit. The polinux/stress image runs the stress tool:
nano memory-hog.yaml
apiVersion: v1
kind: Pod
metadata:
name: memory-hog
namespace: resources-demo
spec:
restartPolicy: Never
containers:
- name: stress
image: polinux/stress
command: ["stress"]
args: ["--vm", "1", "--vm-bytes", "250M", "--vm-hang", "1"]
resources:
requests:
memory: 50Mi
limits:
memory: 100Mi
Apply it and check its status after a few seconds:
kubectl apply -f memory-hog.yaml
kubectl get pod memory-hog -n resources-demo
NAME READY STATUS RESTARTS AGE
memory-hog 0/1 OOMKilled 0 12s
The termination reason and exit code are also recorded in the pod status:
kubectl get pod memory-hog -n resources-demo -o jsonpath='{.status.containerStatuses[0].state.terminated.reason} {.status.containerStatuses[0].state.terminated.exitCode}{"\n"}'
OOMKilled 137
Exit code 137 means the process received SIGKILL. In a Deployment the container would be restarted and, if it keeps exceeding the limit, end up in CrashLoopBackOff. The fix is either a higher memory limit or less memory use in the application, never removing the limit on a shared cluster. Delete the test pod:
kubectl delete pod memory-hog -n resources-demo
Step 5 - Setting namespace defaults with LimitRange
Containers without resources are BestEffort and invisible to the scheduler's capacity math. A LimitRange fills in default requests and limits for containers that do not set them, and rejects values outside a range.
nano limitrange.yaml
apiVersion: v1
kind: LimitRange
metadata:
name: container-defaults
namespace: resources-demo
spec:
limits:
- type: Container
defaultRequest:
cpu: 100m
memory: 128Mi
default:
cpu: 500m
memory: 256Mi
max:
cpu: "2"
memory: 2Gi
defaultRequest and default are applied to containers that omit requests or limits; max rejects any container asking for more. Apply it and test with a pod that has no resources block:
kubectl apply -f limitrange.yaml
kubectl run defaults-test -n resources-demo --image=nginx:stable
kubectl get pod defaults-test -n resources-demo -o jsonpath='{.spec.containers[0].resources}{"\n"}'
{"limits":{"cpu":"500m","memory":"256Mi"},"requests":{"cpu":"100m","memory":"128Mi"}}
A LimitRange only affects pods created after it exists. Existing pods keep their values until they are recreated.
Step 6 - Capping a namespace with ResourceQuota
A ResourceQuota limits the total resources that all pods in a namespace may request. It is how you keep one team or application from reserving an entire cluster.
nano quota.yaml
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-quota
namespace: resources-demo
spec:
hard:
requests.cpu: "2"
requests.memory: 2Gi
limits.cpu: "4"
limits.memory: 4Gi
pods: "10"
Apply it and check current usage:
kubectl apply -f quota.yaml
kubectl describe resourcequota compute-quota -n resources-demo
Name: compute-quota
Namespace: resources-demo
Resource Used Hard
-------- ---- ----
limits.cpu 2500m 4
limits.memory 512Mi 4Gi
pods 3 10
requests.cpu 600m 2
requests.memory 256Mi 2Gi
Once a quota covers CPU or memory, every new pod in the namespace must declare those values. The LimitRange from Step 5 supplies them automatically. Try to exceed the quota by scaling the web Deployment:
kubectl scale deployment web -n resources-demo --replicas=8
kubectl get deployment web -n resources-demo
NAME READY UP-TO-DATE AVAILABLE AGE
web 3/8 3 3 6m
The Deployment cannot create the extra pods. The reason is recorded on its ReplicaSet:
kubectl describe replicaset -n resources-demo -l app=web | grep -m 1 "exceeded quota"
Warning FailedCreate 10s replicaset-controller Error creating: pods "web-6d9f8c7b5d-p2k8w" is forbidden: exceeded quota: compute-quota, requested: limits.cpu=1, used: limits.cpu=3500m, limited: limits.cpu=4
Scale back down:
kubectl scale deployment web -n resources-demo --replicas=2
Step 7 - Right-sizing with real usage
Good values come from measurement, not guesses. Once a workload has run under normal traffic, compare its usage with its requests:
kubectl top pods -n resources-demo --containers
POD NAME CPU(cores) MEMORY(bytes)
web-6d9f8c7b5d-4kq2z nginx 2m 5Mi
web-6d9f8c7b5d-x7m9p nginx 1m 5Mi
kubectl top shows current usage only. For decisions, look at usage over days in your monitoring system (for example, the p95 of container_cpu_usage_seconds_total and the peak of container_memory_working_set_bytes in Prometheus). A practical approach:
- Memory request: around the normal peak working set, plus 10 to 20% headroom.
- Memory limit: equal to the request, or modestly above it. Memory cannot be reclaimed from a process without killing it, so a large gap between request and limit makes node-level OOM more likely.
- CPU request: around the p95 of usage, so the scheduler reserves what the pod really needs.
- CPU limit: optional. Many teams leave CPU limits off for latency-sensitive services to avoid throttling, and rely on requests for fair sharing. Keep limits where you need a hard cap, such as on multi-tenant clusters.
To check whether a container is being throttled by its CPU limit, read its cgroup statistics. On nodes with cgroup v2 (Ubuntu 22.04 and later):
kubectl exec -n resources-demo deploy/web -- cat /sys/fs/cgroup/cpu.stat
usage_usec 1843211
user_usec 1203344
system_usec 639867
nr_periods 5213
nr_throttled 0
throttled_usec 0
If nr_throttled grows steadily relative to nr_periods, the limit is too low for the workload's bursts. Raise it or remove it.
TipThe Vertical Pod Autoscaler, installed separately, can run in recommendation-only mode (
updateMode: "Off") and suggest requests from observed usage without changing your pods.
When you are done, remove the demo namespace:
kubectl delete namespace resources-demo
Troubleshooting
- Pod
PendingwithInsufficient cpuorInsufficient memory. No node has enough unrequested capacity. Lower the requests if they are oversized, or add nodes. Check withkubectl describe pod your_pod. - Container restarts with
OOMKilled. The memory limit is below the application's real peak. Checkkubectl describe podforLast State: Terminated, Reason: OOMKilledand raise the limit, or tune the application's heap (for example, JVM-XX:MaxRAMPercentage). - Pods
EvictedwithThe node was low on resource: memory. The node ran out of memory because the sum of usage exceeded capacity. Pods using far more than their memory request are evicted first; set realistic memory requests. must specify limits.cpuwhen creating a pod. A ResourceQuota covers that resource and the pod has no value for it. Addresourcesto the pod or create aLimitRangewith defaults.
Conclusion
You set requests and limits on a Deployment, checked its QoS class, reproduced an OOMKill, added namespace defaults with a LimitRange, capped a namespace with a ResourceQuota and used real usage data to size containers. Next, you can scale replicas on CPU usage with a Horizontal Pod Autoscaler, which depends on correct CPU requests, protect critical pods with PriorityClass resources, or add alerts for containers that are regularly throttled or OOMKilled.
