Most Kubernetes problems show up as a pod that will not start, keeps restarting or cannot be reached. The good news is that a handful of kubectl commands reveal the cause in almost every case. This tutorial walks through a repeatable diagnostic workflow and then covers the most common failures one by one: CrashLoopBackOff, ImagePullBackOff, Pending pods, OOMKilled containers, Services with no endpoints and DNS failures, with the commands that identify each one and how to fix it.
Prerequisites
To follow this tutorial, you will need:
- A running Kubernetes cluster, such as a CubePath managed Kubernetes cluster or a self-managed cluster on VPS or bare metal.
kubectlconfigured for the cluster, with read access to pods, events and nodes in the namespaces you debug. Node-level commands at the end require SSH access to the nodes.
In the commands below, replace your_namespace and your_pod with your own names. You can avoid repeating -n by setting a default namespace for your context:
kubectl config set-context --current --namespace=your_namespace
Step 1 - Following a diagnostic workflow
Work from the outside in: first the pod's status, then its events, then the container logs. Start by listing pods with their node and IP:
kubectl get pods -o wide
NAME READY STATUS RESTARTS AGE IP NODE
api-7c9d8f6b5d-4kq2z 0/1 CrashLoopBackOff 5 (40s ago) 4m 10.42.1.18 worker-1
web-6d9f8c7b5d-x7m9p 1/1 Running 0 2d 10.42.2.11 worker-2
The STATUS column already tells you which section of this guide applies. Next, describe the pod. The Events section at the bottom and the Last State of each container are the most useful parts:
kubectl describe pod your_pod
Events are also available for the whole namespace, sorted by time. This is often the fastest way to spot a problem affecting several pods:
kubectl get events --sort-by=.lastTimestamp
Finally, read the logs. For a container that is restarting, the current instance may have no output yet, so read the previous one with --previous:
kubectl logs your_pod
kubectl logs your_pod --previous
For pods with several containers, add -c container_name. For a Deployment, kubectl logs deploy/your_deployment picks one of its pods.
Step 2 - Fixing CrashLoopBackOff
CrashLoopBackOff means the container starts, exits, and Kubernetes waits longer and longer before restarting it. The container itself is failing, so the answer is almost always in the logs and the exit code.
Read the previous container's logs:
kubectl logs your_pod --previous
Error: connect ECONNREFUSED 10.43.12.7:5432
Then check the exit code in the container's last state:
kubectl get pod your_pod -o jsonpath='{.status.containerStatuses[0].lastState.terminated.exitCode} {.status.containerStatuses[0].lastState.terminated.reason}{"\n"}'
1 Error
Use the exit code to narrow down the cause:
| Exit code | Meaning | Typical cause |
|---|---|---|
1 or other small numbers | The application exited with an error | Bad configuration, missing environment variable, dependency unreachable |
126 / 127 | Command not executable / not found | Wrong command or args, missing binary in the image |
137 | Killed by SIGKILL | Memory limit exceeded (OOMKilled) or failed liveness probe |
143 | Terminated by SIGTERM | The container was stopped, often by a liveness probe failure |
Common fixes:
- Missing configuration. Check that the ConfigMaps and Secrets the pod references exist with
kubectl get configmap,secret. A missing one usually shows asCreateContainerConfigErrorrather than a crash, but a missing key inside an existing one lets the app start and then fail. - Dependency not ready. An app that exits when its database is unreachable will crash-loop until the database is up. Make the application retry its connection, and verify the dependency's Service as described in Step 6.
- Liveness probe too aggressive. If
kubectl describe podshowsLiveness probe failedbefore each restart, the app is killed while still starting. Add astartupProbeor raiseinitialDelaySecondsandfailureThresholdon the liveness probe.
If the container exits too fast to inspect, start a copy of the pod with a shell instead of the normal entrypoint:
kubectl debug your_pod -it --copy-to=your_pod-debug --container=app -- sh
Replace app with the container's name. Inside the copy you can inspect files and environment variables and run the entrypoint by hand. Delete the copy with kubectl delete pod your_pod-debug when you finish.
Step 3 - Fixing ImagePullBackOff and ErrImagePull
These states mean the kubelet could not download the image. The events show the exact reason:
kubectl describe pod your_pod | grep -A 10 Events
Warning Failed 12s kubelet Failed to pull image "registry.example.com/team/api:1.4.2": rpc error: code = NotFound desc = failed to pull and unpack image "registry.example.com/team/api:1.4.2": not found
Warning Failed 12s kubelet Error: ErrImagePull
Normal BackOff 1s kubelet Back-off pulling image "registry.example.com/team/api:1.4.2"
Match the message to its cause:
not foundormanifest unknown. The repository or tag does not exist. Check the spelling and that the tag was pushed.unauthorizedorauthentication required. The registry is private and the pod has no credentials. Create a registry Secret and reference it from the pod.dial tcp ... i/o timeoutorno such host. The node cannot reach the registry. Check node DNS and outbound firewall rules.
To add credentials for a private registry, create a Secret of type docker-registry in the pod's namespace:
kubectl create secret docker-registry regcred \
--docker-server=registry.example.com \
--docker-username=your_user \
--docker-password='your_registry_token'
Reference it in the pod template of your Deployment, at the same level as containers:
spec:
imagePullSecrets:
- name: regcred
containers:
- name: api
image: registry.example.com/team/api:1.4.2
After applying the change, the new pods pull the image and move to Running.
Step 4 - Fixing Pending pods
A Pending pod has not been placed on a node, or is waiting for a volume. The scheduler writes the reason as a FailedScheduling event:
kubectl describe pod your_pod | grep -A 5 Events
Warning FailedScheduling 30s default-scheduler 0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient memory. preemption: 0/3 nodes are available: 3 No preemption victims found for incoming pod.
The most common reasons are:
Insufficient cpu/Insufficient memory. The pod's requests do not fit in the unrequested capacity of any node. Compare the pod's requests with theAllocated resourcessection ofkubectl describe node your_node_name. Lower oversized requests or add capacity.untolerated taint. The only nodes with room carry a taint the pod does not tolerate. List taints withkubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints. Either add a matchingtolerationsentry to the pod or schedule on other nodes.didn't match Pod's node affinity/selector. AnodeSelectoror affinity rule refers to labels no node has. Check them withkubectl get nodes --show-labels.pod has unbound immediate PersistentVolumeClaims. The volume cannot be provisioned. Check the claim:
kubectl get pvc
kubectl describe pvc your_claim
A claim stuck in Pending usually names a storageClassName that does not exist (compare with kubectl get storageclass) or has no default StorageClass to fall back on.
If a Deployment shows fewer pods than desired but no Pending pods at all, a ResourceQuota may be rejecting them. Look for exceeded quota in kubectl describe replicaset and check usage with kubectl describe resourcequota.
Step 5 - Fixing OOMKilled and evicted pods
When a container exceeds its memory limit, the kernel kills it. kubectl describe pod shows it in the container's last state:
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Check the limit and the actual usage of the pod:
kubectl get pod your_pod -o jsonpath='{.spec.containers[*].resources}{"\n"}'
kubectl top pod your_pod --containers
Raise resources.limits.memory to cover the application's real peak, or reduce its memory use (for runtimes such as the JVM, size the heap relative to the container limit).
Pods with status Evicted or Error and the message The node was low on resource: memory were removed by the kubelet because the whole node ran short. Pods using the most memory above their request go first. Set realistic memory requests so the scheduler does not overpack nodes, and delete the leftover evicted pod objects:
kubectl delete pods --field-selector=status.phase=Failed
Step 6 - Fixing Services with no endpoints
If pods are Running but other pods cannot reach them through a Service, the Service is often not selecting any pods. Check its EndpointSlices:
kubectl get endpointslices -l kubernetes.io/service-name=your_service
NAME ADDRESSTYPE PORTS ENDPOINTS AGE
your_service-9xk2p IPv4 <unset> <unset> 10m
Empty ENDPOINTS means no ready pod matches the selector. Compare the Service's selector with the pod labels:
kubectl get service your_service -o jsonpath='{.spec.selector}{"\n"}'
kubectl get pods --show-labels
A typo in a label (app: api against app: api-server) is the most frequent cause. Pods that fail their readiness probe are also left out, so check READY in kubectl get pods.
If endpoints exist but connections fail, check that the Service targetPort matches the port the container listens on. Test the pod directly, bypassing the Service, with a port-forward:
kubectl port-forward pod/your_pod 8080:your_container_port
In a second terminal:
curl -i http://localhost:8080/
If this works but the Service does not, the problem is the Service definition or a NetworkPolicy blocking traffic. List policies with kubectl get networkpolicy.
Step 7 - Debugging cluster DNS
Pods resolve Service names like your_service.your_namespace.svc.cluster.local through CoreDNS. Start a pod with DNS tools to test resolution from inside the cluster:
kubectl run dnsutils --image=registry.k8s.io/e2e-test-images/jessie-dnsutils:1.3 --restart=Never --command -- sleep infinity
kubectl wait --for=condition=Ready pod/dnsutils --timeout=60s
kubectl exec dnsutils -- nslookup kubernetes.default
Server: 10.43.0.10
Address: 10.43.0.10#53
Name: kubernetes.default.svc.cluster.local
Address: 10.43.0.1
If that fails, check the pod's resolver configuration and the CoreDNS pods:
kubectl exec dnsutils -- cat /etc/resolv.conf
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns
nameserver in resolv.conf must match the cluster IP of the kube-dns Service (kubectl get svc -n kube-system kube-dns). If CoreDNS pods are crashing or logging errors, restart them with kubectl rollout restart deployment coredns -n kube-system and review its configuration in the coredns ConfigMap. If only external names fail, check the upstream resolvers CoreDNS forwards to.
Remove the test pod when done:
kubectl delete pod dnsutils
Step 8 - Checking node health
When many pods on one node fail at once, look at the node:
kubectl get nodes
kubectl describe node your_node_name
In the Conditions section, Ready must be True, and MemoryPressure, DiskPressure and PIDPressure must be False. A node under DiskPressure evicts pods and refuses new ones until disk space is freed, usually from old images or container logs.
If you have SSH access to the node, check the kubelet and container runtime:
sudo systemctl status kubelet
sudo journalctl -u kubelet --since "30 minutes ago"
sudo crictl ps -a
Without SSH, kubectl debug can start a privileged pod on the node with the host filesystem mounted at /host:
kubectl debug node/your_node_name -it --image=ubuntu:24.04
Inside it, df -h /host shows disk usage on the node. Delete the debug pod afterwards; kubectl get pods lists it as node-debugger-....
NoteOn managed Kubernetes, node components are operated by the provider. If a node stays
NotReady, drain it withkubectl drain your_node_name --ignore-daemonsets --delete-emptydir-dataand contact support or replace the node.
Conclusion
You now have a repeatable workflow for failing pods: read the status, then the events, then the logs, and use the exit code or scheduler message to point to the cause. You also covered the fixes for crash loops, image pulls, scheduling, memory, Service selection, DNS and node health. To catch these problems before users do, add alerts on pod restarts and Pending pods, set realistic resource requests and limits, and define readiness and startup probes for every workload.
