Most Kubernetes problems show up as a pod that will not start, keeps restarting or cannot be reached. The good news is that a handful of kubectl commands reveal the cause in almost every case. This tutorial walks through a repeatable diagnostic workflow and then covers the most common failures one by one: CrashLoopBackOff, ImagePullBackOff, Pending pods, OOMKilled containers, Services with no endpoints and DNS failures, with the commands that identify each one and how to fix it.

Prerequisites

To follow this tutorial, you will need:

  • A running Kubernetes cluster, such as a CubePath managed Kubernetes cluster or a self-managed cluster on VPS or bare metal.
  • kubectl configured for the cluster, with read access to pods, events and nodes in the namespaces you debug. Node-level commands at the end require SSH access to the nodes.

In the commands below, replace your_namespace and your_pod with your own names. You can avoid repeating -n by setting a default namespace for your context:

kubectl config set-context --current --namespace=your_namespace

Step 1 - Following a diagnostic workflow

Work from the outside in: first the pod's status, then its events, then the container logs. Start by listing pods with their node and IP:

kubectl get pods -o wide
NAME                   READY   STATUS             RESTARTS      AGE   IP           NODE
api-7c9d8f6b5d-4kq2z   0/1     CrashLoopBackOff   5 (40s ago)   4m    10.42.1.18   worker-1
web-6d9f8c7b5d-x7m9p   1/1     Running            0             2d    10.42.2.11   worker-2

The STATUS column already tells you which section of this guide applies. Next, describe the pod. The Events section at the bottom and the Last State of each container are the most useful parts:

kubectl describe pod your_pod

Events are also available for the whole namespace, sorted by time. This is often the fastest way to spot a problem affecting several pods:

kubectl get events --sort-by=.lastTimestamp

Finally, read the logs. For a container that is restarting, the current instance may have no output yet, so read the previous one with --previous:

kubectl logs your_pod
kubectl logs your_pod --previous

For pods with several containers, add -c container_name. For a Deployment, kubectl logs deploy/your_deployment picks one of its pods.

Step 2 - Fixing CrashLoopBackOff

CrashLoopBackOff means the container starts, exits, and Kubernetes waits longer and longer before restarting it. The container itself is failing, so the answer is almost always in the logs and the exit code.

Read the previous container's logs:

kubectl logs your_pod --previous
Error: connect ECONNREFUSED 10.43.12.7:5432

Then check the exit code in the container's last state:

kubectl get pod your_pod -o jsonpath='{.status.containerStatuses[0].lastState.terminated.exitCode} {.status.containerStatuses[0].lastState.terminated.reason}{"\n"}'
1 Error

Use the exit code to narrow down the cause:

Exit codeMeaningTypical cause
1 or other small numbersThe application exited with an errorBad configuration, missing environment variable, dependency unreachable
126 / 127Command not executable / not foundWrong command or args, missing binary in the image
137Killed by SIGKILLMemory limit exceeded (OOMKilled) or failed liveness probe
143Terminated by SIGTERMThe container was stopped, often by a liveness probe failure

Common fixes:

  • Missing configuration. Check that the ConfigMaps and Secrets the pod references exist with kubectl get configmap,secret. A missing one usually shows as CreateContainerConfigError rather than a crash, but a missing key inside an existing one lets the app start and then fail.
  • Dependency not ready. An app that exits when its database is unreachable will crash-loop until the database is up. Make the application retry its connection, and verify the dependency's Service as described in Step 6.
  • Liveness probe too aggressive. If kubectl describe pod shows Liveness probe failed before each restart, the app is killed while still starting. Add a startupProbe or raise initialDelaySeconds and failureThreshold on the liveness probe.

If the container exits too fast to inspect, start a copy of the pod with a shell instead of the normal entrypoint:

kubectl debug your_pod -it --copy-to=your_pod-debug --container=app -- sh

Replace app with the container's name. Inside the copy you can inspect files and environment variables and run the entrypoint by hand. Delete the copy with kubectl delete pod your_pod-debug when you finish.

Step 3 - Fixing ImagePullBackOff and ErrImagePull

These states mean the kubelet could not download the image. The events show the exact reason:

kubectl describe pod your_pod | grep -A 10 Events
  Warning  Failed     12s   kubelet  Failed to pull image "registry.example.com/team/api:1.4.2": rpc error: code = NotFound desc = failed to pull and unpack image "registry.example.com/team/api:1.4.2": not found
  Warning  Failed     12s   kubelet  Error: ErrImagePull
  Normal   BackOff    1s    kubelet  Back-off pulling image "registry.example.com/team/api:1.4.2"

Match the message to its cause:

  • not found or manifest unknown. The repository or tag does not exist. Check the spelling and that the tag was pushed.
  • unauthorized or authentication required. The registry is private and the pod has no credentials. Create a registry Secret and reference it from the pod.
  • dial tcp ... i/o timeout or no such host. The node cannot reach the registry. Check node DNS and outbound firewall rules.

To add credentials for a private registry, create a Secret of type docker-registry in the pod's namespace:

kubectl create secret docker-registry regcred \
  --docker-server=registry.example.com \
  --docker-username=your_user \
  --docker-password='your_registry_token'

Reference it in the pod template of your Deployment, at the same level as containers:

spec:
  imagePullSecrets:
  - name: regcred
  containers:
  - name: api
    image: registry.example.com/team/api:1.4.2

After applying the change, the new pods pull the image and move to Running.

Step 4 - Fixing Pending pods

A Pending pod has not been placed on a node, or is waiting for a volume. The scheduler writes the reason as a FailedScheduling event:

kubectl describe pod your_pod | grep -A 5 Events
  Warning  FailedScheduling  30s  default-scheduler  0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient memory. preemption: 0/3 nodes are available: 3 No preemption victims found for incoming pod.

The most common reasons are:

  • Insufficient cpu / Insufficient memory. The pod's requests do not fit in the unrequested capacity of any node. Compare the pod's requests with the Allocated resources section of kubectl describe node your_node_name. Lower oversized requests or add capacity.
  • untolerated taint. The only nodes with room carry a taint the pod does not tolerate. List taints with kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints. Either add a matching tolerations entry to the pod or schedule on other nodes.
  • didn't match Pod's node affinity/selector. A nodeSelector or affinity rule refers to labels no node has. Check them with kubectl get nodes --show-labels.
  • pod has unbound immediate PersistentVolumeClaims. The volume cannot be provisioned. Check the claim:
kubectl get pvc
kubectl describe pvc your_claim

A claim stuck in Pending usually names a storageClassName that does not exist (compare with kubectl get storageclass) or has no default StorageClass to fall back on.

If a Deployment shows fewer pods than desired but no Pending pods at all, a ResourceQuota may be rejecting them. Look for exceeded quota in kubectl describe replicaset and check usage with kubectl describe resourcequota.

Step 5 - Fixing OOMKilled and evicted pods

When a container exceeds its memory limit, the kernel kills it. kubectl describe pod shows it in the container's last state:

    Last State:     Terminated
      Reason:       OOMKilled
      Exit Code:    137

Check the limit and the actual usage of the pod:

kubectl get pod your_pod -o jsonpath='{.spec.containers[*].resources}{"\n"}'
kubectl top pod your_pod --containers

Raise resources.limits.memory to cover the application's real peak, or reduce its memory use (for runtimes such as the JVM, size the heap relative to the container limit).

Pods with status Evicted or Error and the message The node was low on resource: memory were removed by the kubelet because the whole node ran short. Pods using the most memory above their request go first. Set realistic memory requests so the scheduler does not overpack nodes, and delete the leftover evicted pod objects:

kubectl delete pods --field-selector=status.phase=Failed

Step 6 - Fixing Services with no endpoints

If pods are Running but other pods cannot reach them through a Service, the Service is often not selecting any pods. Check its EndpointSlices:

kubectl get endpointslices -l kubernetes.io/service-name=your_service
NAME                 ADDRESSTYPE   PORTS     ENDPOINTS   AGE
your_service-9xk2p   IPv4          <unset>   <unset>     10m

Empty ENDPOINTS means no ready pod matches the selector. Compare the Service's selector with the pod labels:

kubectl get service your_service -o jsonpath='{.spec.selector}{"\n"}'
kubectl get pods --show-labels

A typo in a label (app: api against app: api-server) is the most frequent cause. Pods that fail their readiness probe are also left out, so check READY in kubectl get pods.

If endpoints exist but connections fail, check that the Service targetPort matches the port the container listens on. Test the pod directly, bypassing the Service, with a port-forward:

kubectl port-forward pod/your_pod 8080:your_container_port

In a second terminal:

curl -i http://localhost:8080/

If this works but the Service does not, the problem is the Service definition or a NetworkPolicy blocking traffic. List policies with kubectl get networkpolicy.

Step 7 - Debugging cluster DNS

Pods resolve Service names like your_service.your_namespace.svc.cluster.local through CoreDNS. Start a pod with DNS tools to test resolution from inside the cluster:

kubectl run dnsutils --image=registry.k8s.io/e2e-test-images/jessie-dnsutils:1.3 --restart=Never --command -- sleep infinity
kubectl wait --for=condition=Ready pod/dnsutils --timeout=60s
kubectl exec dnsutils -- nslookup kubernetes.default
Server:		10.43.0.10
Address:	10.43.0.10#53

Name:	kubernetes.default.svc.cluster.local
Address: 10.43.0.1

If that fails, check the pod's resolver configuration and the CoreDNS pods:

kubectl exec dnsutils -- cat /etc/resolv.conf
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns

nameserver in resolv.conf must match the cluster IP of the kube-dns Service (kubectl get svc -n kube-system kube-dns). If CoreDNS pods are crashing or logging errors, restart them with kubectl rollout restart deployment coredns -n kube-system and review its configuration in the coredns ConfigMap. If only external names fail, check the upstream resolvers CoreDNS forwards to.

Remove the test pod when done:

kubectl delete pod dnsutils

Step 8 - Checking node health

When many pods on one node fail at once, look at the node:

kubectl get nodes
kubectl describe node your_node_name

In the Conditions section, Ready must be True, and MemoryPressure, DiskPressure and PIDPressure must be False. A node under DiskPressure evicts pods and refuses new ones until disk space is freed, usually from old images or container logs.

If you have SSH access to the node, check the kubelet and container runtime:

sudo systemctl status kubelet
sudo journalctl -u kubelet --since "30 minutes ago"
sudo crictl ps -a

Without SSH, kubectl debug can start a privileged pod on the node with the host filesystem mounted at /host:

kubectl debug node/your_node_name -it --image=ubuntu:24.04

Inside it, df -h /host shows disk usage on the node. Delete the debug pod afterwards; kubectl get pods lists it as node-debugger-....

Conclusion

You now have a repeatable workflow for failing pods: read the status, then the events, then the logs, and use the exit code or scheduler message to point to the cause. You also covered the fixes for crash loops, image pulls, scheduling, memory, Service selection, DNS and node health. To catch these problems before users do, add alerts on pod restarts and Pending pods, set realistic resource requests and limits, and define readiness and startup probes for every workload.