๐Ÿ‡ฎ๐Ÿ‡ณ
๐Ÿ‡ฎ๐Ÿ‡ณ
Limited-Time Offer!Get 20% OFF on all live courses
Enroll Now
PrakalpanaLive online tech training
DevOpsโฑ๏ธ 10 min read๐Ÿ“… Oct 1

Kubernetes Troubleshooting Guide: Fix CrashLoopBackOff, Pending, ImagePullBackOff & OOMKilled

AS
Ankit Sharmaโ€ขDevOps Lead
๐Ÿ“‘ Contents (12 sections)

๐Ÿ“ŒA Method Before the Commands

Most Kubernetes problems are solved by the same three commands, in this order:

  • 1kubectl get pods -o wide โ€” what state is it in, and on which node?
  • 2kubectl describe pod โ€” read the Events at the bottom
  • 3kubectl logs --previous โ€” what did the container say before it died?
  • Add kubectl get events --sort-by=.lastTimestamp for a cluster-wide timeline. This is exactly how we teach troubleshooting in our Kubernetes course: break it, then fix it.

    ๐Ÿ“ŒCrashLoopBackOff

    Meaning: the container starts, exits, and Kubernetes keeps restarting it with increasing back-off delays.

    Diagnose:

    kubectl logs <pod> --previous
    kubectl describe pod <pod>

    Common causes and fixes:

  • The app crashes on startup โ€” a missing environment variable, a bad config or an unreachable database. Fix the config or Secret.
  • The wrong command or entrypoint โ€” the main process exits immediately.
  • A liveness probe that fails too early โ€” add a startup probe or increase initialDelaySeconds.
  • Exit code 137 โ€” killed for memory (see OOMKilled below).
  • ๐Ÿ“ŒPending

    Meaning: the scheduler cannot place the pod on any node.

    Diagnose: kubectl describe pod and read the FailedScheduling event.

    Common causes and fixes:

  • Insufficient cpu/memory โ€” lower the requests, add nodes or enable the Cluster Autoscaler.
  • Unbound PersistentVolumeClaim โ€” check the StorageClass and kubectl get pvc.
  • Taints without tolerations โ€” add a toleration or use a different node pool.
  • A nodeSelector or affinity that matches no node โ€” fix the labels.
  • ๐Ÿ“ŒImagePullBackOff / ErrImagePull

    Meaning: the kubelet cannot pull the image.

    Common causes and fixes:

  • A typo in the image name or a tag that does not exist โ€” check the registry.
  • A private registry โ€” create a docker-registry Secret and reference it in imagePullSecrets.
  • On EKS, the node role lacks ECR pull permissions.
  • Docker Hub rate limits โ€” authenticate or mirror the image.
  • ๐Ÿ“ŒOOMKilled

    Meaning: the container exceeded its memory limit and the kernel killed it (exit code 137).

    Diagnose:

    kubectl describe pod <pod> # Last State: Terminated, Reason: OOMKilled
    kubectl top pod <pod>

    Fixes: raise the memory limit to match real usage, fix the leak, or tune the runtime โ€” for example, make sure the JVM respects container limits with -XX:MaxRAMPercentage.

    ๐Ÿ“ŒCreateContainerConfigError

    Meaning: the pod references a ConfigMap, Secret or key that does not exist. Check the names in the pod spec and the namespace โ€” Secrets are namespace-scoped.

    ๐Ÿ“ŒService Not Reachable

    Diagnose:

    kubectl get endpoints <service>
    kubectl get pods --show-labels

    Common causes:

  • Empty endpoints โ€” the Service selector does not match the pod labels.
  • Wrong targetPort โ€” it must match the port the container listens on.
  • Readiness probe failing โ€” the pod is running but not marked ready.
  • NetworkPolicy blocking the traffic.
  • For Ingress, check the ingress controller logs and the host and path rules.
  • Test from inside the cluster:

    kubectl run tmp --rm -it --image=busybox -- wget -qO- http://<service>.<namespace>:80

    ๐Ÿ“ŒNode NotReady

    Check kubectl describe node for conditions such as MemoryPressure, DiskPressure and PIDPressure, then the kubelet and container runtime logs on the node. A full disk from images and logs is a very common cause.

    ๐Ÿ“ŒEvicted Pods

    Pods are evicted when a node runs low on memory or disk. Set realistic requests, keep critical workloads in the Guaranteed QoS class, and clean up with kubectl delete pod --field-selector=status.phase=Failed.

    ๐Ÿ“ŒRollout Stuck

    kubectl rollout status deployment/<name>
    kubectl rollout history deployment/<name>
    kubectl rollout undo deployment/<name>

    New pods failing readiness will stall a rolling update โ€” roll back first, then debug the new version.

    ๐Ÿ“ŒTroubleshooting Cheat Sheet

    SymptomFirst commandUsual cause

    CrashLoopBackOffkubectl logs --previousApp error, bad config, probe Pendingkubectl describe podResources, PVC, taints ImagePullBackOffkubectl describe podImage name, registry auth OOMKilledkubectl describe podMemory limit too low Service unreachablekubectl get endpointsSelector or port mismatch Node NotReadykubectl describe nodeKubelet, disk or memory

    ๐Ÿ“ŒKeep Learning

    Prepare for interviews with our Kubernetes interview questions, and if containers are still new, start with Docker vs Kubernetes. Prakalpana's Kubernetes course runs these break-and-fix labs on real clusters, live online or 1-on-1, with CKA/CKAD preparation. WhatsApp or call +91 9243078181 for a free demo.

    AS

    Written by

    Ankit Sharma

    DevOps Lead

    ๐Ÿš€ Master DevOps

    Live online + 1-on-1 Kubernetes (K8s) training ยท Join 5000+ developers