Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Appendix C: Common Kubernetes Error Reference

Quick lookup for error messages you’ll encounter in the wild.


Pod Errors

CrashLoopBackOff

What it is: Container starts and exits repeatedly. Kubernetes exponentially backs off restarts.

Immediate check:

kubectl logs <pod> --previous
kubectl describe pod <pod> | grep "Exit Code"

Causes:

  • Exit code 1: Application error (check logs)
  • Exit code 127: Binary not found in image
  • Exit code 137: OOMKilled (memory limit exceeded)
  • Exit code 139: Segfault

ImagePullBackOff / ErrImagePull

What it is: Kubelet can’t pull the container image.

Immediate check:

kubectl describe pod <pod> | grep -A 5 "Events"

Causes and fixes:

ErrorCauseFix
404 Not FoundTag doesn’t existVerify tag exists in registry
401 UnauthorizedNo credentialsAdd imagePullSecret
context deadline exceededNetwork timeoutCheck node → registry connectivity
name unknownWrong image nameCheck for typos in image path

OOMKilled

What it is: Container exceeded limits.memory. Kernel killed it with SIGKILL (exit code 137).

Immediate check:

kubectl describe pod <pod> | grep -A 3 "Last State"
kubectl top pod <pod>

Fix: Increase limits.memory or fix application memory leak.


Evicted

What it is: Node was under memory or disk pressure; pod was evicted to free resources.

Immediate check:

kubectl describe pod <pod> | grep "Reason"
kubectl describe node <node> | grep -A 5 "Conditions"

Fix: Check node disk/memory usage. Free up space or add nodes.


CreateContainerConfigError

What it is: Pod spec references a Secret or ConfigMap that doesn’t exist.

Immediate check:

kubectl describe pod <pod> | grep -A 5 "Events"

Common pattern:

Error: secret "db-credentials" not found

Fix: Create the missing Secret or ConfigMap before the pod.


Terminating (stuck)

What it is: Pod is stuck in Terminating state — usually because a finalizer isn’t clearing or the node is unreachable.

Fix:

# Force delete
kubectl delete pod <pod> --grace-period=0 --force

# If still stuck: remove finalizers
kubectl patch pod <pod> -p '{"metadata":{"finalizers":[]}}' --type=merge

Scheduling Errors

FailedScheduling: Insufficient CPU/Memory

0/3 nodes are available: 3 Insufficient cpu.

Fix: Reduce pod requests, remove unused pods, or add nodes.


FailedScheduling: node(s) had taints

0/1 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane:NoSchedule}

Fix: Add matching toleration to pod spec, or target a worker node.


FailedScheduling: didn’t match node affinity

0/3 nodes are available: 3 node(s) didn't match Pod's node affinity/selector.

Fix: Relax nodeAffinity rules or label a node with the required label.


Storage Errors

ProvisioningFailed: no provisioner found

no provisioner found for "fast-ssd" in storage classes

Fix: Use an existing StorageClass (kubectl get storageclass) or install the CSI driver.


Multi-Attach error for volume

Multi-Attach error: volume is already used by pod(s) on node worker-1

What it is: A ReadWriteOnce volume is still “attached” to a node while another pod on a different node tries to attach it.

Fix:

kubectl delete pod <old-pod> --grace-period=0 --force
# Wait 30-60 seconds for the volume to detach

MountVolume.MountDevice failed

What it is: Volume exists and is attached, but couldn’t be formatted or mounted on the node.

Check: Look at CSI node plugin logs:

kubectl logs -n kube-system -l app=csi-node-driver

Networking Errors

Service — No Endpoints

kubectl describe service my-svc
# Endpoints: <none>

What it is: Service selector labels don’t match any pod labels.

Fix:

kubectl get pods --show-labels     # what labels do pods have?
kubectl edit service my-svc        # fix the selector

DNS Resolution Failure

nslookup: can't resolve 'my-service': Name or service not known

Check:

kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns

Connection Refused

curl: (7) Failed to connect to 10.96.55.100 port 80: Connection refused

Causes:

  • Wrong port in Service (port vs targetPort)
  • App not listening on the port
  • Service selector doesn’t match pod

RBAC Errors

Forbidden: User cannot list resource

Error from server (Forbidden): pods is forbidden: User "jane" cannot list resource "pods"
in API group "" in the namespace "production"

Fix: Create a Role/ClusterRole with the required permissions and bind it to the user.

kubectl create role pod-reader \
  --verb=get,list,watch --resource=pods -n production
kubectl create rolebinding jane-pod-reader \
  --role=pod-reader --user=jane -n production

API Errors

No matches for kind X in version Y

error: unable to recognize "manifest.yaml": no matches for kind "Ingress" in version "extensions/v1beta1"

Fix: You’re using a deprecated API version. Update to the current stable version:

  • Ingress: use networking.k8s.io/v1 (not extensions/v1beta1)
  • CronJob: use batch/v1 (not batch/v1beta1)
# Check current API versions
kubectl api-resources | grep ingress
kubectl api-versions | grep networking

server-side apply: conflict

Apply failed with 1 conflict: conflict with "kubectl-client-side-apply":

Fix:

# Force override (take ownership of fields)
kubectl apply -f manifest.yaml --force-conflicts --server-side

Exit Code Reference

Exit CodeSignalMeaning
0Success
1General application error
2Misuse of shell command
125Container failed to run
126Command not executable
127Command not found
130SIGINT (2)Container interrupted (Ctrl+C)
137SIGKILL (9)OOMKilled or force-killed
139SIGSEGV (11)Segmentation fault
143SIGTERM (15)Graceful termination