Kubernetes Troubleshooting: Health Checks, Readiness, Liveness y Startup Probes

Learn how to troubleshoot Kubernetes readiness, liveness, and startup probes, including timeouts, restart loops, health checks, and cascading failures.

This post is also available in: Español

A Kubernetes Pod can be Running while still being unavailable for traffic.

It can also be running a perfectly functional application and start restarting because a health check is incorrectly configured.

Kubernetes provides three primary types of probes:

Startup
Liveness
Readiness

Although their configuration looks similar, they answer very different questions.

Startup
|
v
Has the application finished starting?
Liveness
|
v
Is the application still functioning,
or should it be restarted?
Readiness
|
v
Is the application ready
to receive traffic?

Confusing these responsibilities can turn a small failure into a much larger incident.

In this fourth article of our Kubernetes Troubleshooting series, we will investigate:

  • readiness probes;

  • liveness probes;

  • startup probes;

  • HTTP, TCP, exec, and gRPC probes;

  • timeouts and thresholds;

  • Pods that are Running but NotReady;

  • probe-induced restart loops;

  • slow-starting applications;

  • external dependencies in health checks;

  • cascading failures;

  • how to design probes that help instead of making incidents worse.


First: the three probes do not do the same thing

This is the most important distinction in this article.

Container
|
+----------------+----------------+
| | |
v v v
Startup Liveness Readiness
| | |
Started? Alive? Can receive
traffic?
| | |
FAIL FAIL FAIL
| | |
v v v
Restart Restart Remove from
normal Service
traffic

A readiness failure does not normally restart the container.

A liveness failure can cause the kubelet to terminate the container and apply its restart policy.

A startup probe protects applications that require additional startup time.


1. Readiness probes

A readiness probe answers:

Is this container able to receive traffic right now?

Suppose we have three Pods:

Service
|
+--------+--------+
| | |
v v v
Pod A Pod B Pod C
Ready Ready NotReady

Normal Service traffic should go to endpoints considered ready.

Pod C can remain:

Running

while not being:

Ready

That is a valid and important Kubernetes state.


Example readiness probe

readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3

Conceptually:

GET /ready
|
+--- Success
| |
| v
| Ready
|
+--- Failure
|
v
NotReady
|
v
No normal Service traffic

The application process continues running.


Pod is Running but 0/1 Ready

A common situation is:

kubectl get pods -n production

returning:

NAME READY STATUS RESTARTS
api-7bd95fcb85-rx7mw 0/1 Running 0

This tells us:

Container
|
+--- Running
|
+--- Not Ready

Start with:

kubectl describe pod \
api-7bd95fcb85-rx7mw \
-n production

Look for Events such as:

Readiness probe failed:
HTTP probe failed with statuscode: 503

or:

dial tcp 10.244.2.15:8080:
connect: connection refused

What is the readiness probe actually checking?

Inspect the workload:

kubectl get deployment api \
-n production \
-o yaml

Find:

readinessProbe:

Then ask:

Correct path?
Correct port?
Correct protocol?
Is the application listening?
Is the timeout realistic?
Is the HTTP response valid?
Does the endpoint depend on other systems?

We can also reproduce the check manually:

kubectl exec \
-n production \
<pod> -- \
curl -v http://127.0.0.1:8080/ready

If the production image does not contain curl, use a debugging container or temporary debug Pod instead.


Readiness and EndpointSlices

When a Pod becomes not ready, Kubernetes reflects this in the backend information used by Services.

Inspect it:

kubectl get endpointslices \
-n production \
-l kubernetes.io/service-name=api-service \
-o yaml

Look for:

conditions:
ready: false

This connects health checks with the networking path we investigated in the previous article:

Readiness Probe
|
v
Pod NotReady
|
v
EndpointSlice condition changes
|
v
Normal Service traffic stops

2. Liveness probes

A liveness probe asks a different question:

Is the application still healthy enough to keep running?

For example:

livenessProbe:
httpGet:
path: /live
port: 8080
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3

If the probe fails enough consecutive times:

Liveness fails
|
v
Failure threshold reached
|
v
kubelet stops container
|
v
Restart policy applies
|
v
Container starts again

This can recover applications from conditions such as certain deadlocks.

But it can also be dangerous.


When liveness becomes the problem

Imagine an application under heavy load.

Normally:

/live -> 20 ms

During a traffic spike:

/live -> 1.5 seconds

But the probe has:

timeoutSeconds: 1
failureThreshold: 3

The result can become:

Traffic increases
|
v
Application slows
|
v
Liveness times out
|
v
Pod restarts
|
v
Less capacity
|
v
Remaining Pods receive more traffic
|
v
They become slower
|
v
More liveness failures
|
v
More restartstxt

We have created a cascading failure.

The application started with a performance degradation.

The health-check configuration progressively removed capacity.


3. Startup probes

Startup probes are especially useful for applications that need significant initialization time.

For example:

Container starts
|
v
Load configuration
|
v
Initialize runtime
|
v
Warm caches
|
v
Application ready

Suppose this takes:

90 seconds

If liveness begins too early, Kubernetes can terminate the process while it is still legitimately starting.


Without a startup probe

Container starts
|
v
Application starting...
|
v
Liveness begins
|
v
FAIL
|
v
Restart
|
v
Application starting...
|
v
FAIL

Eventually we may see:

CrashLoopBackOff

even though the application might have successfully started if given enough time.


With a startup probe

startupProbe:
httpGet:
path: /startup
port: 8080
periodSeconds: 10
failureThreshold: 30

This provides an approximate maximum startup window of:

10 seconds × 30 failures = 300 seconds

While startup has not succeeded:

Startup probe
|
+--- checking
|
Liveness disabled
Readiness disabled

Once startup succeeds:

Startup succeeds
|
+------------+
| |
v v
Liveness Readiness
begins begins

initialDelaySeconds vs startupProbe

We could configure:

livenessProbe:
initialDelaySeconds: 120

but a startup probe usually communicates the intent more precisely for variable startup times.

A fixed delay means:

Wait 120 seconds

A startup probe means:

Check startup
|
+--- ready after 20 seconds
| |
| v
| continue
|
+--- still starting
|
v
keep checking

This protects slow startups without necessarily waiting for the maximum possible delay.


4. Troubleshooting probe-induced restart loops

Suppose:

kubectl get pods -n production

shows:

NAME READY STATUS RESTARTS
api-6b8cff746c-r27ft 0/1 CrashLoopBackOff 8

Do not immediately assume that the application process is crashing on its own.

Run:

kubectl describe pod \
api-6b8cff746c-r27ft \
-n production

and suppose we find:

Warning Unhealthy
Liveness probe failed

Our hypothesis changes immediately.

Check:

kubectl logs \
api-6b8cff746c-r27ft \
-n production \
--previous

and inspect the probe configuration:

kubectl get deployment api \
-n production \
-o yaml

Investigate the timeline

We want to determine:

Container starts
|
v
When does the probe start?
|
v
How long does startup take?
|
v
How frequently is it checked?
|
v
How long until failure?
|
v
Why does the endpoint fail?

For example:

livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 3
periodSeconds: 2
timeoutSeconds: 1
failureThreshold: 2

may be far too aggressive for some applications.


5. Probe parameters we need to understand

initialDelaySeconds

Controls when probing begins after the container starts.

initialDelaySeconds: 10

For workloads with variable or long initialization, it should not automatically replace a proper startup probe.


periodSeconds

Controls how frequently the probe runs.

periodSeconds: 10

Running probes unnecessarily often can create additional overhead.


timeoutSeconds

Controls how long Kubernetes waits for the probe result.

timeoutSeconds: 2

An unrealistically short timeout can produce false failures under load.


failureThreshold

Controls the number of consecutive failures before Kubernetes considers the probe failed.

failureThreshold: 3

Together with periodSeconds, this determines much of the failure tolerance.


successThreshold

Controls how many consecutive successes are required after failure.

Readiness can use values greater than one.

For liveness and startup probes:

successThreshold = 1

is required.


6. HTTP probes

An HTTP probe can look like:

readinessProbe:
httpGet:
path: /ready
port: 8080
scheme: HTTP

We can reproduce it:

curl -v http://<pod-ip>:8080/ready

Investigate:

HTTP status
response time
connection errors
TLS
application logs

7. TCP probes

A TCP probe checks whether a TCP connection can be established.

livenessProbe:
tcpSocket:
port: 8080

Conceptually:

Can kubelet establish TCP connection?
|
+----+----+
| |
YES NO

But a successful connection does not prove that the application is logically healthy.

A service can accept:

TCP -> OK

while returning:

HTTP 500

for every request.

Choose the probe mechanism based on what you actually need to validate.


8. Exec probes

A probe can also execute a command inside the container:

livenessProbe:
exec:
command:
- /bin/sh
- -c
- test -f /tmp/healthy

This allows application-specific checks.

However, exec probes involve executing processes within the container.

Very frequent exec probes across dense clusters can create unnecessary overhead, so they should be used deliberately.


9. gRPC probes

Kubernetes also supports native probes for gRPC workloads.

For example:

livenessProbe:
grpc:
port: 50051
initialDelaySeconds: 10

This can avoid exposing an additional HTTP endpoint solely for Kubernetes health checks.

There are configuration differences compared with HTTP and TCP probes, however.

For example, gRPC probes do not support named ports in the same way HTTP and TCP probes do.

Always validate probe configuration against the protocol being used.


10. Should liveness check the database?

This is one of the most important health-check design questions.

Suppose:

/live
|
+--- application
+--- PostgreSQL
+--- Redis
+--- payment API

Now PostgreSQL has an outage.

PostgreSQL unavailable
|
v
/live fails
|
v
Every application Pod
fails liveness
|
v
Every Pod restarts
|
v
Application capacity collapses

Restarting application Pods does not repair PostgreSQL.

We are using liveness to react to a failure that restarting the local process cannot fix.


A better separation

A useful model is:

/live
|
+--- Is my process internally healthy?
/ready
|
+--- Can I currently serve useful traffic?

For example:

Database unavailable
|
v
Readiness fails
|
v
Pod removed from normal traffic
|
v
Container remains running

But readiness also requires careful design.


11. The danger of checking every dependency in readiness

Suppose the API uses:

API
|
+-- PostgreSQL
+-- Redis
+-- External Payment API

and readiness requires all three to be healthy.

The payment provider has a small outage:

Payment API unavailable
|
v
Every API Pod becomes NotReady
|
v
Service has zero ready backends
|
v
Entire API unavailable

But perhaps:

GET /products
GET /profile
GET /orders

could still work.

When designing readiness, ask:

Does failure of this dependency actually make the entire Pod unable to serve useful traffic?

The answer depends on the application architecture.


12. Probes should be cheap

A health endpoint should not execute an expensive database query every few seconds.

Avoid designs such as:

/readiness
|
v
Complex DB query
|
v
Multiple downstream calls
|
v
Storage operation

multiplied by:

500 Pods
×
probe every 5 seconds

The health-check infrastructure itself can become significant load.

A good probe should normally be:

fast
cheap
predictable
purpose-specific

13. Observe probe failures

Events are an excellent first signal:

kubectl get events \
-n production \
--sort-by='.lastTimestamp'

Look for:

Unhealthy
Readiness probe failed
Liveness probe failed
Startup probe failed

But Events are not historical monitoring.

Correlate:

Probe failures
+
Restarts
+
CPU
+
Memory
+
Latency
+
Traffic
+
Deployments

For example:

10:00 Deployment
10:03 Traffic increases
10:05 CPU saturation
10:06 Liveness failures
10:07 Restarts increase
10:08 Available replicas decrease
10:09 5xx spike

That is much more informative than simply seeing Unhealthy.


14. Logs

When a probe fails:

kubectl logs <pod> \
-n production

If the container restarted:

kubectl logs <pod> \
-n production \
--previous

Correlate timestamps between Events and application logs.

You might discover:

probe timeout
|
v
same timestamp
|
v
long GC pause

or:

readiness returns 503
|
v
connection pool exhausted

or:

startup failure
|
v
database migration still running

15. Metrics and traces

Probes indicate whether an instance is available.

They do not necessarily explain why the instance is degraded.

Suppose:

Readiness failures increasing

while:

CPU normal
Memory normal
Latency high

A distributed trace could reveal:

API 20 ms
|
+-- Inventory 15 ms
|
+-- Database 4.2 sec

Now we know which dependency deserves further investigation.

Health checks are one signal.

Logs, metrics, and traces provide the context.


16. Rolling updates and readiness

Readiness becomes especially important during deployments.

Old Pods
|
+--- Ready
New Pod starts
|
v
NotReady
|
v
Initialize
|
v
Ready
|
v
Receives traffic

Without an appropriate readiness check, a new application instance can begin receiving traffic before it is genuinely prepared to serve requests correctly.

This can produce failures that appear only during rollouts.


17. Readiness Gates

For more advanced scenarios, Kubernetes allows additional conditions to participate in Pod readiness through:

readinessGates:

Conceptually:

Containers Ready
+
Custom condition Ready
|
v
Pod Ready

This allows controllers or external conditions to influence readiness.

Most workloads do not need readiness gates, but they are useful when availability depends on something that cannot be accurately represented by a normal container probe.


18. Probe failures isolated to one node

Suppose:

node-01
Pods -> probes OK
node-02
Pods -> probes FAIL

Now the application itself may not be the problem.

Check placement:

kubectl get pods \
-n production \
-o wide

Then:

kubectl describe node node-02

Possible causes include:

node CPU saturation
network failure
CNI problem
kubelet problem
disk pressure
packet loss

Blast radius is again an important diagnostic signal.


19. What to inspect with kubectl describe

During troubleshooting:

kubectl describe pod <pod> \
-n <namespace>

inspect:

State
Last State
Ready
Restart Count
Conditions
Liveness
Readiness
Startup
Events

Then:

kubectl get pod <pod> \
-n <namespace> \
-o yaml

to inspect the complete probe configuration.


20. A practical configuration

A workload might use:

startupProbe:
httpGet:
path: /startup
port: 8080
periodSeconds: 5
failureThreshold: 30
livenessProbe:
httpGet:
path: /live
port: 8080
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: 8080
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3

But these values are not universal recommendations.

They should be derived from actual data about:

startup time
normal response time
load behavior
recovery time
dependency behavior
traffic patterns

Copying another service's probe configuration without understanding those characteristics can create new problems.


Troubleshooting flow

Pod problem
|
v
Is container Running?
|
v
Is Pod Ready?
|
+----------+----------+
| |
NO YES
| |
v v
Readiness Events Restarting?
| |
v v
Probe endpoint works? Liveness / Startup?
| |
v v
Path / port / timeout Probe timing
dependency / overload startup duration
| |
+----------+----------+
|
v
Metrics + Logs + Events
|
v
Check dependencies
|
v
Root cause
|
v
Change
|
v
Validate

Command reference

# Pod state
kubectl get pods \
-n <namespace>
kubectl get pods \
-n <namespace> \
-o wide
# Pod details and probe Events
kubectl describe pod <pod> \
-n <namespace>
# Complete configuration
kubectl get pod <pod> \
-n <namespace> \
-o yaml
kubectl get deployment <deployment> \
-n <namespace> \
-o yaml
# Logs
kubectl logs <pod> \
-n <namespace>
kubectl logs <pod> \
-n <namespace> \
--previous
# Events
kubectl get events \
-n <namespace> \
--sort-by='.lastTimestamp'
# EndpointSlices
kubectl get endpointslices \
-n <namespace> \
-l kubernetes.io/service-name=<service> \
-o yaml
# Resource metrics
kubectl top pods \
-n <namespace>
kubectl top pod <pod> \
-n <namespace> \
--containers
# Deployment history
kubectl rollout history \
deployment/<deployment> \
-n <namespace>
# Test endpoints
curl -v http://<pod-ip>:<port>/ready
curl -v http://<pod-ip>:<port>/live
# Debug when application image lacks tools
kubectl debug <pod> \
-n <namespace> \
-it \
--image=ubuntu

What not to do

Avoid automatically responding with:

Readiness failing
|
v
Restart Pod

or:

Liveness failing
|
v
Increase timeout

or:

Slow startup
|
v
Set initialDelaySeconds = 300

Understand the failure first.

Failure
|
v
Understand probe purpose
|
v
Reproduce probe
|
v
Correlate Events
|
v
Check metrics + logs
|
v
Identify application/dependency issue
|
v
Adjust probe or application
|
v
Validate under realistic load

Conclusion

Kubernetes probes are not simply health checks.

They can directly change cluster behavior.

Readiness failure
|
v
Traffic changes
Liveness failure
|
v
Container lifecycle changes
Startup failure
|
v
Container lifecycle changes

That is why a poorly designed probe can be dangerous.

Each probe should answer one specific question:

Startup:
Has the application finished starting?
Liveness:
Can restarting this process recover it?
Readiness:
Can this instance serve useful traffic now?

Before adding any dependency to a health check, ask:

What will Kubernetes do when this dependency fails, and will that action actually help?

That question prevents many dangerous probe configurations.

In the next article we will troubleshoot Kubernetes Nodes:

  • NotReady;

  • MemoryPressure;

  • DiskPressure;

  • PIDPressure;

  • kubelet;

  • container runtime;

  • node networking;

  • failures concentrated on a single node;

  • determining when an incident has moved from the workload layer to the host layer.

Share this article

Keep reading

Related articles

DevOps

Dynamic Jenkins Parameters with Groovy: GitHub Branches, Builds, and Dependencies

Learn how to create dynamic Jenkins parameters with Groovy and shared libraries, retrieve GitHub branches, select builds from other pipelines, and define parameter dependencies with Active Choices. Includes practical Jenkinsfile examples and security best

AWS

AWS, Azure, GCP & OCI Weekly Cloud Updates — September 15–20, 2026

Explore this week's AWS, Azure, Google Cloud and OCI updates, including Kubernetes, database recovery, Prometheus compatibility, cloud security and networking.

DevOps

Kubernetes Troubleshooting: Networking, Services, DNS, NetworkPolicy, Ingress y CNI

Learn how to troubleshoot Kubernetes networking from Pods and Services through EndpointSlices, DNS, NetworkPolicy, Ingress, Gateway API, LoadBalancers, and CNI.

Comments