In this page
This post is also available in: Español
A Kubernetes Pod can be Running while still being unavailable for traffic.
It can also be running a perfectly functional application and start restarting because a health check is incorrectly configured.
Kubernetes provides three primary types of probes:
Startup
Liveness
Readiness
Although their configuration looks similar, they answer very different questions.
Startup
|
v
Has the application finished starting?
Liveness
|
v
Is the application still functioning,
or should it be restarted?
Readiness
|
v
Is the application ready
to receive traffic?
Confusing these responsibilities can turn a small failure into a much larger incident.
In this fourth article of our Kubernetes Troubleshooting series, we will investigate:
readiness probes;
liveness probes;
startup probes;
HTTP, TCP, exec, and gRPC probes;
timeouts and thresholds;
Pods that are
RunningbutNotReady;probe-induced restart loops;
slow-starting applications;
external dependencies in health checks;
cascading failures;
how to design probes that help instead of making incidents worse.
First: the three probes do not do the same thing
This is the most important distinction in this article.
Container
|
+----------------+----------------+
| | |
v v v
Startup Liveness Readiness
| | |
Started? Alive? Can receive
traffic?
| | |
FAIL FAIL FAIL
| | |
v v v
Restart Restart Remove from
normal Service
traffic
A readiness failure does not normally restart the container.
A liveness failure can cause the kubelet to terminate the container and apply its restart policy.
A startup probe protects applications that require additional startup time.
1. Readiness probes
A readiness probe answers:
Is this container able to receive traffic right now?
Suppose we have three Pods:
Service
|
+--------+--------+
| | |
v v v
Pod A Pod B Pod C
Ready Ready NotReady
Normal Service traffic should go to endpoints considered ready.
Pod C can remain:
Running
while not being:
Ready
That is a valid and important Kubernetes state.
Example readiness probe
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
Conceptually:
GET /ready
|
+--- Success
| |
| v
| Ready
|
+--- Failure
|
v
NotReady
|
v
No normal Service traffic
The application process continues running.
Pod is Running but 0/1 Ready
A common situation is:
kubectl get pods -n production
returning:
NAME READY STATUS RESTARTS
api-7bd95fcb85-rx7mw 0/1 Running 0
This tells us:
Container
|
+--- Running
|
+--- Not Ready
Start with:
kubectl describe pod \
api-7bd95fcb85-rx7mw \
-n production
Look for Events such as:
Readiness probe failed:
HTTP probe failed with statuscode: 503
or:
dial tcp 10.244.2.15:8080:
connect: connection refused
What is the readiness probe actually checking?
Inspect the workload:
kubectl get deployment api \
-n production \
-o yaml
Find:
readinessProbe:
Then ask:
Correct path?
Correct port?
Correct protocol?
Is the application listening?
Is the timeout realistic?
Is the HTTP response valid?
Does the endpoint depend on other systems?
We can also reproduce the check manually:
kubectl exec \
-n production \
<pod> -- \
curl -v http://127.0.0.1:8080/ready
If the production image does not contain curl, use a debugging container or temporary debug Pod instead.
Readiness and EndpointSlices
When a Pod becomes not ready, Kubernetes reflects this in the backend information used by Services.
Inspect it:
kubectl get endpointslices \
-n production \
-l kubernetes.io/service-name=api-service \
-o yaml
Look for:
conditions:
ready: false
This connects health checks with the networking path we investigated in the previous article:
Readiness Probe
|
v
Pod NotReady
|
v
EndpointSlice condition changes
|
v
Normal Service traffic stops
2. Liveness probes
A liveness probe asks a different question:
Is the application still healthy enough to keep running?
For example:
livenessProbe:
httpGet:
path: /live
port: 8080
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
If the probe fails enough consecutive times:
Liveness fails
|
v
Failure threshold reached
|
v
kubelet stops container
|
v
Restart policy applies
|
v
Container starts again
This can recover applications from conditions such as certain deadlocks.
But it can also be dangerous.
When liveness becomes the problem
Imagine an application under heavy load.
Normally:
/live -> 20 ms
During a traffic spike:
/live -> 1.5 seconds
But the probe has:
timeoutSeconds: 1
failureThreshold: 3
The result can become:
Traffic increases
|
v
Application slows
|
v
Liveness times out
|
v
Pod restarts
|
v
Less capacity
|
v
Remaining Pods receive more traffic
|
v
They become slower
|
v
More liveness failures
|
v
More restartstxt
We have created a cascading failure.
The application started with a performance degradation.
The health-check configuration progressively removed capacity.
3. Startup probes
Startup probes are especially useful for applications that need significant initialization time.
For example:
Container starts
|
v
Load configuration
|
v
Initialize runtime
|
v
Warm caches
|
v
Application ready
Suppose this takes:
90 seconds
If liveness begins too early, Kubernetes can terminate the process while it is still legitimately starting.
Without a startup probe
Container starts
|
v
Application starting...
|
v
Liveness begins
|
v
FAIL
|
v
Restart
|
v
Application starting...
|
v
FAIL
Eventually we may see:
CrashLoopBackOff
even though the application might have successfully started if given enough time.
With a startup probe
startupProbe:
httpGet:
path: /startup
port: 8080
periodSeconds: 10
failureThreshold: 30
This provides an approximate maximum startup window of:
10 seconds × 30 failures = 300 seconds
While startup has not succeeded:
Startup probe
|
+--- checking
|
Liveness disabled
Readiness disabled
Once startup succeeds:
Startup succeeds
|
+------------+
| |
v v
Liveness Readiness
begins begins
initialDelaySeconds vs startupProbe
We could configure:
livenessProbe:
initialDelaySeconds: 120
but a startup probe usually communicates the intent more precisely for variable startup times.
A fixed delay means:
Wait 120 seconds
A startup probe means:
Check startup
|
+--- ready after 20 seconds
| |
| v
| continue
|
+--- still starting
|
v
keep checking
This protects slow startups without necessarily waiting for the maximum possible delay.
4. Troubleshooting probe-induced restart loops
Suppose:
kubectl get pods -n production
shows:
NAME READY STATUS RESTARTS
api-6b8cff746c-r27ft 0/1 CrashLoopBackOff 8
Do not immediately assume that the application process is crashing on its own.
Run:
kubectl describe pod \
api-6b8cff746c-r27ft \
-n production
and suppose we find:
Warning Unhealthy
Liveness probe failed
Our hypothesis changes immediately.
Check:
kubectl logs \
api-6b8cff746c-r27ft \
-n production \
--previous
and inspect the probe configuration:
kubectl get deployment api \
-n production \
-o yaml
Investigate the timeline
We want to determine:
Container starts
|
v
When does the probe start?
|
v
How long does startup take?
|
v
How frequently is it checked?
|
v
How long until failure?
|
v
Why does the endpoint fail?
For example:
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 3
periodSeconds: 2
timeoutSeconds: 1
failureThreshold: 2
may be far too aggressive for some applications.
5. Probe parameters we need to understand
initialDelaySeconds
Controls when probing begins after the container starts.
initialDelaySeconds: 10
For workloads with variable or long initialization, it should not automatically replace a proper startup probe.
periodSeconds
Controls how frequently the probe runs.
periodSeconds: 10
Running probes unnecessarily often can create additional overhead.
timeoutSeconds
Controls how long Kubernetes waits for the probe result.
timeoutSeconds: 2
An unrealistically short timeout can produce false failures under load.
failureThreshold
Controls the number of consecutive failures before Kubernetes considers the probe failed.
failureThreshold: 3
Together with periodSeconds, this determines much of the failure tolerance.
successThreshold
Controls how many consecutive successes are required after failure.
Readiness can use values greater than one.
For liveness and startup probes:
successThreshold = 1
is required.
6. HTTP probes
An HTTP probe can look like:
readinessProbe:
httpGet:
path: /ready
port: 8080
scheme: HTTP
We can reproduce it:
curl -v http://<pod-ip>:8080/ready
Investigate:
HTTP status
response time
connection errors
TLS
application logs
7. TCP probes
A TCP probe checks whether a TCP connection can be established.
livenessProbe:
tcpSocket:
port: 8080
Conceptually:
Can kubelet establish TCP connection?
|
+----+----+
| |
YES NO
But a successful connection does not prove that the application is logically healthy.
A service can accept:
TCP -> OK
while returning:
HTTP 500
for every request.
Choose the probe mechanism based on what you actually need to validate.
8. Exec probes
A probe can also execute a command inside the container:
livenessProbe:
exec:
command:
- /bin/sh
- -c
- test -f /tmp/healthy
This allows application-specific checks.
However, exec probes involve executing processes within the container.
Very frequent exec probes across dense clusters can create unnecessary overhead, so they should be used deliberately.
9. gRPC probes
Kubernetes also supports native probes for gRPC workloads.
For example:
livenessProbe:
grpc:
port: 50051
initialDelaySeconds: 10
This can avoid exposing an additional HTTP endpoint solely for Kubernetes health checks.
There are configuration differences compared with HTTP and TCP probes, however.
For example, gRPC probes do not support named ports in the same way HTTP and TCP probes do.
Always validate probe configuration against the protocol being used.
10. Should liveness check the database?
This is one of the most important health-check design questions.
Suppose:
/live
|
+--- application
+--- PostgreSQL
+--- Redis
+--- payment API
Now PostgreSQL has an outage.
PostgreSQL unavailable
|
v
/live fails
|
v
Every application Pod
fails liveness
|
v
Every Pod restarts
|
v
Application capacity collapses
Restarting application Pods does not repair PostgreSQL.
We are using liveness to react to a failure that restarting the local process cannot fix.
A better separation
A useful model is:
/live
|
+--- Is my process internally healthy?
/ready
|
+--- Can I currently serve useful traffic?
For example:
Database unavailable
|
v
Readiness fails
|
v
Pod removed from normal traffic
|
v
Container remains running
But readiness also requires careful design.
11. The danger of checking every dependency in readiness
Suppose the API uses:
API
|
+-- PostgreSQL
+-- Redis
+-- External Payment API
and readiness requires all three to be healthy.
The payment provider has a small outage:
Payment API unavailable
|
v
Every API Pod becomes NotReady
|
v
Service has zero ready backends
|
v
Entire API unavailable
But perhaps:
GET /products
GET /profile
GET /orders
could still work.
When designing readiness, ask:
Does failure of this dependency actually make the entire Pod unable to serve useful traffic?
The answer depends on the application architecture.
12. Probes should be cheap
A health endpoint should not execute an expensive database query every few seconds.
Avoid designs such as:
/readiness
|
v
Complex DB query
|
v
Multiple downstream calls
|
v
Storage operation
multiplied by:
500 Pods
×
probe every 5 seconds
The health-check infrastructure itself can become significant load.
A good probe should normally be:
fast
cheap
predictable
purpose-specific
13. Observe probe failures
Events are an excellent first signal:
kubectl get events \
-n production \
--sort-by='.lastTimestamp'
Look for:
Unhealthy
Readiness probe failed
Liveness probe failed
Startup probe failed
But Events are not historical monitoring.
Correlate:
Probe failures
+
Restarts
+
CPU
+
Memory
+
Latency
+
Traffic
+
Deployments
For example:
10:00 Deployment
10:03 Traffic increases
10:05 CPU saturation
10:06 Liveness failures
10:07 Restarts increase
10:08 Available replicas decrease
10:09 5xx spike
That is much more informative than simply seeing Unhealthy.
14. Logs
When a probe fails:
kubectl logs <pod> \
-n production
If the container restarted:
kubectl logs <pod> \
-n production \
--previous
Correlate timestamps between Events and application logs.
You might discover:
probe timeout
|
v
same timestamp
|
v
long GC pause
or:
readiness returns 503
|
v
connection pool exhausted
or:
startup failure
|
v
database migration still running
15. Metrics and traces
Probes indicate whether an instance is available.
They do not necessarily explain why the instance is degraded.
Suppose:
Readiness failures increasing
while:
CPU normal
Memory normal
Latency high
A distributed trace could reveal:
API 20 ms
|
+-- Inventory 15 ms
|
+-- Database 4.2 sec
Now we know which dependency deserves further investigation.
Health checks are one signal.
Logs, metrics, and traces provide the context.
16. Rolling updates and readiness
Readiness becomes especially important during deployments.
Old Pods
|
+--- Ready
New Pod starts
|
v
NotReady
|
v
Initialize
|
v
Ready
|
v
Receives traffic
Without an appropriate readiness check, a new application instance can begin receiving traffic before it is genuinely prepared to serve requests correctly.
This can produce failures that appear only during rollouts.
17. Readiness Gates
For more advanced scenarios, Kubernetes allows additional conditions to participate in Pod readiness through:
readinessGates:
Conceptually:
Containers Ready
+
Custom condition Ready
|
v
Pod Ready
This allows controllers or external conditions to influence readiness.
Most workloads do not need readiness gates, but they are useful when availability depends on something that cannot be accurately represented by a normal container probe.
18. Probe failures isolated to one node
Suppose:
node-01
Pods -> probes OK
node-02
Pods -> probes FAIL
Now the application itself may not be the problem.
Check placement:
kubectl get pods \
-n production \
-o wide
Then:
kubectl describe node node-02
Possible causes include:
node CPU saturation
network failure
CNI problem
kubelet problem
disk pressure
packet loss
Blast radius is again an important diagnostic signal.
19. What to inspect with kubectl describe
During troubleshooting:
kubectl describe pod <pod> \
-n <namespace>
inspect:
State
Last State
Ready
Restart Count
Conditions
Liveness
Readiness
Startup
Events
Then:
kubectl get pod <pod> \
-n <namespace> \
-o yaml
to inspect the complete probe configuration.
20. A practical configuration
A workload might use:
startupProbe:
httpGet:
path: /startup
port: 8080
periodSeconds: 5
failureThreshold: 30
livenessProbe:
httpGet:
path: /live
port: 8080
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: 8080
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
But these values are not universal recommendations.
They should be derived from actual data about:
startup time
normal response time
load behavior
recovery time
dependency behavior
traffic patterns
Copying another service's probe configuration without understanding those characteristics can create new problems.
Troubleshooting flow
Pod problem
|
v
Is container Running?
|
v
Is Pod Ready?
|
+----------+----------+
| |
NO YES
| |
v v
Readiness Events Restarting?
| |
v v
Probe endpoint works? Liveness / Startup?
| |
v v
Path / port / timeout Probe timing
dependency / overload startup duration
| |
+----------+----------+
|
v
Metrics + Logs + Events
|
v
Check dependencies
|
v
Root cause
|
v
Change
|
v
Validate
Command reference
# Pod state
kubectl get pods \
-n <namespace>
kubectl get pods \
-n <namespace> \
-o wide
# Pod details and probe Events
kubectl describe pod <pod> \
-n <namespace>
# Complete configuration
kubectl get pod <pod> \
-n <namespace> \
-o yaml
kubectl get deployment <deployment> \
-n <namespace> \
-o yaml
# Logs
kubectl logs <pod> \
-n <namespace>
kubectl logs <pod> \
-n <namespace> \
--previous
# Events
kubectl get events \
-n <namespace> \
--sort-by='.lastTimestamp'
# EndpointSlices
kubectl get endpointslices \
-n <namespace> \
-l kubernetes.io/service-name=<service> \
-o yaml
# Resource metrics
kubectl top pods \
-n <namespace>
kubectl top pod <pod> \
-n <namespace> \
--containers
# Deployment history
kubectl rollout history \
deployment/<deployment> \
-n <namespace>
# Test endpoints
curl -v http://<pod-ip>:<port>/ready
curl -v http://<pod-ip>:<port>/live
# Debug when application image lacks tools
kubectl debug <pod> \
-n <namespace> \
-it \
--image=ubuntu
What not to do
Avoid automatically responding with:
Readiness failing
|
v
Restart Pod
or:
Liveness failing
|
v
Increase timeout
or:
Slow startup
|
v
Set initialDelaySeconds = 300
Understand the failure first.
Failure
|
v
Understand probe purpose
|
v
Reproduce probe
|
v
Correlate Events
|
v
Check metrics + logs
|
v
Identify application/dependency issue
|
v
Adjust probe or application
|
v
Validate under realistic load
Conclusion
Kubernetes probes are not simply health checks.
They can directly change cluster behavior.
Readiness failure
|
v
Traffic changes
Liveness failure
|
v
Container lifecycle changes
Startup failure
|
v
Container lifecycle changes
That is why a poorly designed probe can be dangerous.
Each probe should answer one specific question:
Startup:
Has the application finished starting?
Liveness:
Can restarting this process recover it?
Readiness:
Can this instance serve useful traffic now?
Before adding any dependency to a health check, ask:
What will Kubernetes do when this dependency fails, and will that action actually help?
That question prevents many dangerous probe configurations.
In the next article we will troubleshoot Kubernetes Nodes:
NotReady;MemoryPressure;DiskPressure;PIDPressure;kubelet;
container runtime;
node networking;
failures concentrated on a single node;
determining when an incident has moved from the workload layer to the host layer.
Keep reading
Related articles
Dynamic Jenkins Parameters with Groovy: GitHub Branches, Builds, and Dependencies
Learn how to create dynamic Jenkins parameters with Groovy and shared libraries, retrieve GitHub branches, select builds from other pipelines, and define parameter dependencies with Active Choices. Includes practical Jenkinsfile examples and security best
AWS, Azure, GCP & OCI Weekly Cloud Updates — September 15–20, 2026
Explore this week's AWS, Azure, Google Cloud and OCI updates, including Kubernetes, database recovery, Prometheus compatibility, cloud security and networking.
Kubernetes Troubleshooting: Networking, Services, DNS, NetworkPolicy, Ingress y CNI
Learn how to troubleshoot Kubernetes networking from Pods and Services through EndpointSlices, DNS, NetworkPolicy, Ingress, Gateway API, LoadBalancers, and CNI.
Comments