常见的三种探针是:
可以先用一句话理解:
| 探针 | 关注问题 | 失败后的主要动作 |
|---|---|---|
| Startup Probe | 应用启动完成了吗? | 重启容器 |
| Liveness Probe | 应用还活着吗? | 重启容器 |
| Readiness Probe | 应用能接收流量吗? | 从 Service 后端摘除,不接收流量 |
仅仅判断容器进程是否存在,并不能说明应用是正常的。
例如:
容器进程还在
但是:
- 应用死锁
- HTTP 服务无法响应
- 数据库连接池耗尽
- 应用还在启动
- 依赖服务未就绪
- 只能处理健康检查,无法处理业务请求如果没有探针,Kubernetes 可能认为容器是正常的。
探针就是让应用告诉 Kubernetes:
我是否已经启动?
我是否还活着?
我是否可以接收业务流量?Startup Probe 用来判断:
应用是否已经完成启动。
它特别适合启动时间比较长的应用,例如:
假设一个 Java 应用启动需要 120 秒。
如果只配置 Liveness Probe:
那么应用还没启动完成,探针就开始检查:
第 10 秒:检查失败
第 20 秒:检查失败
第 30 秒:检查失败
达到 failureThreshold
容器被重启结果:
应用还没启动完成,就被 Kubernetes 重启并且重启后又重新开始启动,形成:
启动 → 探针失败 → 重启 → 启动 → 探针失败 → 重启这叫做启动失败循环。
Startup Probe 的作用是:
在应用启动成功之前,暂时不要执行 Liveness Probe 和 Readiness Probe。
如果配置了 Startup Probe:
容器启动
↓
只执行 Startup Probe
如果 Startup Probe 一直失败:
达到 failureThreshold
↓
容器被重启因此,Startup Probe 主要解决的是:
应用启动比较慢,但启动完成后运行比较稳定。
含义:
注意:
periodSeconds × failureThreshold可以粗略估算启动允许时间,但实际还会受到请求耗时和调度时间影响。
Liveness Probe 用来判断:
容器中的应用是否仍然处于可恢复运行状态。
如果 Liveness Probe 持续失败,Kubernetes 会认为应用已经失去活性,并重启容器。
例如应用进程没有退出,但已经:
此时 Kubernetes 可以通过重启容器自动恢复。
Liveness Probe 失败
↓
达到 failureThreshold
注意:
Liveness Probe 失败通常是重启容器,不一定是删除 Pod。
如果 Pod 属于 Deployment,容器重启次数会增加,但 Pod 名称可能不变。
查看重启次数:
kubectl get pod查看详情:
kubectl describe pod <pod-name>含义:
successThreshold 必须为 1Readiness Probe 用来判断:
当前 Pod 是否已经准备好接收业务流量。
Readiness Probe 失败时,Kubernetes 通常不会重启容器,而是:
Pod 仍然运行
↓
从 Service 的 Endpoints 中移除
↓
Service 不再向该 Pod 转发新请求例如:
这种情况下,应用不一定需要重启,只需要暂时停止接收流量。
Readiness Probe 失败
↓
Pod 标记为 NotReady
容器通常仍然继续运行。
查看 Pod 状态:
kubectl get pod可能显示:
READY STATUS
0/1 Running这表示:
回答:
应用启动完成了吗?失败后:
重启容器回答:
应用是不是已经卡死或失去恢复能力?失败后:
重启容器回答:
应用现在能不能接收业务请求?失败后:
从 Service 后端摘除| 特性 | Startup | Liveness | Readiness |
|---|---|---|---|
| 判断目标 | 是否启动完成 | 是否仍然存活 | 是否可以接收流量 |
| 失败后是否重启 | 是 | 是 | 否 |
| 失败后是否摘除流量 | 间接 | 通常会,因为容器重启 | 是 |
| 启动阶段是否执行 | 最先执行 | Startup 成功后执行 | Startup 成功后执行 |
| 常见用途 | 慢启动应用 | 死锁、卡死恢复 | 流量控制 |
successThreshold | 必须为 1 | 必须为 1 | 可以大于 1 |
Kubernetes 常用以下几种探针检测方式:
HTTP 探针向容器发送 HTTP 请求,根据返回状态码判断是否成功。
示例:
httpGet:
path: /healthz
port: 8080默认情况下:
HTTP 状态码 200 到 399:成功
其他状态码:失败有些应用需要特定请求头,可以使用这种方式。
注意:
容器端口可以命名:
ports:
- name: http
containerPort: 8080探针引用:
readinessProbe:
这样比直接写数字更容易维护。
TCP 探针尝试连接指定端口。
如果 TCP 连接成功,则认为探针成功。
tcpSocket:
port: 3306适合:
TCP 连接成功只能说明:
端口有人监听不能说明:
应用业务正常
数据库可以执行查询
Redis 可以读写
服务可以处理有效请求例如:
进程仍然监听 8080
但内部线程池全部死锁TCP 探针可能仍然成功。
所以对于 HTTP 服务,通常优先选择具有业务含义的 HTTP 健康接口。
Exec 探针在容器内部执行命令。
exec:
command:
- cat
- /tmp/healthy命令退出码为 0:
探针成功非 0:
探针失败容器内如果存在:
/tmp/healthy则探针成功。
如果 redis-cli ping 返回成功,则 Pod 就绪。
注意:
Exec 探针会在容器内启动进程,可能产生额外开销。
如果探针执行频率太高,例如:
periodSeconds: 1可能导致:
因此,能用 HTTP 探针时,通常优先使用 HTTP 探针。
Kubernetes 支持对实现 gRPC Health Checking Protocol 的服务进行探测。
示例:
也可以指定服务名:
readinessProbe
应用需要实现标准 gRPC 健康检查接口。
注意:
port 必须是数字端口下面是一个同时配置三种探针的例子:
三个接口可以有不同含义:
/health/startup:应用是否完成启动
/health/live:应用是否还活着
/health/ready:应用是否可以接收流量Spring Boot 使用 Actuator 时,可以配置:
management.endpoint.health.probes.enabled=true
management.endpoints.web.exposure.include=health常见接口:
/actuator/health/liveness
/actuator/health/readinessKubernetes 配置:
生产中不建议简单把所有外部依赖都放入 Liveness 检查。
例如:
数据库暂时不可用这更可能意味着应用暂时“不应接收流量”,而不是应用本身已经死掉。
通常可以:
initialDelaySeconds: 30容器启动后等待多少秒,再开始第一次探测。
不过配置 Startup Probe 后,通常不需要过度依赖 initialDelaySeconds。
periodSeconds: 10两次探测之间的间隔,单位是秒。
常见建议:
具体要根据应用特点调整。
timeoutSeconds: 3单次探测最多等待多长时间。
如果健康接口本身只需要几十毫秒,通常不需要设置很大。
但也不要设置得过小,例如:
timeoutSeconds: 1在高负载或网络抖动时可能产生误判。
failureThreshold: 3连续失败多少次后认为探针失败。
例如:
periodSeconds: 10
failureThreshold: 3理论上大约连续失败 30 秒后触发动作。
对于 Liveness,不建议设置得过于敏感,否则短暂抖动就会重启应用。
successThreshold: 2连续成功多少次后认为探针恢复成功。
限制:
Readiness 使用大于 1 可以减少短暂抖动造成的流量反复切换。
可以在探针级别指定终止宽限时间:
容器因 Liveness Probe 失败而被重启时,Kubernetes 会尝试优雅终止。
一般优先使用 Pod 级别的:
spec:
terminationGracePeriodSeconds: 30Readiness Probe 会影响 Service 的后端列表。
例如:
Pod A:Ready
Pod B:NotReady
Pod C:ReadyService 通常只会把流量转发给:
Pod A
Pod C查看 Endpoints:
kubectl get endpoints <service-name>较新的 Kubernetes 也可以查看:
kubectl get endpointslice如果所有 Pod 的 Readiness 都失败:
Service 没有可用后端客户端可能收到:
Connection refused
503 Service UnavailableDeployment 滚动更新时,Readiness Probe 非常重要。
过程大致是:
如果没有正确配置 Readiness Probe,Deployment 可能认为新 Pod 已经可以接收流量,导致:
Pod 被删除时,通常会经历:
应用应该能够正确处理:
SIGTERM例如:
可以配合:
但不建议只依赖 sleep,更重要的是应用本身要实现优雅关闭。
除了内置 Readiness Probe,Pod 还可以使用自定义 Readiness Gate。
它允许外部控制器向 Pod 增加额外的就绪条件。
示例:
spec:
readinessGates:
- conditionType: example.com/readyPod 只有同时满足:
容器 Readiness Probe 成功
+
自定义 condition 为 True才会被认为是 Ready。
适合:
这是比较进阶的用法。
例如:
Liveness Probe 检查数据库当数据库短暂故障时:
更合理的方式通常是:
Readiness 检查数据库连接
Liveness 只检查应用自身是否还能工作不建议健康接口每次都执行:
健康接口应该:
应用启动需要 2 分钟,但配置:
initialDelaySeconds: 5
failureThreshold: 3
periodSeconds: 大约 20 秒后就可能重启,导致应用永远启动不了。
改用 Startup Probe:
如果应用已经不能正常处理业务,但 Readiness 仍返回 200:
Pod 仍在 Service 后端
客户端继续访问
请求持续失败Readiness 必须反映应用当前是否真的可以处理流量。
容器监听:
8080探针却检查:
port: 80结果必然失败。
可以检查:
kubectl exec -it <pod-name> -- sh然后测试:
wget -qO- http://127.0.0.1:8080/healthz容器内的 HTTP 探针通常从 Pod 网络命名空间访问。
应用应该监听:
0.0.0.0而不是只监听某个不可访问地址。
kubectl describe pod <pod-name>常见事件:
Liveness probe failed
Readiness probe failed
Startup probe failedkubectl logs <pod-name>如果容器已经重启过:
kubectl logs <pod-name> --previous进入容器:
kubectl exec -it <pod-name> -- sh测试 HTTP:
wget -S -O- http://127.0.0.1:8080/healthz如果镜像中有 curl:
curl -v http://127.0.0.1:8080/healthz测试端口:
nc -zv 127.0.0.1 8080kubectl get pod <pod-name> -o wide重点关注:
READY
RESTARTS
STATUS例如:
READY STATUS RESTARTS
0/1 Running 5可能表示:
可以只配置:
前提是应用实现 gRPC Health Checking Protocol。
脚本要:
这个示例实现了:
Liveness 失败会重启容器,配置过于严格会造成:
短暂抖动 → 重启
网络延迟 → 重启
数据库短暂故障 → 重启Readiness 失败只会摘除流量,通常比重启更安全。
因此可以让 Readiness 更快反映应用状态。
如果应用启动时间不稳定或较长,优先使用 Startup Probe,而不是把 Liveness 的 initialDelaySeconds 设置得特别大。
推荐分别设计:
/live:应用本身是否还能运行
/ready:应用是否可以接收流量
/startup:应用是否完成初始化不要让三个探针都简单访问同一个接口,却表达不同含义。
健康检查应该:
最重要的记忆方式是:
Startup:
你启动好了吗?
Liveness:
失败后的行为:
Startup 失败:
重启容器
常见组合:
生产环境中最推荐的思路是:
Startup 负责启动保护
Liveness 负责故障自愈
Readiness 负责流量控制还有一些探针的进阶细节,但核心内容已经讲完了。补充几个生产中很实用的点:
探针配置在 Pod/Container 上,通常访问:
Pod IP + 容器端口它不会经过:
所以探针正常,不代表 Service、Ingress 或外部网络一定正常。
HTTP、TCP、Exec 探针由节点上的 kubelet 执行。不同网络插件实现可能略有差异,但不能简单把探针当作普通外部客户端流量。
如果健康检查依赖其他服务,还要考虑:
每次探测都会产生资源开销:
不要把探测周期设置得过短,例如:
periodSeconds: 1尤其不要让健康接口执行复杂 SQL 或大量远程调用。
可能是:
timeoutSeconds 太小排查时不要只看 RestartCount,还要看:
terminationGracePeriodSeconds 配合优雅退出spec:
terminationGracePeriodSeconds: 30应用收到 SIGTERM 后,应当:
否则即使 Readiness 配置正确,仍可能出现请求中断或数据丢失。
如果应用遇到短暂问题,推荐:
Readiness 失败而不是:
Liveness 失败例如数据库短暂不可用时:
通常前者更安全,避免所有副本同时重启造成雪崩。
推荐定义:
/startup:初始化完成了吗?
/live:进程和核心运行循环正常吗?
/ready:依赖和业务条件满足,可以接流量吗?典型判断:
| 接口 | 可以检查 |
|---|---|
| startup | 配置加载、端口启动、初始化任务 |
| live | 主线程、事件循环、关键内部状态 |
| ready | 数据库连接、缓存连接、配置加载、流量开关 |
但不要把所有外部依赖都放进 /live。
容器被探针杀死后,是否重新启动还与:
restartPolicy有关。
Deployment、StatefulSet 中通常使用:
restartPolicy: Always普通 Pod 常见:
restartPolicy: Never如果是一次性 Job,不应该简单使用 Liveness Probe 代替任务状态判断。
Job 的目标是:
任务执行完成并退出如果给长时间运行的任务配置不合适的 Liveness,可能导致任务被反复重启。
Job 更关注:
backoffLimitactiveDeadlineSecondscompletionsparallelism探针解决:
应用是否存活、是否接流量PDB 解决:
节点维护或主动驱逐时,最多允许多少副本同时不可用二者可以配合:
Readiness Probe + PDB + 多副本实现更安全的滚动升级和节点维护。
不推荐只写:
initialDelaySeconds: 300更推荐:
这样启动完成后,Liveness 和 Readiness 可以立即按正常频率检查,而不是始终等待很长时间。
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: periodSeconds: 5
timeoutSeconds: 3
failureThreshold: startupProbe:
periodSeconds: 10
failureThreshold: 60可以适当增大:
timeoutSeconds: 5
failureThreshold: 5参数没有绝对标准,应结合启动耗时、请求延迟和故障恢复策略测试。
| 现象 | 适合的处理 |
|---|---|
| 应用启动很慢 | Startup Probe |
| 应用进程死锁 | Liveness Probe |
| 数据库暂时不可用 | Readiness Probe |
| 应用预热未完成 | Readiness Probe |
| 端口是否监听 | TCP Probe |
| HTTP 接口是否正常 | HTTP Probe |
| gRPC 健康状态 | gRPC Probe |
| 没有健康接口 | Exec Probe |
| 一次性任务执行完成 | 依靠进程退出码和 Job 状态 |
总结就是:
Startup:控制启动阶段
Liveness:负责自动重启
Readiness:负责流量摘除
HTTP/TCP/Exec/gRPC:负责定义“如何检查”探针配置没有脱离应用独立存在,最终要结合:
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 10startupProbe:
httpGet:
path: /actuator/health
port: 8080
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 30livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3
successThreshold: 1readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
successThreshold: 1livenessProbe:
httpGet:
scheme: HTTP
path: /healthz
port: 8080
initialDelaySeconds: 20
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3readinessProbe:
httpGet:
path: /ready
port: 8080
httpHeaders:
- name: X-Health-Check
value: KuberneteslivenessProbe:
httpGet:
scheme: HTTPS
path: /healthz
port: 8443livenessProbe:
tcpSocket:
port: 8080
initialDelaySeconds: 15
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3livenessProbe:
exec:
command:
- sh
- -c
- test -f /tmp/healthy
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3readinessProbe:
exec:
command:
- sh
- -c
- mysqladmin ping -h 127.0.0.1 -uroot -p"$MYSQL_ROOT_PASSWORD"livenessProbe:
grpc:
port: 50051
initialDelaySeconds: 10
periodSeconds: 10apiVersion: apps/v1
kind: Deployment
metadata:
name: web-app
spec:
replicas: 3
selector:
matchLabels:
app: web-app
template:
metadata:
labels:
app: web-app
spec:
containers:
- name: web
image: example/web-app:1.0
ports:
- name: http
containerPort: 8080
startupProbe:
httpGet:
path: /health/startup
port: http
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 30
livenessProbe:
httpGet:
path: /health/live
port: http
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3
readinessProbe:
httpGet:
path: /health/ready
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
successThreshold: 1startupProbe:
httpGet:
path: /actuator/health
port: 8080
failureThreshold: 30
periodSeconds: 10
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
periodSeconds: 5
failureThreshold: 3livenessProbe:
httpGet:
path: /healthz
port: 8080
terminationGracePeriodSeconds: 30创建新 Pod
↓
新 Pod 启动
↓
Startup Probe 成功
↓
Readiness Probe 成功
↓
新 Pod 加入 Service
↓
旧 Pod 逐步退出Pod 开始终止
↓
从 Service 中移除
↓
执行 preStop
↓
发送 SIGTERM
↓
等待 terminationGracePeriodSeconds
↓
发送 SIGKILLlifecycle:
preStop:
exec:
command:
- sh
- -c
- sleep 10数据库故障
↓
所有 Pod Liveness 失败
↓
所有 Pod 重启
↓
应用重连数据库
↓
数据库压力更大
↓
故障扩大startupProbe:
httpGet:
path: /health/startup
port: 8080
periodSeconds: 10
failureThreshold: 30startupProbe:
httpGet:
path: /health/startup
port: 8080
periodSeconds: 10
failureThreshold: 30
livenessProbe:
httpGet:
path: /health/live
port: 8080
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /health/ready
port: 8080
periodSeconds: 5
failureThreshold: 3livenessProbe:
httpGet:
path: /healthz
port: 8080
readinessProbe:
httpGet:
path: /ready
port: 8080livenessProbe:
tcpSocket:
port: 9000
readinessProbe:
tcpSocket:
port: 9000livenessProbe:
grpc:
port: 50051
readinessProbe:
grpc:
port: 50051readinessProbe:
exec:
command:
- /bin/sh
- -c
- /app/check-ready.shapiVersion: apps/v1
kind: Deployment
metadata:
name: order-service
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels:
app: order-service
template:
metadata:
labels:
app: order-service
spec:
terminationGracePeriodSeconds: 30
containers:
- name: order-service
image: example/order-service:1.2.0
ports:
- name: http
containerPort: 8080
startupProbe:
httpGet:
path: /health/startup
port: http
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 30
livenessProbe:
httpGet:
path: /health/live
port: http
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3
readinessProbe:
httpGet:
path: /health/ready
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
successThreshold: 2
lifecycle:
preStop:
exec:
command:
- sh
- -c
- sleep 5
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi慢启动应用:
Startup + Liveness + Readiness
普通 HTTP 服务:
Liveness + Readiness
纯 TCP 服务:
TCP Liveness + TCP Readiness
gRPC 服务:
gRPC Liveness + gRPC Readiness
无健康接口的特殊服务:
Exec Probekubectl describe pod <pod-name>
kubectl logs <pod-name>
kubectl logs <pod-name> --previousstartupProbe:
httpGet:
path: /health/startup
port: 8080
periodSeconds: 10
failureThreshold: 30