🐳 Section 14 · Question #19

Why are health checks needed

Kubernetes provides three complementary health probe types:


🟢 Junior Level

30-Second Summary

Health Checks (Probes) are a built-in diagnostic subsystem in Kubernetes through which the node agent (kubelet) continuously monitors the physical and logical state of containerized workloads to deliver autonomous self-healing and zero-downtime rolling updates.

Kubernetes provides three complementary health probe types:

  1. startupProbe: Answers: “Has the application finished initializing?” It grants slow-starting applications (such as enterprise Java Spring Boot workloads) ample time to boot by temporarily disabling all other probes.
  2. livenessProbe: Answers: “Is the container process healthy or deadlocked/frozen?” If this probe fails repeatedly, Kubernetes forcibly kills and restarts the container.
  3. readinessProbe: Answers: “Is the Pod ready to accept incoming client network traffic?” If this probe fails, the container is NOT restarted, but its IP is dynamically withdrawn from Service load-balancing pools (Endpoints / EndpointSlice).

Real-World Analogy

Consider the grand opening of a specialty coffee shop:

  • startupProbe: A sign on the door reads “Under Renovation & Machine Calibration”. City inspectors wait patiently without issuing fines for not serving coffee yet.
  • livenessProbe: A vital signs check: “Is the barista conscious and breathing, or collapsed from exhaustion?” If the barista collapses (Deadlock), management calls in a substitute immediately (container restart).
  • readinessProbe: A counter sign reads “5-Minute Break to Grind Beans”. The barista is alive and well, but currently unavailable to take orders (incoming customers are diverted to the adjacent checkout register).

Comprehensive Manifest with the Complete Probe Triad (Java 21 / Spring Boot 3)

apiVersion: v1
kind: Pod
metadata:
  name: order-service
spec:
  containers:
  - name: order-app
    image: order-service:1.0
    # 1. Startup Protection: Grants up to 30 * 5s = 150 seconds for JVM startup
    startupProbe:
      httpGet:
        path: /actuator/health/liveness
        port: 8080
      periodSeconds: 5
      failureThreshold: 30

    # 2. Liveness Check: Detects deadlocks and fatal thread freezes
    livenessProbe:
      httpGet:
        path: /actuator/health/liveness
        port: 8080
      periodSeconds: 10
      failureThreshold: 3

    # 3. Readiness Check: Gates user traffic based on operational readiness
    readinessProbe:
      httpGet:
        path: /actuator/health/readiness
        port: 8080
      periodSeconds: 5
      failureThreshold: 2

🟡 Middle Level

Probe Lifecycle and State Machine

The Kubelet prober engine coordinates probe execution using a deterministic state machine:

               [ Container Spawned in Linux Kernel ]
                                │
                                ▼
                   [ 1. startupProbe Active ]
               (liveness & readiness frozen/idle)
                                │
                ┌───────────────┴───────────────┐
                │ Success                       │ Failure (failureThreshold exceeded)
                ▼                               ▼
    [ startupProbe Disarmed Permanently ]   [ 🔴 Container Killed & Restarted ]
                │
                ├───────────────────────────────┐
                ▼                               ▼
     [ 2. livenessProbe Active ]     [ 3. readinessProbe Active ]
        (Continuous polling)            (Continuous polling)
                │                               │
          Failure:                        Failure:
          🔴 Restart container            🟡 Remove IP from Endpoints

Comparison of the Probe Triad

Evaluation Metric startupProbe livenessProbe readinessProbe
Core Question Has the boot process finished? Is the process deadlocked? Can it handle client RPCs?
Active Window Container startup only After startup success, continuous After startup success, continuous
Action on Failure 🔴 Container terminated/restarted 🔴 Container restarted 🟡 Removed from Service traffic
Cross-Probe Impact Blocks Liveness and Readiness Independent evaluation Independent evaluation
Primary Scope Spring context loading, migrations Deadlocks, infinite event loops Cache warm-up, connection pools

The Catastrophic Anti-Pattern: A Single Unified /health Endpoint

Many developers mistakenly point all three probes to a generic /health endpoint:

# ❌ CATASTROPHIC ARCHITECTURAL MISTAKE
livenessProbe:
  httpGet: { path: /health, port: 8080 } # Evaluates PostgreSQL and Redis!
readinessProbe:
  httpGet: { path: /health, port: 8080 }

Why this destroys production reliability:

  1. If an external Redis cache experiences a transient 10-second failover, /health returns HTTP 503.
  2. Because the single endpoint is bound to livenessProbe, Kubernetes assumes the Java application process itself is dead.
  3. Kubelet simultaneously terminates and restarts all backend pods across the cluster.
  4. The service plunges into a cluster-wide CrashLoop cascade, turning a minor cache blip into a catastrophic full-scale outage.

🔴 Senior Level

Kubelet Prober Architecture Internals

Inside the kubelet daemon, the pkg/kubelet/prober subsystem manages health evaluations:

  • Asynchronous Worker Goroutines: For every defined probe across all containers on a node, Kubelet spawns an isolated Go worker goroutine (worker.go).
  • Dedicated WorkQueues: Probe executions run independently of Kubelet’s core Pod lifecycle reconciliation loop, preventing slow probes from blocking node operations.
  • Strict Network Socket Timeouts: If a container socket fails to reply within timeoutSeconds, the goroutine cancels the net.Dialer context and registers a failure, preventing socket handle leaks on the host node.

JVM (HotSpot) Production Tuning

Running enterprise Java workloads inside containers requires specific probe accommodations:

1. Tolerating Garbage Collector Stop-The-World (STW) Pauses

When using G1 GC or Parallel GC with heaps between 8 GB and 32 GB:

  • A Full GC pause can temporarily suspend application threads for 1–3 seconds.
  • An aggressive probe timeout (timeoutSeconds: 1) will falsely record a failure.
  • Rule: Set timeoutSeconds to at least 3–5 seconds, ensuring it exceeds the maximum anticipated STW pause duration.

2. JIT Warmup & Ahead-Of-Time (AOT) Compilation

In Java 21 with Spring Boot 3, early incoming requests trigger heavy C1/C2 JIT compilation:

  • Configure startupProbe with ample headroom (failureThreshold: 30, periodSeconds: 5 = 150 seconds).
  • For sub-second startup times, compile applications with GraalVM Native Image or leverage CRaC (Coordinated Restore at Checkpoint) to eliminate JIT compilation during container boot.

Orchestration During Graceful Shutdown

During workload termination, Kubernetes coordinates shutdown to avoid packet loss:

  1. kubectl delete pod $\to$ Pod status switches to Terminating.
  2. Kubelet executes the configured preStop hook (sleep 10).
  3. The EndpointSlice controller removes the Pod IP from cluster iptables/IPVS routing tables.
  4. Kubelet sends SIGTERM to the container process.
  5. Spring Boot waits for active HTTP requests to complete (spring.lifecycle.timeout-per-shutdown-phase=30s).
  6. The container exits cleanly with return code 0. During the Terminating phase, Kubelet suppresses livenessProbe restart actions, even if probe requests begin failing.

4 Tricky Questions

1. What happens if a container exhausts its failureThreshold in startupProbe, and how does modern Kubernetes behavior compare to legacy versions?

Answer:

  • Modern Kubernetes (1.20+): If startupProbe fails failureThreshold consecutive times, Kubelet determines that the container failed to initialize properly. It issues a SIGKILL to terminate the container and triggers a restart in accordance with the Pod’s restartPolicy (incrementing the restart counter).
  • Legacy Kubernetes (pre-1.20): Due to an internal prober subsystem bug, an exhausted startupProbe occasionally left the Pod permanently stuck in an unready state without triggering a container restart (resolved in PR #95190).

2. Why is using the exec mechanism for frequent health checks strongly discouraged in high-throughput production clusters?

Answer: Every time an exec probe executes, Kubelet instructs the container runtime (containerd / CRI-O) to make Linux kernel clone() / fork() and execve() system calls:

  • A new operating system process (e.g., /bin/sh or curl) is spawned inside the container cgroup namespace.
  • This creates significant CPU kernel overhead, forces context switches, generates cgroup lock contention, and churns the host’s PID table.
  • In resource-constrained environments, fork() calls can time out, producing false probe failures.
  • Furthermore, modern Distroless container images exclude shells and binaries like curl. Native httpGet or tcpSocket probes avoid process creation entirely by validating open network sockets from the host.

3. What is the operational state of livenessProbe and readinessProbe while startupProbe is actively running and has not yet succeeded?

Answer: While startupProbe is executing, both livenessProbe and readinessProbe are completely disabled (frozen). Kubelet does not dispatch HTTP or TCP probe requests to their configured endpoints. This design prevents slow-booting applications (running Liquibase migrations, hydrating caches, or warming JIT caches) from being prematurely terminated by an aggressive liveness probe.


4. Can a failing readinessProbe protect an application from being terminated by the Linux OOM Killer (Exit Code 137) during a traffic surge?

Answer: Yes, it can. When an application implements adaptive load shedding:

  • If Tomcat’s thread pool saturates or internal memory pressure mounts, the application programmatically transitions ReadinessState to REFUSING_TRAFFIC (returning HTTP 503 on the readiness endpoint).
  • Kubelet dynamically removes the Pod IP from the Service EndpointSlice, cutting off new incoming client requests.
  • The Pod continues executing in-flight requests, allows the JVM Garbage Collector to reclaim heap memory, and stabilizes without exceeding memory limits, successfully evading the host kernel’s OOM Killer (SIGKILL 137). Once memory stabilizes, readiness returns to HTTP 200.

🎯 Interview Cheat Sheet

Core Health Check Principles

  • The Probe Triad:
    • startupProbe: Protects slow initialization (JVM classloading, DB migrations); disarms after initial success.
    • livenessProbe: Recovers stuck processes (deadlocks, infinite loops); triggers container restart.
    • readinessProbe: Regulates client traffic routing; isolates the Pod from Service Endpoints without restarting.
  • Anti-Pattern Warning: Never use a single /health endpoint checking external dependencies across all probes.
  • JVM Timeout Standard: Keep timeoutSeconds between 3–5 seconds to prevent false alerts during GC STW pauses.
  • Probe Mechanisms: Prefer lightweight httpGet and grpc over resource-intensive exec forks.
  • Kubelet Engine: Managed asynchronously via pkg/kubelet/prober goroutines.

Action on Failure Summary

| Probe | Kubernetes Action | Traffic Routing Impact | Primary Objective | | :— | :— | :— | :— | | startup | Kills and restarts container | Traffic suppressed | Safely boots heavy Spring Boot apps | | liveness | Restarts container | Active connections dropped | Recovers deadlocks and frozen processes | | readiness | Detaches from Service | Traffic diverted to healthy peers | Isolates overloaded pods during warm-up |

Red Flags (What NOT to Say)

  • ❌ “We use the same health endpoint for Liveness and Readiness probes.” — An external outage will cascade into cluster-wide container restart storms.
  • ❌ “A failing Readiness Probe causes Kubernetes to kill the container.” — Readiness probes never terminate containers; they only detach network endpoints.
  • ❌ “Java containers do not need startup probes if you set initialDelaySeconds: 5.” — Slow Spring Boot startups will crash repeatedly in an infinite boot-loop.