🐳 Section 14 · Question #15

What is HorizontalPodAutoscaler

HPA acts like a restaurant floor manager on a busy Friday evening:


🟢 Junior Level

30-Second Summary

HorizontalPodAutoscaler (HPA) is a built-in Kubernetes controller that automatically adjusts the number of Pod replicas in a Deployment or StatefulSet based on observed application resource metrics or custom external telemetry.

  • How it works: HPA polls the Metrics Server every 15 seconds, calculates average resource utilization across existing pods, and compares it to a declared target.
  • Traffic surge: CPU consumption exceeds the threshold $\to$ HPA increases replica count (Scale Out / Scale Up), distributing incoming load across more instances.
  • Traffic drop: Traffic subsides during off-peak hours $\to$ HPA gradually terminates excess pods down to minReplicas (Scale In / Scale Down), conserving cloud infrastructure costs.

Real-World Analogy

HPA acts like a restaurant floor manager on a busy Friday evening:

  • When the dining room fills up and waiters are overloaded (CPU utilization > 70%), the manager calls in backup staff from the standby roster.
  • After midnight when patrons leave (CPU utilization < 20%), off-duty staff are released home so the establishment does not pay idle wages.

Basic HPA Manifest (autoscaling/v2)

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: payment-service-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: payment-service
  minReplicas: 2  # High Availability baseline
  maxReplicas: 10 # Hard ceiling to prevent runaway cloud costs
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 60 # Maintain average CPU utilization around 60%

[!IMPORTANT] Two Mandatory Prerequisites for HPA Operation:

  1. The Metrics Server must be deployed and healthy within the cluster to scrape container cgroup metrics.
  2. The target Pod specification MUST define resources.requests.cpu, because percentage utilization is computed strictly relative to requested resources!

🟡 Middle Level

Replica Calculation Formula & Tolerance Threshold

The HPA controller determines the desired replica count using the following mathematical formula:

\[\mathbf{DesiredReplicas} = \left\lceil \mathbf{CurrentReplicas} \times \left( \frac{\mathbf{CurrentMetricValue}}{\mathbf{TargetMetricValue}} \right) \right\rceil\]

The 10% Built-in Tolerance Threshold

To prevent rapid thrashing caused by trivial 1–2% metric fluctuations:

  • Kubernetes enforces the --horizontal-pod-autoscaler-tolerance parameter (default: 0.1, or 10%).
  • If $\left \frac{\text{Current}}{\text{Target}} - 1.0 \right \le 0.1$, HPA takes no scaling action!
  • Example: If target CPU utilization is 60%, and current utilization fluctuates between 55% and 65%, HPA remains idle.

The 4 Metric Types Supported in autoscaling/v2

Metric Type Data Source Production Use Case
Resource Metrics Server (cgroups) Standard container cpu or memory.
Pods Prometheus Adapter / Custom Metrics Pod-aggregated application metrics (e.g., HTTP RPS per pod).
Object Custom Metrics API / K8s Objects Metrics tied to cluster objects (e.g., Ingress connection queue depth).
External External Metrics API / Cloud Providers Metrics outside the cluster (e.g., AWS SQS queue depth, Kafka topic lag).

Fine-Tuning Scaling Velocity: The behavior Block

To prevent flapping (rapid oscillation between scaling up and down), configure stabilization windows and rate limits:

spec:
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0 # Scale out immediately upon traffic spikes
      policies:
      - type: Percent
        value: 100 # At most double the replica count in a single evaluation
        periodSeconds: 15
    scaleDown:
      stabilizationWindowSeconds: 300 # Wait 5 minutes of sustained low traffic before terminating pods
      policies:
      - type: Pods
        value: 1 # Terminate at most 1 pod every 60 seconds
        periodSeconds: 60

🔴 Senior Level

Multi-Metric Resolution Algorithm (The MAX Rule)

When an HPA manifest specifies multiple metrics (e.g., CPU + RPS + Latency):

metrics:
- type: Resource
  resource: { name: cpu, target: { type: Utilization, averageUtilization: 60 } } # Calculated: 4 pods
- type: Pods
  pods: { metric: { name: http_requests_per_second }, target: { averageValue: "100" } } # Calculated: 8 pods

Resolution Rule: The HPA controller evaluates each metric rule independently and selects the MAXIMUM (highest) calculated replica count across all metrics. In this scenario, HPA scales the Deployment to 8 pods. This fail-safe design guarantees capacity for the most heavily saturated resource.


The Java JIT Warmup Trap & False Overshooting

In enterprise Java Spring Boot deployments, HPA can trigger a catastrophic scaling avalanche:

  1. Traffic surges $\to$ HPA adds 2 new pods.
  2. New JVM containers boot. During the classloading and tiered C1/C2 JIT compilation phase, the JVM consumes 100% CPU for 30–60 seconds.
  3. HPA scrapes the Metrics Server, observes the newly added pods at 100% CPU, and assumes the cluster remains severely under-provisioned.
  4. HPA schedules 4 more pods $\to$ they also hit 100% JIT warmup $\to$ HPA pins the Deployment at maxReplicas.

Architectural Mitigations:

  • Proper Probe Separation: Configure startupProbe with ample failure thresholds (failureThreshold: 30, periodSeconds: 2) so pods are not marked ready before warm initialization completes.
  • Scaling Policies: Set behavior.scaleUp.stabilizationWindowSeconds: 60 or limit step growth with policies.type: Pods.
  • Ahead-Of-Time Compilation: Adopt Spring Boot 3 AOT / GraalVM Native Image or CRaC (Coordinated Restore at Checkpoint) to eliminate JIT compilation overhead at container startup.

Event-Driven Autoscaling with KEDA

For asynchronous consumers processing message queues, standard CPU metrics fail: idle consumers polling an empty queue consume 0% CPU, and when 500,000 messages arrive, CPU stays low while consumers process items sequentially.

KEDA (Kubernetes Event-Driven Autoscaling) solves this:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: payment-consumer-scaler
spec:
  scaleTargetRef:
    name: payment-consumer
  minReplicaCount: 0 # Full Scale-to-Zero support!
  maxReplicaCount: 20
  triggers:
  - type: kafka
    metadata:
      bootstrapServers: kafka:9092
      consumerGroup: payment-processors
      topic: payments
      lagThreshold: "100" # Add 1 pod for every 100 messages of consumer lag

Behind the scenes, KEDA registers an External Metrics API server and dynamically provisions a native HPA object, abstracting Prometheus Adapter configuration.


4 Tricky Questions

1. How does the built-in 10% tolerance threshold work, and why does HPA take no action if CPU increases from 50% to 54% when the target is 50%?

Answer: To prevent constant pod churn caused by micro-variations in traffic, the HPA controller applies a tolerance check before altering replica counts: \(\left| \frac{\text{CurrentMetricValue}}{\text{TargetMetricValue}} - 1.0 \right| \le \text{tolerance}\) The default tolerance flag (--horizontal-pod-autoscaler-tolerance) is 0.1 (10%). In this scenario: \(\frac{54}{50} = 1.08 \implies |1.08 - 1.0| = 0.08 \ (8\%)\) Because 8% is strictly less than or equal to the 10% tolerance threshold, the controller considers the workload within acceptable bounds and leaves the current replica count unchanged.


2. Why does kubectl get hpa display <unknown>/60% in the TARGETS column, and what are the 3 most common root causes?

Answer: The <unknown> status indicates that the HPA controller cannot retrieve valid pod metric telemetry from the Kubernetes Resource Metrics API (v1beta1.metrics.k8s.io). The three primary causes are:

  1. Missing resources.requests: The container specification in the Deployment template lacks resources.requests.cpu. Without requested limits, percentage utilization cannot mathematically be computed.
  2. Metrics Server Unhealthy or Missing: The metrics-server pod is not deployed, in a CrashLoopBackOff, or blocked by kubelet TLS certificate validation (often requiring --kubelet-insecure-tls in development clusters).
  3. Cold Pod Startup Period: Pods were created within the last 30–60 seconds, and the Metrics Server has not yet completed its initial metric scrape and averaging interval.

3. If an HPA defines 3 metrics: CPU (requires 4 pods), Memory (requires 6 pods), and Custom RPS (requires 8 pods), what final replica count will the controller set?

Answer: The controller will set 8 pods. Under the Kubernetes autoscaling/v2 specification, when multiple metrics are listed in the metrics array, HPA computes the desired replica count for each metric rule independently and executes a max() operation: \(\text{Replicas} = \max(\text{Replicas}_{\text{CPU}}, \text{Replicas}_{\text{Memory}}, \text{Replicas}_{\text{RPS}}) = \max(4, 6, 8) = 8\) This conservative behavior ensures the application remains fully provisioned to handle the most saturated metric, preventing service degradation.


4. What happens when an engineer runs kubectl scale deployment payment-service --replicas=5 while an HPA is actively managing that Deployment?

Answer: An imperative-versus-declarative conflict occurs:

  1. The kubectl scale command imperatively updates spec.replicas: 5 in the Deployment manifest.
  2. The Deployment controller creates pods to match the new count of 5.
  3. Within 15 seconds, the HPA controller’s reconciliation loop evaluates current metrics against its target.
  4. If current metrics justify only 2 pods, HPA immediately overwrites spec.replicas back to 2, terminating the 3 newly spawned pods. Rule: When HPA is attached to a workload, imperative scaling via kubectl scale is futile; adjustments must be made by modifying minReplicas, maxReplicas, or metric targets directly on the HPA object.

🎯 Interview Cheat Sheet

Core HPA Architecture

  • Purpose: Native controller for horizontally scaling Pod replicas.
  • Evaluation Loop: Every 15 seconds via kube-controller-manager.
  • Hard Prerequisites: Operational Metrics Server + explicit resources.requests.
  • Scaling Formula: $\text{Desired} = \lceil \text{Current} \times (\text{CurrentMetric} / \text{TargetMetric}) \rceil$.
  • Multiple Metrics: Evaluated independently; controller always applies MAX().
  • Java Golden Rule: Never autoscale Java on memory; scale strictly on CPU, RPS, or queue lag.
  • Stabilization Windows: Quick scaleUp (0–15s), conservative scaleDown (300s) to prevent flapping.

HPA Metric Types Comparison

| Type | Telemetry Origin | Common Usage | | :— | :— | :— | | Resource | Container cgroups via Metrics Server | CPU utilization, Memory bytes | | Pods | Prometheus Adapter / Pod endpoints | HTTP RPS per pod, Active connections | | Object | K8s Cluster Objects | Ingress backlog, Job count | | External | External APIs / KEDA | Kafka topic consumer lag, SQS length |

Red Flags (What NOT to Say)

  • ❌ “HPA calculates CPU percentage from resources.limits.” — Percentage is strictly computed against resources.requests.
  • ❌ “I can temporarily scale an HPA-controlled deployment using kubectl scale.” — HPA reconciliation overwrites manual changes within 15 seconds.
  • ❌ “Standard HPA can scale a Deployment down to 0 replicas on idle.” — Native HPA enforces minReplicas >= 1; Scale-to-Zero requires KEDA or Knative.