What is HorizontalPodAutoscaler
HPA acts like a restaurant floor manager on a busy Friday evening:
🟢 Junior Level
30-Second Summary
HorizontalPodAutoscaler (HPA) is a built-in Kubernetes controller that automatically adjusts the number of Pod replicas in a Deployment or StatefulSet based on observed application resource metrics or custom external telemetry.
- How it works: HPA polls the
Metrics Serverevery 15 seconds, calculates average resource utilization across existing pods, and compares it to a declared target. - Traffic surge: CPU consumption exceeds the threshold $\to$ HPA increases replica count (Scale Out / Scale Up), distributing incoming load across more instances.
- Traffic drop: Traffic subsides during off-peak hours $\to$ HPA gradually terminates excess pods down to
minReplicas(Scale In / Scale Down), conserving cloud infrastructure costs.
Real-World Analogy
HPA acts like a restaurant floor manager on a busy Friday evening:
- When the dining room fills up and waiters are overloaded (CPU utilization > 70%), the manager calls in backup staff from the standby roster.
- After midnight when patrons leave (CPU utilization < 20%), off-duty staff are released home so the establishment does not pay idle wages.
Basic HPA Manifest (autoscaling/v2)
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: payment-service-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: payment-service
minReplicas: 2 # High Availability baseline
maxReplicas: 10 # Hard ceiling to prevent runaway cloud costs
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60 # Maintain average CPU utilization around 60%
[!IMPORTANT] Two Mandatory Prerequisites for HPA Operation:
- The
Metrics Servermust be deployed and healthy within the cluster to scrape container cgroup metrics.- The target Pod specification MUST define
resources.requests.cpu, because percentage utilization is computed strictly relative to requested resources!
🟡 Middle Level
Replica Calculation Formula & Tolerance Threshold
The HPA controller determines the desired replica count using the following mathematical formula:
\[\mathbf{DesiredReplicas} = \left\lceil \mathbf{CurrentReplicas} \times \left( \frac{\mathbf{CurrentMetricValue}}{\mathbf{TargetMetricValue}} \right) \right\rceil\]The 10% Built-in Tolerance Threshold
To prevent rapid thrashing caused by trivial 1–2% metric fluctuations:
- Kubernetes enforces the
--horizontal-pod-autoscaler-toleranceparameter (default:0.1, or 10%). -
If $\left \frac{\text{Current}}{\text{Target}} - 1.0 \right \le 0.1$, HPA takes no scaling action! - Example: If target CPU utilization is 60%, and current utilization fluctuates between 55% and 65%, HPA remains idle.
The 4 Metric Types Supported in autoscaling/v2
| Metric Type | Data Source | Production Use Case |
|---|---|---|
Resource |
Metrics Server (cgroups) | Standard container cpu or memory. |
Pods |
Prometheus Adapter / Custom Metrics | Pod-aggregated application metrics (e.g., HTTP RPS per pod). |
Object |
Custom Metrics API / K8s Objects | Metrics tied to cluster objects (e.g., Ingress connection queue depth). |
External |
External Metrics API / Cloud Providers | Metrics outside the cluster (e.g., AWS SQS queue depth, Kafka topic lag). |
Fine-Tuning Scaling Velocity: The behavior Block
To prevent flapping (rapid oscillation between scaling up and down), configure stabilization windows and rate limits:
spec:
behavior:
scaleUp:
stabilizationWindowSeconds: 0 # Scale out immediately upon traffic spikes
policies:
- type: Percent
value: 100 # At most double the replica count in a single evaluation
periodSeconds: 15
scaleDown:
stabilizationWindowSeconds: 300 # Wait 5 minutes of sustained low traffic before terminating pods
policies:
- type: Pods
value: 1 # Terminate at most 1 pod every 60 seconds
periodSeconds: 60
🔴 Senior Level
Multi-Metric Resolution Algorithm (The MAX Rule)
When an HPA manifest specifies multiple metrics (e.g., CPU + RPS + Latency):
metrics:
- type: Resource
resource: { name: cpu, target: { type: Utilization, averageUtilization: 60 } } # Calculated: 4 pods
- type: Pods
pods: { metric: { name: http_requests_per_second }, target: { averageValue: "100" } } # Calculated: 8 pods
Resolution Rule: The HPA controller evaluates each metric rule independently and selects the MAXIMUM (highest) calculated replica count across all metrics. In this scenario, HPA scales the Deployment to 8 pods. This fail-safe design guarantees capacity for the most heavily saturated resource.
The Java JIT Warmup Trap & False Overshooting
In enterprise Java Spring Boot deployments, HPA can trigger a catastrophic scaling avalanche:
- Traffic surges $\to$ HPA adds 2 new pods.
- New JVM containers boot. During the classloading and tiered C1/C2 JIT compilation phase, the JVM consumes 100% CPU for 30–60 seconds.
- HPA scrapes the Metrics Server, observes the newly added pods at 100% CPU, and assumes the cluster remains severely under-provisioned.
- HPA schedules 4 more pods $\to$ they also hit 100% JIT warmup $\to$ HPA pins the Deployment at
maxReplicas.
Architectural Mitigations:
- Proper Probe Separation: Configure
startupProbewith ample failure thresholds (failureThreshold: 30,periodSeconds: 2) so pods are not marked ready before warm initialization completes. - Scaling Policies: Set
behavior.scaleUp.stabilizationWindowSeconds: 60or limit step growth withpolicies.type: Pods. - Ahead-Of-Time Compilation: Adopt Spring Boot 3 AOT / GraalVM Native Image or CRaC (Coordinated Restore at Checkpoint) to eliminate JIT compilation overhead at container startup.
Event-Driven Autoscaling with KEDA
For asynchronous consumers processing message queues, standard CPU metrics fail: idle consumers polling an empty queue consume 0% CPU, and when 500,000 messages arrive, CPU stays low while consumers process items sequentially.
KEDA (Kubernetes Event-Driven Autoscaling) solves this:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: payment-consumer-scaler
spec:
scaleTargetRef:
name: payment-consumer
minReplicaCount: 0 # Full Scale-to-Zero support!
maxReplicaCount: 20
triggers:
- type: kafka
metadata:
bootstrapServers: kafka:9092
consumerGroup: payment-processors
topic: payments
lagThreshold: "100" # Add 1 pod for every 100 messages of consumer lag
Behind the scenes, KEDA registers an External Metrics API server and dynamically provisions a native HPA object, abstracting Prometheus Adapter configuration.
4 Tricky Questions
1. How does the built-in 10% tolerance threshold work, and why does HPA take no action if CPU increases from 50% to 54% when the target is 50%?
Answer:
To prevent constant pod churn caused by micro-variations in traffic, the HPA controller applies a tolerance check before altering replica counts:
\(\left| \frac{\text{CurrentMetricValue}}{\text{TargetMetricValue}} - 1.0 \right| \le \text{tolerance}\)
The default tolerance flag (--horizontal-pod-autoscaler-tolerance) is 0.1 (10%).
In this scenario:
\(\frac{54}{50} = 1.08 \implies |1.08 - 1.0| = 0.08 \ (8\%)\)
Because 8% is strictly less than or equal to the 10% tolerance threshold, the controller considers the workload within acceptable bounds and leaves the current replica count unchanged.
2. Why does kubectl get hpa display <unknown>/60% in the TARGETS column, and what are the 3 most common root causes?
Answer:
The <unknown> status indicates that the HPA controller cannot retrieve valid pod metric telemetry from the Kubernetes Resource Metrics API (v1beta1.metrics.k8s.io).
The three primary causes are:
- Missing
resources.requests: The container specification in the Deployment template lacksresources.requests.cpu. Without requested limits, percentage utilization cannot mathematically be computed. - Metrics Server Unhealthy or Missing: The
metrics-serverpod is not deployed, in aCrashLoopBackOff, or blocked by kubelet TLS certificate validation (often requiring--kubelet-insecure-tlsin development clusters). - Cold Pod Startup Period: Pods were created within the last 30–60 seconds, and the Metrics Server has not yet completed its initial metric scrape and averaging interval.
3. If an HPA defines 3 metrics: CPU (requires 4 pods), Memory (requires 6 pods), and Custom RPS (requires 8 pods), what final replica count will the controller set?
Answer:
The controller will set 8 pods.
Under the Kubernetes autoscaling/v2 specification, when multiple metrics are listed in the metrics array, HPA computes the desired replica count for each metric rule independently and executes a max() operation:
\(\text{Replicas} = \max(\text{Replicas}_{\text{CPU}}, \text{Replicas}_{\text{Memory}}, \text{Replicas}_{\text{RPS}}) = \max(4, 6, 8) = 8\)
This conservative behavior ensures the application remains fully provisioned to handle the most saturated metric, preventing service degradation.
4. What happens when an engineer runs kubectl scale deployment payment-service --replicas=5 while an HPA is actively managing that Deployment?
Answer: An imperative-versus-declarative conflict occurs:
- The
kubectl scalecommand imperatively updatesspec.replicas: 5in the Deployment manifest. - The Deployment controller creates pods to match the new count of 5.
- Within 15 seconds, the HPA controller’s reconciliation loop evaluates current metrics against its target.
- If current metrics justify only 2 pods, HPA immediately overwrites
spec.replicasback to 2, terminating the 3 newly spawned pods. Rule: When HPA is attached to a workload, imperative scaling viakubectl scaleis futile; adjustments must be made by modifyingminReplicas,maxReplicas, or metric targets directly on the HPA object.
🎯 Interview Cheat Sheet
Core HPA Architecture
- Purpose: Native controller for horizontally scaling Pod replicas.
- Evaluation Loop: Every 15 seconds via
kube-controller-manager. - Hard Prerequisites: Operational
Metrics Server+ explicitresources.requests. - Scaling Formula: $\text{Desired} = \lceil \text{Current} \times (\text{CurrentMetric} / \text{TargetMetric}) \rceil$.
- Multiple Metrics: Evaluated independently; controller always applies
MAX(). - Java Golden Rule: Never autoscale Java on memory; scale strictly on CPU, RPS, or queue lag.
- Stabilization Windows: Quick scaleUp (0–15s), conservative scaleDown (300s) to prevent flapping.
HPA Metric Types Comparison
| Type | Telemetry Origin | Common Usage |
| :— | :— | :— |
| Resource | Container cgroups via Metrics Server | CPU utilization, Memory bytes |
| Pods | Prometheus Adapter / Pod endpoints | HTTP RPS per pod, Active connections |
| Object | K8s Cluster Objects | Ingress backlog, Job count |
| External | External APIs / KEDA | Kafka topic consumer lag, SQS length |
Red Flags (What NOT to Say)
- ❌ “HPA calculates CPU percentage from
resources.limits.” — Percentage is strictly computed againstresources.requests. - ❌ “I can temporarily scale an HPA-controlled deployment using
kubectl scale.” — HPA reconciliation overwrites manual changes within 15 seconds. - ❌ “Standard HPA can scale a Deployment down to 0 replicas on idle.” — Native HPA enforces
minReplicas >= 1; Scale-to-Zero requires KEDA or Knative.
Related Topics
- How Does Scaling Work in Kubernetes — Horizontal vs Vertical autoscaling
- What is Pod in Kubernetes — Pod specs and resource requests
- What is ReplicaSet — Pod replica management
- Why Are Health Checks Needed — Startup, liveness, and readiness probes