SkillByAIOpen interactive version →

Lesson 18 / 25

Horizontal Pod Autoscaling

More pods when busy, fewer when quiet.

Scale on metrics relative to requests

A HorizontalPodAutoscaler (HPA) adjusts a Deployment's replica count from metrics, most commonly average CPU utilisation as a percentage of requests (so requests must be set), or memory, or custom metrics such as queue length or requests per second. The core rule is desired = ceil(current replicas x current metric / target), ignored within a tolerance (10% by default), clamped to min and max replicas, and smoothed by stabilisation windows (by default slow to scale down). Pair it with a cluster autoscaler so new pods have nodes to run on.

The HPA scaling formula, run

I ran this with Python 3. It is a simplified model of Kubernetes behaviour for learning, not the real controller code. With a 60% CPU target: 65% is within tolerance (no change); 90% scales 4 pods to 6; 150% hits the maximum of 10; 20% on 6 pods scales down to the minimum of 2. The real controller also applies stabilisation windows and scaling policies.

import math
def desired(current_replicas, current_metric, target, tolerance=0.1, min_r=2, max_r=10):
    ratio = current_metric / target
    if abs(ratio - 1) <= tolerance:            # within tolerance: do nothing
        return current_replicas, "within tolerance"
    want = math.ceil(current_replicas * ratio)
    return max(min_r, min(max_r, want)), f"ratio {ratio:.2f}"
for replicas, cpu in [(4, 65), (4, 90), (4, 150), (6, 20), (8, 400)]:
    r, why = desired(replicas, cpu, target=60)
    print(f"{replicas} pods at {cpu:>3}% avg CPU (target 60%) -> {r} pods ({why})")

Output:

4 pods at  65% avg CPU (target 60%) -> 4 pods (within tolerance)
4 pods at  90% avg CPU (target 60%) -> 6 pods (ratio 1.50)
4 pods at 150% avg CPU (target 60%) -> 10 pods (ratio 2.50)
6 pods at  20% avg CPU (target 60%) -> 2 pods (ratio 0.33)
8 pods at 400% avg CPU (target 60%) -> 10 pods (ratio 6.67)

Scale on the real bottleneck

For queue workers, scale on queue depth (for example with KEDA) rather than CPU.

Quick check: Why must CPU requests be set for CPU-based autoscaling?

  • Utilisation is measured as a percentage of the requested CPU
  • Requests make pods faster
  • The HPA ignores CPU
  • Limits are used instead and requests are optional
Answer

Utilisation is measured as a percentage of the requested CPU — No requests, no utilisation percentage.