Lesson 6 / 25

Autoscaling

Configure autoscaling that reacts in time without flapping.

Matching capacity to demand automatically

Autoscaling adds instances when load rises and removes them when it falls. Target tracking keeps a metric near a target, such as 60% CPU or 100 requests per second per instance; step scaling adds fixed amounts at thresholds; scheduled scaling prepares for known peaks (an exam result day, a festival sale); predictive scaling uses historical patterns. Choose a metric that tracks load: CPU for compute-bound services, request rate or concurrency for I/O-bound services, queue depth for workers. Autoscaling is never instant: new instances need to boot, pull images, warm caches and pass health checks, so keep headroom and scale out early. Use cooldowns or stabilisation windows to prevent flapping, scale in slowly, and drain connections before terminating instances. Remember that downstream systems such as databases do not scale automatically with your web tier.

A Kubernetes HorizontalPodAutoscaler

Scale between 3 and 30 pods at 60% CPU, scaling up quickly and down cautiously.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: checkout
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: checkout
  minReplicas: 3
  maxReplicas: 30
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 60
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0
      policies: [{ type: Percent, value: 100, periodSeconds: 60 }]
    scaleDown:
      stabilizationWindowSeconds: 300
      policies: [{ type: Percent, value: 20, periodSeconds: 60 }]

Pre-scale for known events

If a marketing push email goes out at 10:00, traffic arrives at 10:01, faster than any reactive autoscaler. Schedule extra capacity before the event and let autoscaling handle the rest.

Quick check: Which metric is usually best for autoscaling background workers that read from a queue?

  • Queue depth or age of the oldest message
  • Disk size
  • Number of deployments per day
  • Memory of the load balancer
Answer

Queue depth or age of the oldest message — Queue backlog directly reflects how far workers are behind demand.