Lesson 14 / 25

Requests, Limits and Units

CPU in cores, memory in bytes.

What each setting does

Each container can set requests (the amount reserved for scheduling) and limits (the maximum it may use). CPU is measured in cores (250m = a quarter core); a container over its CPU limit is throttled. Memory is measured in bytes with binary (Mi, Gi) or decimal (M, G) suffixes; a container over its memory limit is OOM-killed and restarted. Set requests close to typical usage and memory limits with headroom; many teams leave CPU limits unset to avoid throttling, relying on requests for fair sharing.

Requests, limits and placement

Requests drive scheduling, limits cap usage, and together they decide QoS and eviction.

Three ideas: requests and limits, scheduling, QoS and eviction.
Figure 5.1 — Resources, scheduling and eviction.

Converting resource units, run

I ran this with Python 3. It is a simplified model of Kubernetes behaviour for learning, not the real controller code. Four containers request 1.80 cores and 1.79 GiB in total. 512M (decimal megabytes) is about 488 MiB, roughly 5% less than 512Mi, a common source of confusion.

def cpu(v):   # "250m" -> 0.25 cores
    return float(v[:-1]) / 1000 if v.endswith("m") else float(v)
def mem(v):   # binary (Mi, Gi) and decimal (M, G) suffixes -> bytes
    units = {"Ki": 2**10, "Mi": 2**20, "Gi": 2**30, "K": 10**3, "M": 10**6, "G": 10**9}
    for u in sorted(units, key=len, reverse=True):
        if v.endswith(u): return float(v[:-len(u)]) * units[u]
    return float(v)
containers = [("web", "250m", "256Mi"), ("sidecar", "50m", "64Mi"), ("worker", "1", "1Gi"), ("cache", "0.5", "512M")]
for name, c, m in containers:
    print(f"{name:<8} cpu {c:>5} = {cpu(c):.3f} cores   memory {m:>6} = {mem(m) / 2**20:,.1f} MiB")
print(f"pod total requests: {sum(cpu(c) for _, c, _ in containers):.2f} cores, {sum(mem(m) for *_, m in containers) / 2**30:.2f} GiB")
print("note: 512M (decimal) is", mem("512M") / mem("512Mi"), "x 512Mi")

Output:

web      cpu  250m = 0.250 cores   memory  256Mi = 256.0 MiB
sidecar  cpu   50m = 0.050 cores   memory   64Mi = 64.0 MiB
worker   cpu     1 = 1.000 cores   memory    1Gi = 1,024.0 MiB
cache    cpu   0.5 = 0.500 cores   memory   512M = 488.3 MiB
pod total requests: 1.80 cores, 1.79 GiB
note: 512M (decimal) is 0.95367431640625 x 512Mi

Measure before setting

Use real usage metrics (for example from metrics-server, Prometheus or VPA recommendations) to set requests, then revisit them regularly.

Quick check: What happens when a container exceeds its memory limit?

  • Nothing
  • It is throttled
  • It gets more memory automatically
  • It is OOM-killed and restarted
Answer

It is OOM-killed and restarted — Memory is not compressible; CPU is.