SkillByAIOpen interactive version →

Lesson 15 / 25

How the Scheduler Places Pods

Filter nodes, then score them.

Requests must fit

For each pending pod, the scheduler filters nodes that can run it (enough unrequested CPU and memory, matching node selectors and affinity, tolerated taints, available ports) and then scores the rest to pick one (spreading, packing, affinity preferences). Scheduling uses requests, not actual usage: a node can be busy but still accept pods whose requests fit, or idle yet refuse pods whose requests do not. A pod that fits nowhere stays Pending until capacity appears (for example via a cluster autoscaler).

Placing seven pods on three nodes, run

I ran this with Python 3. It is a simplified model of Kubernetes behaviour for learning, not the real controller code. Six pods are placed using a simple "most free CPU" score. The batch pod requesting 3 cores stays Pending: 2.0 cores are free in total, but not 3 on any single node. Real scheduling adds affinity, taints and spreading rules.

nodes = {"node-a": {"cpu": 4.0, "mem": 16}, "node-b": {"cpu": 4.0, "mem": 16}, "node-c": {"cpu": 2.0, "mem": 8}}
free = {n: dict(v) for n, v in nodes.items()}
pods = [("api-1", 1.5, 3), ("api-2", 1.5, 3), ("worker-1", 2.0, 6), ("web-1", 0.5, 1), ("web-2", 0.5, 1),
        ("worker-2", 2.0, 6), ("batch-1", 3.0, 4)]
for name, c, m in pods:
    fits = [n for n in free if free[n]["cpu"] >= c and free[n]["mem"] >= m]
    if not fits:
        print(f"{name:<9} requests {c} cpu / {m} GiB -> Pending (no node has enough free requests)")
        continue
    node = max(fits, key=lambda n: free[n]["cpu"])       # simple "most free CPU" scoring
    free[node]["cpu"] -= c; free[node]["mem"] -= m
    print(f"{name:<9} requests {c} cpu / {m} GiB -> {node}")
print("free after scheduling:", {n: (round(v['cpu'], 1), v['mem']) for n, v in free.items()})

Output:

api-1     requests 1.5 cpu / 3 GiB -> node-a
api-2     requests 1.5 cpu / 3 GiB -> node-b
worker-1  requests 2.0 cpu / 6 GiB -> node-a
web-1     requests 0.5 cpu / 1 GiB -> node-b
web-2     requests 0.5 cpu / 1 GiB -> node-b
worker-2  requests 2.0 cpu / 6 GiB -> node-c
batch-1   requests 3.0 cpu / 4 GiB -> Pending (no node has enough free requests)
free after scheduling: {'node-a': (0.5, 7), 'node-b': (1.5, 11), 'node-c': (0.0, 2)}

Read Pending events

kubectl describe pod shows the scheduler's reason (for example "Insufficient cpu"), which tells you what to fix.

Quick check: What does the scheduler compare against node capacity?

  • The pods' resource requests
  • Their actual current usage only
  • Their image size
  • Their names
Answer

The pods' resource requests — Requests drive placement.