SkillByAIOpen interactive version →

Lesson 20 / 25

Design a Distributed Job Scheduler

Run jobs on time, once in effect.

Scheduling, dispatch and execution

Requirements: users submit one-off or recurring (cron-like) jobs; jobs run close to their scheduled time; failures are retried; no job is lost; status is visible. Design: a job store (jobs and their next run time, indexed by next_run_at), scheduler nodes that poll for due jobs and claim them, a queue that buffers runnable tasks, and workers that execute them and report results. Claiming must be exclusive: use a conditional update or SELECT ... FOR UPDATE SKIP LOCKED in a relational database, or partition jobs across scheduler nodes. Delivery is usually at-least-once with a lease: if a worker dies, its lease expires and the task is retried, so job handlers should be idempotent. Add retries with exponential backoff, a dead-letter queue and per-tenant limits.

Claiming due jobs safely

SQL-style sketch (PostgreSQL syntax for SKIP LOCKED).

-- scheduler loop, every second
BEGIN;
SELECT id, payload FROM jobs
 WHERE status = 'SCHEDULED' AND next_run_at <= now()
 ORDER BY next_run_at
 LIMIT 100
 FOR UPDATE SKIP LOCKED;          -- other schedulers skip these rows
UPDATE jobs SET status = 'QUEUED' WHERE id = ANY(:ids);
COMMIT;
-> push ids to task queue

-- worker
receive task -> lease 60s (assumed) -> run handler (idempotent)
  success -> status DONE; if recurring: next_run_at = next cron time, SCHEDULED
  failure -> attempts++ ; next_run_at = now + backoff ; or DEAD after max attempts

Exactly-once is an illusion here

Say at-least-once delivery plus idempotent handlers. Claiming exactly-once execution across failures is a red flag unless you explain how side effects are deduplicated.

Quick check: Why should job handlers be idempotent?

  • Because queues reorder all jobs randomly
  • Because a job may run more than once after a worker failure and retry
  • Because cron syntax requires it
  • Because databases cannot store results
Answer

Because a job may run more than once after a worker failure and retry — At-least-once delivery implies duplicates.