# Design a Distributed Job Scheduler — High-Level & Low-Level Design Interview Problems

Source: https://www.skillbyai.com/en/design-interviews/job-scheduler

> Run jobs on time, once in effect.

## Scheduling, dispatch and execution

Requirements: users submit one-off or recurring (cron-like) jobs; jobs run close to their scheduled time; failures are retried; no job is lost; status is visible. Design: a **job store** (jobs and their next run time, indexed by `next_run_at`), **scheduler** nodes that poll for due jobs and claim them, a **queue** that buffers runnable tasks, and **workers** that execute them and report results. Claiming must be exclusive: use a conditional update or `SELECT ... FOR UPDATE SKIP LOCKED` in a relational database, or partition jobs across scheduler nodes. Delivery is usually **at-least-once** with a lease: if a worker dies, its lease expires and the task is retried, so job handlers should be **idempotent**. Add retries with exponential backoff, a dead-letter queue and per-tenant limits.

## Claiming due jobs safely

SQL-style sketch (PostgreSQL syntax for SKIP LOCKED).

```text
-- scheduler loop, every second
BEGIN;
SELECT id, payload FROM jobs
 WHERE status = 'SCHEDULED' AND next_run_at <= now()
 ORDER BY next_run_at
 LIMIT 100
 FOR UPDATE SKIP LOCKED;          -- other schedulers skip these rows
UPDATE jobs SET status = 'QUEUED' WHERE id = ANY(:ids);
COMMIT;
-> push ids to task queue

-- worker
receive task -> lease 60s (assumed) -> run handler (idempotent)
  success -> status DONE; if recurring: next_run_at = next cron time, SCHEDULED
  failure -> attempts++ ; next_run_at = now + backoff ; or DEAD after max attempts
```

## Exactly-once is an illusion here

Say at-least-once delivery plus idempotent handlers. Claiming exactly-once execution across failures is a red flag unless you explain how side effects are deduplicated.

**Quiz:** Why should job handlers be idempotent?

- [ ] Because queues reorder all jobs randomly
- [x] Because a job may run more than once after a worker failure and retry
- [ ] Because cron syntax requires it
- [ ] Because databases cannot store results

*Answer:* Because a job may run more than once after a worker failure and retry. At-least-once delivery implies duplicates.
