Lesson 9 / 25
Ordering, Retries and Dead Letters
Keep related events in order and handle messages that keep failing.
Order per key, not globally
Global ordering across a whole system does not scale, and you rarely need it. What you usually need is ordering per entity: all events for order o-1001 in sequence. Log brokers give ordering within a partition, so use the entity ID as the partition key; queues offer features like Service Bus sessions or SQS FIFO message groups. Retries can break ordering: if event 5 fails and is retried later while event 6 succeeds, the consumer sees 6 before 5. Options include blocking the partition while retrying (simple, but one bad message stalls a key), carrying a version number so consumers ignore stale updates, or making handlers order-independent. A poison message that always fails (bad data, a bug) must not block forever: after N attempts with backoff, move it to a dead-letter queue (DLQ), alert someone, fix the cause and replay it.
Ignoring stale updates with a version
If an older event arrives late, the condition fails and nothing changes.
-- each OrderStatusChanged event carries the aggregate version that produced it
UPDATE order_view
SET status = :status,
version = :event_version
WHERE order_id = :order_id
AND version < :event_version; -- stale or duplicate events update 0 rowsA DLQ nobody watches is a black hole
Dead-lettered messages are business transactions that did not happen. Alert on DLQ depth, record why each message failed, and build a tool to replay messages after a fix.
Quick check: How do log-based brokers usually guarantee that events for one order are processed in sequence?
- They order every event globally
- They retry until order is correct
- Events with the same key go to the same partition, which is ordered
- Consumers sort events by timestamp
Answer
Events with the same key go to the same partition, which is ordered — Partitioning by entity key keeps that entity's events in one ordered partition.