Lesson 24 / 25

Keeping Elasticsearch in Sync With the Database

Dual writes versus change data capture.

Getting changes from the source of truth

The tempting approach is a dual write: the application saves to the database and then indexes into Elasticsearch. It breaks in subtle ways: if the second write fails or the process crashes between them, the two drift apart, and concurrent updates can arrive out of order. More robust patterns: the transactional outbox (write the change and an outbox event in the same database transaction, then a relay publishes events to a worker that indexes them) or change data capture (CDC), which reads the database's log (for example with Debezium into Kafka) and streams changes to an indexer. Either way, make indexing idempotent by using the database primary key as the document _id, protect against out-of-order updates with external versioning (version_type=external with a monotonically increasing version from the database) and keep a full rebuild path for disaster recovery.

Idempotent, ordered indexing from change events

Sketch of an indexer consuming change events, followed by the request it sends (console syntax).

for event in change_stream():                # outbox relay or CDC topic
    doc_id  = event.primary_key               # same id every time -> idempotent
    version = event.row_version               # e.g. updated_at in microseconds
    if event.op == "delete":
        DELETE /products/_doc/{doc_id}?version={version}&version_type=external
    else:
        PUT /products/_doc/{doc_id}?version={version}&version_type=external
        { ...document built from the row... }
    # a 409 conflict means a newer version is already indexed: safe to skip
    commit_offset(event)

Reconcile periodically

Even good pipelines drift after bugs and outages. Run a periodic job that compares counts or checksums between the database and the index, and make a full reindex from the database a routine, tested operation.

Quick check: What is the main weakness of dual writes from application code?

  • It requires Kafka
  • Elasticsearch cannot accept writes from applications
  • A failure between the two writes leaves the database and the index inconsistent
  • It makes documents immutable
Answer

A failure between the two writes leaves the database and the index inconsistent — Outbox or CDC ties index updates to committed database changes.