Lesson 4 / 25

Document CRUD and the Bulk API

Index, get, update, delete, in batches.

Writing documents

PUT /index/_doc/<id> creates or replaces a document with a known id; POST /index/_doc lets Elasticsearch generate one; PUT /index/_create/<id> fails if the id already exists. GET /index/_doc/<id> reads a document by id in real time, while searches only see changes after a refresh. POST /index/_update/<id> applies a partial update (internally Elasticsearch re-indexes the whole document, since segments are immutable), and DELETE removes one. Every write returns a _seq_no and _primary_term, which you can pass back as if_seq_no/if_primary_term for optimistic concurrency control. For many documents use the bulk API: a newline-delimited JSON (NDJSON) body of action lines and source lines, which must end with a newline.

Getting data in with the right shape

Documents are written with CRUD and bulk APIs, and the mapping decides how every field is indexed.

Three ideas: document APIs, field types, explicit versus dynamic mapping.
Figure 2.1 — Documents flow through the mapping into indexed fields.

Single writes and a bulk request

Kibana Dev Tools console syntax; send the same requests with curl or a client library.

PUT /products/_doc/42
{ "name": "Rain jacket", "price": 120, "tags": ["outdoor"] }

POST /products/_update/42
{ "doc": { "price": 99 } }

GET /products/_doc/42

DELETE /products/_doc/42

# bulk: one action line, then (for index/create/update) one source line
POST /_bulk
{ "index":  { "_index": "products", "_id": "1" } }
{ "name": "Trail shoe", "price": 89.9 }
{ "create": { "_index": "products", "_id": "2" } }
{ "name": "Wool socks", "price": 12 }
{ "update": { "_index": "products", "_id": "1" } }
{ "doc": { "price": 79.9 } }
{ "delete": { "_index": "products", "_id": "3" } }

Always check the bulk errors flag

A bulk call can return HTTP 200 while individual items failed. Check the top-level "errors" field and inspect each item's status, then retry only the failed items (for example 429 rejections) with backoff.

Quick check: What format does the bulk API body use?

  • A single JSON array of documents
  • Newline-delimited JSON with action and source lines
  • CSV with a header row
  • XML with one element per document
Answer

Newline-delimited JSON with action and source lines — NDJSON lets the coordinating node split the request without parsing it as one big document.