Lesson 22 / 25
Performance Tuning
Indexing throughput and query latency.
Where the time goes
Indexing: send documents with the bulk API in batches sized by bytes (a few MB to a few tens of MB is a common starting point; measure on your hardware), use several concurrent clients, and back off on 429 responses. For a big initial load, set refresh_interval to -1 (or a longer value such as 30s) and replicas to 0, then restore both afterwards. Queries: put exact conditions in filter context so they can be cached, query keyword fields for exact matches, avoid leading wildcards and heavy scripts on hot paths, return only needed fields, avoid deep from pagination, and keep aggregation size reasonable. Nodes: give the JVM heap no more than about half the RAM (leaving the rest for the filesystem cache Lucene relies on) and keep it below the compressed-pointers threshold, around 31 GB; recent versions size the heap automatically by default.
Fast, secure and in sync
Production readiness means tuned indexing and queries, locked-down access and a reliable pipeline from the primary database.
Settings for a bulk load
Kibana Dev Tools console syntax; send the same requests with curl or a client library.
# before the load
PUT /products_v2/_settings
{
"index": { "refresh_interval": "-1", "number_of_replicas": 0 }
}
# ... run bulk requests from several workers ...
# after the load
PUT /products_v2/_settings
{
"index": { "refresh_interval": "1s", "number_of_replicas": 1 }
}
POST /products_v2/_refresh
# find slow queries: profile one request
GET /products/_search
{
"profile": true,
"query": { "match": { "name": "jacket" } }
}Use slow logs and profile
Enable search and indexing slow logs with thresholds, and use "profile": true on suspicious queries. Fix the worst query patterns before adding hardware.
Quick check: Which change typically speeds up a large one-off bulk load?
- Running every condition in must instead of filter
- Using one document per request
- Setting refresh_interval to 1ms
- Disabling refresh and replicas during the load, then restoring them
Answer
Disabling refresh and replicas during the load, then restoring them — Fewer refreshes and no replication during the load reduce work.