Lesson 9 / 25
Debugging With the _analyze API
See exactly which terms are produced.
Test before you index
The _analyze API returns the tokens an analyzer produces, with positions and offsets. You can test a built-in analyzer, an ad-hoc combination of tokenizer and filters, or the analyzer configured for a specific field of an existing index ("field": "title"). When a query does not match a document you expect, analyse both the document text and the query text with the field's analyzers and compare the terms; the mismatch is usually obvious. The explain option shows the output of each stage.
Three ways to call _analyze
Kibana Dev Tools console syntax; send the same requests with curl or a client library.
# a built-in analyzer
POST /_analyze
{
"analyzer": "english",
"text": "The runners were running quickly"
}
# an ad-hoc chain
POST /_analyze
{
"tokenizer": "standard",
"filter": [ "lowercase", "asciifolding" ],
"text": "Crème Brûlée"
}
# whatever the mapping uses for a field
POST /catalog/_analyze
{
"field": "title",
"text": "Smart TV 55 inch"
}Use the field form in tests
Calling _analyze with "field" tests the real configuration, including custom filters, so add it to automated checks whenever you change analysis settings.
Quick check: A match query fails to find a document. What is a good first debugging step?
- Switch the field to dense_vector
- Increase the number of shards
- Run _analyze on both the document text and the query text with the field's analyzers
- Restart the node
Answer
Run _analyze on both the document text and the query text with the field's analyzers — Comparing produced terms reveals analysis mismatches.