Lesson 7 / 25
How Analyzers Work
Character filters, a tokenizer, token filters.
The three stages
An analyzer has zero or more character filters (rewrite the raw string, for example html_strip or a mapping that turns "&" into "and"), exactly one tokenizer (splits the string into tokens, for example standard, whitespace, keyword or pattern) and zero or more token filters (change, remove or add tokens, for example lowercase, stop, asciifolding, stemmer, synonym_graph). The default standard analyzer uses the standard tokenizer (Unicode word boundaries) plus the lowercase filter; stop words are off by default. The analyzer for a text field runs at index time, and the same analyzer (or a configured search_analyzer) runs on query text, so both sides produce comparable terms. Language analyzers such as english add stop words and stemming.
Turning text into searchable terms
An analyzer runs character filters, a tokenizer and token filters, at index time and at search time.
Following a string through the pipeline
Conceptual trace of a custom analyzer.
input: "<p>Running SHOES & Café</p>"
html_strip: "Running SHOES & Café"
standard tok.: [Running] [SHOES] [Café] (punctuation dropped)
lowercase: [running] [shoes] [café]
asciifolding: [running] [shoes] [cafe]
english stem: [run] [shoe] [cafe]
These terms go into the inverted index with their positions.A kitchen prep line
Character filters wash the vegetables, the tokenizer chops them into pieces, and token filters season, trim or throw away pieces before they go into the pot.
Quick check: How many tokenizers does an analyzer have?
- Exactly one
- Zero or more
- At least two
- One per token filter
Answer
Exactly one — Character and token filters are optional and can be many; the tokenizer is required and single.