Lesson 1 / 25
What Elasticsearch Is
Search and analytics over a REST API.
A search and analytics engine
Elasticsearch is a distributed search and analytics engine built on Apache Lucene, the Java library that does the low-level indexing and scoring. You send JSON documents over a REST API, and Elasticsearch makes them searchable in near real time (by default a new segment becomes visible after a refresh, roughly every second). Typical uses are product and site search, log and metrics analytics (the "ELK" stack with Logstash/Beats and Kibana), security analytics and, more recently, vector and hybrid search. OpenSearch is a community fork of Elasticsearch 7.10 created after Elastic changed its licence in 2021; the core concepts in this course apply to both, but APIs and features have diverged since, so check the docs for the engine and version you run.
A distributed search engine on Lucene
Elasticsearch stores JSON documents in inverted indexes, spreads them over shards and answers search and analytics queries over HTTP.
Talking to Elasticsearch
Kibana Dev Tools console syntax; send the same requests with curl or a client library.
# cluster and version information
GET /
# store a document (the index is created on first write)
PUT /products/_doc/1
{
"name": "Trail running shoe",
"brand": "Northpeak",
"price": 89.90
}
# full-text search
GET /products/_search
{
"query": { "match": { "name": "running shoes" } }
}It is not your primary database
Elasticsearch is usually a secondary, rebuildable copy of data that lives in a system of record such as PostgreSQL. It has no multi-document transactions, so keep the source of truth elsewhere and design for reindexing.
Quick check: Which library provides the core indexing and scoring inside Elasticsearch?
- RocksDB
- Apache Kafka
- Apache Lucene
- Apache Spark
Answer
Apache Lucene — Each Elasticsearch shard is a Lucene index.