Lesson 18 / 25

Vector Search and Hybrid Retrieval

Semantic matching with dense_vector and knn.

Embeddings next to text

An embedding model turns text (or images) into a vector so that similar meanings are close together. Elasticsearch stores these in a dense_vector field with dims and a similarity (cosine, dot_product, l2_norm), indexed with an HNSW graph for approximate k-nearest-neighbour search. A search request's knn option takes a query_vector, the number of results k and num_candidates (more candidates per shard means better recall but more work), and can include a filter. Hybrid search combines a BM25 query with knn in one request, either by adding boosted scores or with reciprocal rank fusion (RRF), which merges rankings rather than raw scores. Newer versions add features such as semantic_text fields, retrievers and quantised vectors, and some depend on licence tier, so check the docs for your version.

A vector field and a hybrid query

Kibana Dev Tools console syntax; send the same requests with curl or a client library.

PUT /docs
{
  "mappings": {
    "properties": {
      "title":     { "type": "text" },
      "lang":      { "type": "keyword" },
      "embedding": { "type": "dense_vector", "dims": 384, "index": true, "similarity": "cosine" }
    }
  }
}

GET /docs/_search
{
  "query": {
    "match": { "title": { "query": "reset my password", "boost": 0.3 } }
  },
  "knn": {
    "field": "embedding",
    "query_vector": [ 0.012, -0.044, 0.093 ],
    "k": 10,
    "num_candidates": 100,
    "filter": { "term": { "lang": "en" } },
    "boost": 0.7
  },
  "size": 10
}

# query_vector is shortened here; it must have exactly 384 values
# and come from the same model used to embed the documents.

Keep BM25 in the mix

Vectors are great at paraphrases but weak at exact identifiers, SKUs and rare names. Hybrid retrieval usually beats either method alone; evaluate on your own queries.

Quick check: What does increasing num_candidates in a knn search generally do?

  • Improves recall at the cost of more work per shard
  • Changes the vector dimensions
  • Disables filters
  • Switches scoring to BM25
Answer

Improves recall at the cost of more work per shard — More candidates are explored in the HNSW graph.