Lesson 13 / 25

Evaluating Rerankers, Hybrid Search and Query Rewriting

Measure each added stage against the pipeline without it.

Every stage must earn its latency

Pipelines grow stages: hybrid search (keyword plus vector), query rewriting (an LLM reformulates the question), rerankers (a cross-encoder or LLM reorders candidates). Each adds latency, cost and failure modes. Evaluate with an ablation table: the baseline, then each stage added alone, then combinations, on the same queries, measuring recall at the candidate pool, ranking quality at the sent k, latency and cost. Keep a stage only if it improves the primary metric beyond the noise, on the slices you care about, at acceptable latency.

An ablation table template

Fill in measured values; these are blanks, not results.

variant                          recall@50   nDCG@5   p95 ms   cost/1k queries
vector only                      ___         ___      ___      ___
+ keyword (hybrid, RRF)          ___         ___      ___      ___
+ reranker                       ___         ___      ___      ___
+ query rewriting                ___         ___      ___      ___
hybrid + reranker                ___         ___      ___      ___

Test the stage that failed most

Pick the next experiment from the failure analysis: if most misses are vocabulary mismatches, try hybrid or rewriting before a reranker.

Quick check: When should an added retrieval stage be kept?

  • Always, more stages are better
  • When it improves the primary metric beyond noise at acceptable latency and cost
  • Only if it is the newest technique
  • When it lowers recall
Answer

When it improves the primary metric beyond noise at acceptable latency and cost — Ablations show what each stage contributes.