SkillByAIOpen interactive version →

Lesson 6 / 25

Choosing k and the Primary Metric

Pick metrics that match how the context is used.

Match the metric to the pipeline

Choose one primary metric for decisions and a few guardrail metrics. If the generator receives the top 5 passages in any order, recall@5 (or hit@5 for single-answer questions) is primary. If a reranker selects the top 3 from 50 candidates, track recall@50 for the first stage and nDCG@3 or recall@3 for the reranker. Add precision@k or context token count as a guardrail so you notice when you buy recall by flooding the context. Report metrics at a few values of k to see the curve, but decide on one.

Metric choices per pipeline shape

A starting guide.

pipeline                                 primary            guardrails
retrieve top 5 -> LLM                    recall@5           precision@5, context tokens
retrieve 50 -> rerank -> top 3 -> LLM    recall@50 (stage 1) nDCG@3, latency
single-answer FAQ lookup                 hit@1 / MRR        not-found accuracy
multi-hop questions                      recall@k (all hops) answer correctness

Write the decision rule down

Agree in advance, for example: ship if recall@5 improves and precision@5 does not drop by more than 0.02.

Quick check: In a retrieve-then-rerank pipeline, what should the first stage be measured on?

  • Precision@1 of the first stage
  • Only nDCG@1
  • Answer fluency
  • Recall over the full candidate pool passed to the reranker
Answer

Recall over the full candidate pool passed to the reranker — The reranker can only reorder what stage one found.