Lesson 6 / 25
Choosing k and the Primary Metric
Pick metrics that match how the context is used.
Match the metric to the pipeline
Choose one primary metric for decisions and a few guardrail metrics. If the generator receives the top 5 passages in any order, recall@5 (or hit@5 for single-answer questions) is primary. If a reranker selects the top 3 from 50 candidates, track recall@50 for the first stage and nDCG@3 or recall@3 for the reranker. Add precision@k or context token count as a guardrail so you notice when you buy recall by flooding the context. Report metrics at a few values of k to see the curve, but decide on one.
Metric choices per pipeline shape
A starting guide.
pipeline primary guardrails
retrieve top 5 -> LLM recall@5 precision@5, context tokens
retrieve 50 -> rerank -> top 3 -> LLM recall@50 (stage 1) nDCG@3, latency
single-answer FAQ lookup hit@1 / MRR not-found accuracy
multi-hop questions recall@k (all hops) answer correctnessWrite the decision rule down
Agree in advance, for example: ship if recall@5 improves and precision@5 does not drop by more than 0.02.
Quick check: In a retrieve-then-rerank pipeline, what should the first stage be measured on?
- Precision@1 of the first stage
- Only nDCG@1
- Answer fluency
- Recall over the full candidate pool passed to the reranker
Answer
Recall over the full candidate pool passed to the reranker — The reranker can only reorder what stage one found.