# Choosing k and the Primary Metric — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/m-choose

> Pick metrics that match how the context is used.

## Match the metric to the pipeline

Choose one **primary metric** for decisions and a few **guardrail metrics**. If the generator receives the top 5 passages in any order, **recall@5** (or hit@5 for single-answer questions) is primary. If a reranker selects the top 3 from 50 candidates, track **recall@50** for the first stage and **nDCG@3** or recall@3 for the reranker. Add **precision@k** or context token count as a guardrail so you notice when you buy recall by flooding the context. Report metrics at a few values of k to see the curve, but decide on one.

## Metric choices per pipeline shape

A starting guide.

```text
pipeline                                 primary            guardrails
retrieve top 5 -> LLM                    recall@5           precision@5, context tokens
retrieve 50 -> rerank -> top 3 -> LLM    recall@50 (stage 1) nDCG@3, latency
single-answer FAQ lookup                 hit@1 / MRR        not-found accuracy
multi-hop questions                      recall@k (all hops) answer correctness
```

## Write the decision rule down

Agree in advance, for example: ship if recall@5 improves and precision@5 does not drop by more than 0.02.

**Quiz:** In a retrieve-then-rerank pipeline, what should the first stage be measured on?

- [ ] Precision@1 of the first stage
- [ ] Only nDCG@1
- [ ] Answer fluency
- [x] Recall over the full candidate pool passed to the reranker

*Answer:* Recall over the full candidate pool passed to the reranker. The reranker can only reorder what stage one found.
