# Recall@k, Precision@k, Hit Rate and MRR — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/m-basic

> The core metrics for a ranked list.

## Four questions about the top k

**Recall@k**: what fraction of the relevant passages appear in the top k (did we find what exists?). **Precision@k**: what fraction of the top k are relevant (how much noise did we send?). **Hit rate@k** (or success@k): did at least one relevant passage appear. **Reciprocal rank**: 1 divided by the position of the first relevant passage, averaged across queries as **MRR** (how early does the first good result appear?). For RAG, recall at the k you actually send is usually the first metric to watch, because a passage outside the context cannot be used.

## Did we find it, and how high?

Different metrics answer different questions about a ranked list.

![Four ideas: recall and precision, rank, graded gain, context noise.](assets/figures/rag-evaluation/section-2-map.svg) — Figure 2.1 — Recall, rank, gain and noise.

## Computing the four metrics on three queries, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. q1 finds its one relevant document at rank 3; q2 finds one of two relevant documents at rank 1; q3 finds nothing. Mean recall@3 is 0.50, precision@3 0.22, hit@3 0.67 and MRR 0.44.

```python
runs = {  # query -> (ranked doc ids returned, set of relevant ids)
    "q1": (["d3", "d7", "d1", "d9", "d2"], {"d1"}),
    "q2": (["d4", "d5", "d8", "d6", "d0"], {"d4", "d6"}),
    "q3": (["d2", "d0", "d5", "d1", "d3"], {"d8"}),
}
k = 3
def recall(r, rel): return len(set(r[:k]) & rel) / len(rel)
def precision(r, rel): return len(set(r[:k]) & rel) / k
def hit(r, rel): return 1.0 if set(r[:k]) & rel else 0.0
def rr(r, rel): return next((1 / (i + 1) for i, d in enumerate(r) if d in rel), 0.0)
print(f"query  recall@{k}  precision@{k}  hit@{k}  reciprocal rank")
for q, (r, rel) in runs.items():
    print(f"{q:<6} {recall(r, rel):>8.2f} {precision(r, rel):>12.2f} {hit(r, rel):>7.0f} {rr(r, rel):>16.2f}")
n = len(runs)
print(f"mean   {sum(recall(*v) for v in runs.values()) / n:>8.2f} {sum(precision(*v) for v in runs.values()) / n:>12.2f}"
      f" {sum(hit(*v) for v in runs.values()) / n:>7.2f} {sum(rr(*v) for v in runs.values()) / n:>16.2f}  (MRR)")
```

Output:

```
query  recall@3  precision@3  hit@3  reciprocal rank
q1         1.00         0.33       1             0.33
q2         0.50         0.33       1             1.00
q3         0.00         0.00       0             0.00
mean       0.50         0.22    0.67             0.44  (MRR)
```

## Report the k you send

Measure recall at the number of passages that actually reach the model, not at a convenient k from a benchmark.

**Quiz:** Which metric tells you how early the first relevant passage appears?

- [ ] Recall@k
- [x] Mean reciprocal rank (MRR)
- [ ] Precision@k
- [ ] Total tokens

*Answer:* Mean reciprocal rank (MRR). MRR rewards a relevant result at rank 1.
