पाठ 4 / 25
Recall@k, Precision@k, Hit Rate and MRR
The core metrics for a ranked list.
Four questions about the top k
Recall@k: what fraction of the relevant passages appear in the top k (did we find what exists?). Precision@k: what fraction of the top k are relevant (how much noise did we send?). Hit rate@k (or success@k): did at least one relevant passage appear. Reciprocal rank: 1 divided by the position of the first relevant passage, averaged across queries as MRR (how early does the first good result appear?). For RAG, recall at the k you actually send is usually the first metric to watch, because a passage outside the context cannot be used.
Did we find it, and how high?
Different metrics answer different questions about a ranked list.
Computing the four metrics on three queries, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. q1 finds its one relevant document at rank 3; q2 finds one of two relevant documents at rank 1; q3 finds nothing. Mean recall@3 is 0.50, precision@3 0.22, hit@3 0.67 and MRR 0.44.
runs = { # query -> (ranked doc ids returned, set of relevant ids)
"q1": (["d3", "d7", "d1", "d9", "d2"], {"d1"}),
"q2": (["d4", "d5", "d8", "d6", "d0"], {"d4", "d6"}),
"q3": (["d2", "d0", "d5", "d1", "d3"], {"d8"}),
}
k = 3
def recall(r, rel): return len(set(r[:k]) & rel) / len(rel)
def precision(r, rel): return len(set(r[:k]) & rel) / k
def hit(r, rel): return 1.0 if set(r[:k]) & rel else 0.0
def rr(r, rel): return next((1 / (i + 1) for i, d in enumerate(r) if d in rel), 0.0)
print(f"query recall@{k} precision@{k} hit@{k} reciprocal rank")
for q, (r, rel) in runs.items():
print(f"{q:<6} {recall(r, rel):>8.2f} {precision(r, rel):>12.2f} {hit(r, rel):>7.0f} {rr(r, rel):>16.2f}")
n = len(runs)
print(f"mean {sum(recall(*v) for v in runs.values()) / n:>8.2f} {sum(precision(*v) for v in runs.values()) / n:>12.2f}"
f" {sum(hit(*v) for v in runs.values()) / n:>7.2f} {sum(rr(*v) for v in runs.values()) / n:>16.2f} (MRR)")
Output:
query recall@3 precision@3 hit@3 reciprocal rank q1 1.00 0.33 1 0.33 q2 0.50 0.33 1 1.00 q3 0.00 0.00 0 0.00 mean 0.50 0.22 0.67 0.44 (MRR)
Report the k you send
Measure recall at the number of passages that actually reach the model, not at a convenient k from a benchmark.
त्वरित जाँच: Which metric tells you how early the first relevant passage appears?
- Recall@k
- Mean reciprocal rank (MRR)
- Precision@k
- Total tokens
Answer
Mean reciprocal rank (MRR) — MRR rewards a relevant result at rank 1.