पाठ 4 / 25

Recall@k, Precision@k, Hit Rate and MRR

The core metrics for a ranked list.

Four questions about the top k

Recall@k: what fraction of the relevant passages appear in the top k (did we find what exists?). Precision@k: what fraction of the top k are relevant (how much noise did we send?). Hit rate@k (or success@k): did at least one relevant passage appear. Reciprocal rank: 1 divided by the position of the first relevant passage, averaged across queries as MRR (how early does the first good result appear?). For RAG, recall at the k you actually send is usually the first metric to watch, because a passage outside the context cannot be used.

Did we find it, and how high?

Different metrics answer different questions about a ranked list.

Four ideas: recall and precision, rank, graded gain, context noise.
Figure 2.1 — Recall, rank, gain and noise.

Computing the four metrics on three queries, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. q1 finds its one relevant document at rank 3; q2 finds one of two relevant documents at rank 1; q3 finds nothing. Mean recall@3 is 0.50, precision@3 0.22, hit@3 0.67 and MRR 0.44.

runs = {  # query -> (ranked doc ids returned, set of relevant ids)
    "q1": (["d3", "d7", "d1", "d9", "d2"], {"d1"}),
    "q2": (["d4", "d5", "d8", "d6", "d0"], {"d4", "d6"}),
    "q3": (["d2", "d0", "d5", "d1", "d3"], {"d8"}),
}
k = 3
def recall(r, rel): return len(set(r[:k]) & rel) / len(rel)
def precision(r, rel): return len(set(r[:k]) & rel) / k
def hit(r, rel): return 1.0 if set(r[:k]) & rel else 0.0
def rr(r, rel): return next((1 / (i + 1) for i, d in enumerate(r) if d in rel), 0.0)
print(f"query  recall@{k}  precision@{k}  hit@{k}  reciprocal rank")
for q, (r, rel) in runs.items():
    print(f"{q:<6} {recall(r, rel):>8.2f} {precision(r, rel):>12.2f} {hit(r, rel):>7.0f} {rr(r, rel):>16.2f}")
n = len(runs)
print(f"mean   {sum(recall(*v) for v in runs.values()) / n:>8.2f} {sum(precision(*v) for v in runs.values()) / n:>12.2f}"
      f" {sum(hit(*v) for v in runs.values()) / n:>7.2f} {sum(rr(*v) for v in runs.values()) / n:>16.2f}  (MRR)")

Output:

query  recall@3  precision@3  hit@3  reciprocal rank
q1         1.00         0.33       1             0.33
q2         0.50         0.33       1             1.00
q3         0.00         0.00       0             0.00
mean       0.50         0.22    0.67             0.44  (MRR)

Report the k you send

Measure recall at the number of passages that actually reach the model, not at a convenient k from a benchmark.

त्वरित जाँच: Which metric tells you how early the first relevant passage appears?

  • Recall@k
  • Mean reciprocal rank (MRR)
  • Precision@k
  • Total tokens
Answer

Mean reciprocal rank (MRR) — MRR rewards a relevant result at rank 1.