पाठ 5 / 25
Graded Relevance and nDCG
When some passages are better than others, order matters.
Discounted cumulative gain
nDCG@k (normalised discounted cumulative gain) uses graded labels: each passage contributes a gain based on its grade, discounted by how far down the list it appears, and the total is divided by the best possible ordering so scores range from 0 to 1. Two systems that retrieve the same passages can get very different nDCG if one puts the best passage first. Use nDCG when you rank with a reranker or when position in the context matters; use recall@k when you simply need the right passage somewhere in the context.
Same passages, different order, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. Both systems return the same four passages. System A puts the perfect passage first (nDCG@4 0.993); system B puts it last (0.587).
import math
grades = {"d1": 3, "d2": 2, "d5": 1} # 3 = perfect, 2 = useful, 1 = marginal, missing = 0
def dcg(ranked, k):
return sum((2 ** grades.get(d, 0) - 1) / math.log2(i + 2) for i, d in enumerate(ranked[:k]))
ideal = sorted(grades, key=lambda d: -grades[d])
for name, ranked in {"system A": ["d1", "d2", "d9", "d5"], "system B": ["d5", "d9", "d2", "d1"]}.items():
print(name, "nDCG@4 =", round(dcg(ranked, 4) / dcg(ideal, 4), 3), "| same docs found, different order")
Output:
system A nDCG@4 = 0.993 | same docs found, different order system B nDCG@4 = 0.587 | same docs found, different order
Use nDCG to evaluate rerankers
A reranker changes order, not the set of candidates, so recall@candidate-pool stays the same while nDCG shows the gain.
त्वरित जाँच: What does nDCG capture that recall@k does not?
- Whether the answer is fluent
- Whether the best passages are ranked near the top
- The latency of the retriever
- The number of documents in the index
Answer
Whether the best passages are ranked near the top — nDCG discounts gains lower in the list.