पाठ 14 / 25
Reference-Based Metrics: Exact Match and Token F1
Cheap scores when you have a reference answer.
Useful for short factual answers
If each question has a short reference answer, simple metrics work: exact match (EM) after normalising case and punctuation, and token F1, which measures word overlap between prediction and reference. They are cheap, deterministic and good for short facts (dates, numbers, names). They are poor for long or freely worded answers: a correct paraphrase can score low and a wrong answer sharing many words can score high. Use them for short-answer slices and pair them with judgement-based checks for the rest.
Correct, grounded, cited
After retrieval, check that answers are correct, supported by the context and cite the right sources.
Exact match and token F1, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. Against the reference answer, an identical answer scores EM 1 and F1 1.00, a correct paraphrase scores F1 0.77 with EM 0, a wrong 30-day answer F1 0.25, and a partial answer 0.40.
import re, collections
def norm(s):
return re.sub(r"[^a-z0-9 ]", " ", s.lower()).split()
def em(pred, gold): return float(norm(pred) == norm(gold))
def f1(pred, gold):
p, g = norm(pred), norm(gold)
common = sum((collections.Counter(p) & collections.Counter(g)).values())
if common == 0: return 0.0
prec, rec = common / len(p), common / len(g)
return 2 * prec * rec / (prec + rec)
gold = "14 days after receiving the item"
for pred in ["14 days after receiving the item", "Within 14 days of receiving the item.", "30 days", "Refunds take 14 days."]:
print(f"EM={em(pred, gold):.0f} F1={f1(pred, gold):.2f} | {pred}")
Output:
EM=1 F1=1.00 | 14 days after receiving the item EM=0 F1=0.77 | Within 14 days of receiving the item. EM=0 F1=0.25 | 30 days EM=0 F1=0.40 | Refunds take 14 days.
Extract the key value first
For numeric questions, extract the number from the answer and compare it exactly; word overlap is too forgiving.
त्वरित जाँच: Where does token F1 work poorly?
- Single names
- Short dates
- Exact numbers
- Long, freely worded answers where correct paraphrases share few words with the reference
Answer
Long, freely worded answers where correct paraphrases share few words with the reference — Overlap is a proxy, not understanding.