# Reference-Based Metrics: Exact Match and Token F1 — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/a-ref

> Cheap scores when you have a reference answer.

## Useful for short factual answers

If each question has a short **reference answer**, simple metrics work: **exact match (EM)** after normalising case and punctuation, and **token F1**, which measures word overlap between prediction and reference. They are cheap, deterministic and good for short facts (dates, numbers, names). They are poor for long or freely worded answers: a correct paraphrase can score low and a wrong answer sharing many words can score high. Use them for short-answer slices and pair them with judgement-based checks for the rest.

## Correct, grounded, cited

After retrieval, check that answers are correct, supported by the context and cite the right sources.

![Three ideas: reference match, faithfulness, citations.](assets/figures/rag-evaluation/section-5-map.svg) — Figure 5.1 — Reference match, faithfulness and citations.

## Exact match and token F1, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. Against the reference answer, an identical answer scores EM 1 and F1 1.00, a correct paraphrase scores F1 0.77 with EM 0, a wrong 30-day answer F1 0.25, and a partial answer 0.40.

```python
import re, collections
def norm(s):
    return re.sub(r"[^a-z0-9 ]", " ", s.lower()).split()
def em(pred, gold): return float(norm(pred) == norm(gold))
def f1(pred, gold):
    p, g = norm(pred), norm(gold)
    common = sum((collections.Counter(p) & collections.Counter(g)).values())
    if common == 0: return 0.0
    prec, rec = common / len(p), common / len(g)
    return 2 * prec * rec / (prec + rec)
gold = "14 days after receiving the item"
for pred in ["14 days after receiving the item", "Within 14 days of receiving the item.", "30 days", "Refunds take 14 days."]:
    print(f"EM={em(pred, gold):.0f}  F1={f1(pred, gold):.2f}  | {pred}")
```

Output:

```
EM=1  F1=1.00  | 14 days after receiving the item
EM=0  F1=0.77  | Within 14 days of receiving the item.
EM=0  F1=0.25  | 30 days
EM=0  F1=0.40  | Refunds take 14 days.
```

## Extract the key value first

For numeric questions, extract the number from the answer and compare it exactly; word overlap is too forgiving.

**Quiz:** Where does token F1 work poorly?

- [ ] Single names
- [ ] Short dates
- [ ] Exact numbers
- [x] Long, freely worded answers where correct paraphrases share few words with the reference

*Answer:* Long, freely worded answers where correct paraphrases share few words with the reference. Overlap is a proxy, not understanding.
