Lesson 11 / 25
Comparing Retrievers on Your Queries
Benchmarks do not know your users' queries.
Your data decides
Public leaderboards rank embedding models and retrievers on public datasets, but your queries have their own vocabulary, typos, languages and document styles. Compare candidates (keyword, different embedding models, hybrid) on your evaluation set, with the same chunks and the same k, reporting overall and per-slice metrics plus latency and cost. Small differences need paired statistics; large differences on a meaningful slice can decide alone. The toy below shows how one property of real queries, typos, can flip the ranking of two simple retrievers.
Change one component, measure the effect
Retrievers, chunking and rerankers are design choices you can test, not opinions.
Word versus character n-gram TF-IDF on typo queries, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. Eight help articles and sixteen queries, half with typos. Word-level TF-IDF finds the right article first for 0.88 of clean queries but only 0.25 of typo queries; character 3 to 5-gram TF-IDF gets all sixteen right. A toy set this small only illustrates the method.
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
docs = [
"How to reset your account password from the login page",
"Refunds are issued within fourteen days of receiving the item",
"Shipping to international addresses takes ten business days",
"Change the billing address in account settings",
"Two factor authentication setup with an authenticator app",
"Cancel a subscription before the renewal date to avoid charges",
"Export your invoices as PDF from the billing page",
"Delete your account permanently and erase personal data",
]
queries = [ # (query, index of relevant doc); some contain typos as real users type
("reset password", 0), ("pasword resett", 0), ("refund how many days", 1), ("refnd timeline", 1),
("international shipping time", 2), ("shiping abroad", 2), ("update billing address", 3), ("biling adress", 3),
("set up two factor", 4), ("2fa authenticater", 4), ("cancel subscription", 5), ("cancle subscripton", 5),
("download invoice pdf", 6), ("invoce pdf", 6), ("delete account", 7), ("delet acount", 7),
]
def hit_at(vec, k=1):
D = vec.fit_transform(docs); Q = vec.transform([q for q, _ in queries])
sims = (Q @ D.T).toarray()
top = np.argsort(-sims, axis=1)[:, :k]
hits = [rel in top[i] for i, (_, rel) in enumerate(queries)]
clean = np.mean(hits[0::2]); typo = np.mean(hits[1::2])
return np.mean(hits), clean, typo
for name, vec in [("word TF-IDF", TfidfVectorizer()),
("char 3-5gram TF-IDF", TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5)))]:
allq, clean, typo = hit_at(vec)
print(f"{name:<20} hit@1 all={allq:.2f} clean queries={clean:.2f} typo queries={typo:.2f}")
Output:
word TF-IDF hit@1 all=0.56 clean queries=0.88 typo queries=0.25 char 3-5gram TF-IDF hit@1 all=1.00 clean queries=1.00 typo queries=1.00
Freeze chunks while comparing retrievers
Change one variable at a time; re-chunking and switching retriever together makes the result unexplainable.
Quick check: Why compare retrievers on your own evaluation set rather than a public leaderboard?
- Retrievers cannot be compared
- Leaderboards are always fake
- Your set has no labels
- Your queries and documents differ from benchmark data
Answer
Your queries and documents differ from benchmark data — Measure on the distribution you serve.