Lesson 21 / 25
Online Signals and A/B Tests
Offline sets miss what real users do.
Implicit and explicit feedback
Production gives signals the offline set cannot: explicit feedback (thumbs up/down, comments), implicit signals (follow-up rephrasing, copying the answer, opening a cited source, escalation to a human), not-found rates and latency. Signals are biased (few users rate, unhappy users rate more), so use them for trends and for finding failures, not as precise quality scores. For big changes, run an A/B test: split traffic, compare pre-registered metrics, and run long enough for a meaningful sample. Feed bad examples back into the offline set.
Signals to log per answer
Combine them into dashboards and review queues.
answer_id, template_version, retriever_version
retrieved ids + scores, cited ids
user feedback: thumbs, comment
implicit: rephrased within 60 s?, cited source opened?, escalated?
not_found: true/false
latency_ms, tokens_in, tokens_outReview a random sample, not only complaints
Complaints show failures users noticed; random samples show failures they did not.
Quick check: Why are thumbs-up/down ratings not a precise quality score?
- Ratings are always positive
- Few users rate and those who do are not representative
- Ratings cannot be stored
- They measure latency
Answer
Few users rate and those who do are not representative — Use them for trends and failure discovery.