SkillByAIOpen interactive version →

Lesson 21 / 25

Online Signals and A/B Tests

Offline sets miss what real users do.

Implicit and explicit feedback

Production gives signals the offline set cannot: explicit feedback (thumbs up/down, comments), implicit signals (follow-up rephrasing, copying the answer, opening a cited source, escalation to a human), not-found rates and latency. Signals are biased (few users rate, unhappy users rate more), so use them for trends and for finding failures, not as precise quality scores. For big changes, run an A/B test: split traffic, compare pre-registered metrics, and run long enough for a meaningful sample. Feed bad examples back into the offline set.

Signals to log per answer

Combine them into dashboards and review queues.

answer_id, template_version, retriever_version
retrieved ids + scores, cited ids
user feedback: thumbs, comment
implicit: rephrased within 60 s?, cited source opened?, escalated?
not_found: true/false
latency_ms, tokens_in, tokens_out

Review a random sample, not only complaints

Complaints show failures users noticed; random samples show failures they did not.

Quick check: Why are thumbs-up/down ratings not a precise quality score?

  • Ratings are always positive
  • Few users rate and those who do are not representative
  • Ratings cannot be stored
  • They measure latency
Answer

Few users rate and those who do are not representative — Use them for trends and failure discovery.