SkillByAIOpen interactive version →

Lesson 5 / 25

Why Longer Context Is Not Always Better

More tokens cost money and time and can bury the answer.

Relevance beats volume

Filling a large window with everything that might help has costs: price scales with input tokens, latency grows (the model must process all of it before answering), and quality can drop because relevant facts are surrounded by distractors. Research and practice have found that models can use information near the start and end of a long input more reliably than information buried in the middle, though newer models vary, so test your own. The practical rule: send the smallest set of highly relevant content that answers the question, and measure whether adding more actually helps.

A briefing pack, not a warehouse

Handing someone a 300-page binder for a one-line question makes them slower and more likely to miss the one page that matters.

Plot quality against k

Run your evaluation with 2, 5, 10 and 20 retrieved chunks; stop at the point where quality stops rising.

Quick check: Why can a much longer context hurt answers?

  • Models refuse long inputs always
  • Relevant facts get buried among distractors, and cost and latency rise
  • Long inputs change the model weights
  • Tokens become free
Answer

Relevant facts get buried among distractors, and cost and latency rise — Send the smallest highly relevant set.