Lesson 7 / 25
Packing Retrieved Chunks Into a Budget
Choose the best chunks that fit, with a relevance floor.
Rank, threshold, fill
Retrieval returns candidates with a relevance score. Packing turns them into context: sort by score, drop anything below a minimum score (a weak match is often worse than nothing because it invites the model to use it), and add chunks while they fit the token budget. Skip a chunk that would overflow and try smaller ones. Also consider freshness and authority: an old policy version with a high similarity score can be more harmful than helpful, so filter by metadata (date, status, source) before or during packing.
Pick, de-duplicate, compress, exemplify
Good context is chosen, not dumped: rank, remove repeats, shrink, and pick examples that match the request.
Greedy packing with a threshold, run
I ran this with plain Python 3 (standard library only); the data is made-up example data. Six candidate chunks, a 1,600-token budget and a 0.6 score floor: the code packs the refund policy, refund FAQ and returns form (1,450 tokens), skips an old 2019 policy that would overflow, and skips two low-relevance chunks.
chunks = [ # (id, relevance score, tokens)
("refund-policy", 0.91, 700), ("refund-faq", 0.88, 450), ("shipping", 0.52, 600),
("refund-2019-old", 0.85, 900), ("contact", 0.40, 150), ("returns-form", 0.77, 300),
]
budget, min_score = 1600, 0.6
picked, used = [], 0
for cid, score, tok in sorted(chunks, key=lambda c: -c[1]):
if score < min_score:
print(f"skip {cid:<16} score {score} below threshold")
continue
if used + tok > budget:
print(f"skip {cid:<16} would exceed budget ({used}+{tok})")
continue
picked.append(cid); used += tok
print("packed:", picked, "| tokens:", used, "of", budget)
Output:
skip refund-2019-old would exceed budget (1150+900) skip shipping score 0.52 below threshold skip contact score 0.4 below threshold packed: ['refund-policy', 'refund-faq', 'returns-form'] | tokens: 1450 of 1600
Filter stale versions explicitly
Mark superseded documents in metadata and exclude them; similarity search cannot tell that a 2019 policy is out of date.
Quick check: Why use a minimum relevance score when packing?
- Weak matches can mislead the model, so leaving them out is often better
- It makes every chunk longer
- It increases the context window
- It is required by JSON
Answer
Weak matches can mislead the model, so leaving them out is often better — Only include content that genuinely helps.