Lesson 7 / 25

Packing Retrieved Chunks Into a Budget

Choose the best chunks that fit, with a relevance floor.

Rank, threshold, fill

Retrieval returns candidates with a relevance score. Packing turns them into context: sort by score, drop anything below a minimum score (a weak match is often worse than nothing because it invites the model to use it), and add chunks while they fit the token budget. Skip a chunk that would overflow and try smaller ones. Also consider freshness and authority: an old policy version with a high similarity score can be more harmful than helpful, so filter by metadata (date, status, source) before or during packing.

Pick, de-duplicate, compress, exemplify

Good context is chosen, not dumped: rank, remove repeats, shrink, and pick examples that match the request.

Four ideas: pack, dedupe, compress, select examples.
Figure 3.1 — Pack, dedupe, compress and select examples.

Greedy packing with a threshold, run

I ran this with plain Python 3 (standard library only); the data is made-up example data. Six candidate chunks, a 1,600-token budget and a 0.6 score floor: the code packs the refund policy, refund FAQ and returns form (1,450 tokens), skips an old 2019 policy that would overflow, and skips two low-relevance chunks.

chunks = [  # (id, relevance score, tokens)
    ("refund-policy", 0.91, 700), ("refund-faq", 0.88, 450), ("shipping", 0.52, 600),
    ("refund-2019-old", 0.85, 900), ("contact", 0.40, 150), ("returns-form", 0.77, 300),
]
budget, min_score = 1600, 0.6
picked, used = [], 0
for cid, score, tok in sorted(chunks, key=lambda c: -c[1]):
    if score < min_score:
        print(f"skip {cid:<16} score {score} below threshold")
        continue
    if used + tok > budget:
        print(f"skip {cid:<16} would exceed budget ({used}+{tok})")
        continue
    picked.append(cid); used += tok
print("packed:", picked, "| tokens:", used, "of", budget)

Output:

skip refund-2019-old  would exceed budget (1150+900)
skip shipping         score 0.52 below threshold
skip contact          score 0.4 below threshold
packed: ['refund-policy', 'refund-faq', 'returns-form'] | tokens: 1450 of 1600

Filter stale versions explicitly

Mark superseded documents in metadata and exclude them; similarity search cannot tell that a 2019 policy is out of date.

Quick check: Why use a minimum relevance score when packing?

  • Weak matches can mislead the model, so leaving them out is often better
  • It makes every chunk longer
  • It increases the context window
  • It is required by JSON
Answer

Weak matches can mislead the model, so leaving them out is often better — Only include content that genuinely helps.