पाठ 12 / 25
Chunk Size Experiments
Too small splits answers from their context; too large wastes the budget.
The chunking trade-off
Chunk size controls what retrieval can match and what reaches the model. Tiny chunks can separate the words the user searches for (a heading, a product name) from the sentence containing the answer, so the matched chunk lacks the answer. Huge chunks usually contain the answer but send many irrelevant words, raising cost and distraction, and one chunk may mix several topics. Evaluate chunk sizes with the same queries, measuring both whether the answer reaches the model and how many tokens are sent; also test overlap and structure-aware chunking (by heading or section).
Answer coverage and words sent per chunk size, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. In a synthetic corpus where each policy heading is followed by a separate answer sentence, one-sentence chunks never return the answer (0.00) because the heading matches but the answer sits in another chunk. Two-sentence chunks reach 1.00 with 27 words sent; larger chunks also reach 1.00 but send up to 727 words. The aligned layout makes this cleaner than real documents.
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
rng = np.random.default_rng(0)
products = ["laptop", "phone", "tablet", "camera", "speaker", "monitor", "printer", "router"]
topics = ["refund", "warranty", "shipping", "repair", "exchange"]
filler = ["Our team reviews every request carefully.", "Customers can contact support at any time.",
"Policies may change and are published online.", "We aim to respond within one working day.",
"Please keep your order number ready.", "Thank you for choosing our service."]
sentences, items = [], []
for p in products:
for t in topics:
start = len(sentences)
sentences.append(f"Policy section: {p} {t}.") # has the query words
sentences.append(f"The limit is {rng.integers(10, 90)} days.") # has the answer, no query words
sentences += list(rng.choice(filler, size=6))
items.append((f"{p} {t} limit", start + 1))
def evaluate(size, k=3):
chunks = [" ".join(sentences[i:i + size]) for i in range(0, len(sentences), size)]
vec = TfidfVectorizer().fit(chunks)
C = vec.transform(chunks); Q = vec.transform([q for q, _ in items])
top = np.argsort(-(Q @ C.T).toarray(), axis=1)[:, :k]
found = np.mean([ans // size in top[i] for i, (_, ans) in enumerate(items)])
words = np.mean([sum(len(chunks[c].split()) for c in row) for row in top])
return found, words
print("chunk size | answer in top-3 | words sent to the model")
for size in [1, 2, 4, 8, 16, 40]:
f, w = evaluate(size)
print(f"{size:>10} | {f:>15.2f} | {w:>8.0f}")
Output:
chunk size | answer in top-3 | words sent to the model
1 | 0.00 | 12
2 | 1.00 | 27
4 | 1.00 | 66
8 | 1.00 | 146
16 | 1.00 | 292
40 | 1.00 | 727Keep headings with their content
Prepend the section heading to each chunk so small chunks still carry the words people search for.
त्वरित जाँच: What can go wrong with very small chunks?
- They always send too many tokens
- The matched text can be separated from the sentence holding the answer
- They make indexing impossible
- They remove the need for embeddings
Answer
The matched text can be separated from the sentence holding the answer — Retrieval matches one chunk; the answer must be inside it.