पाठ 7 / 25

Creating a Knowledge Base and Chunking

How documents are split decides what can be found.

Chunk size, overlap, cleaning

A knowledge base stores your documents (PDFs, web pages, text) split into chunks that are embedded and indexed. Settings include the chunking mode (general or parent-child, where small chunks are matched but larger parent sections are returned), maximum chunk length, overlap between chunks, delimiters and text cleaning rules. Smaller chunks match precisely but may lose context; larger ones carry context but dilute relevance. Overlap reduces facts being cut in half, at the cost of storing more text. Preview chunks before indexing. Dify's node names, menus and options change between versions; check the current Dify documentation.

Upload, chunk, retrieve

Knowledge bases let workflows answer from your documents; chunking and retrieval settings decide quality.

Three ideas: chunking, retrieval settings, RAG answers.
Figure 3.1 — Chunking, retrieval and answers.

Chunk counts for three settings, run

I ran this with plain Python 3 (scikit-learn 1.9.1 where imported). It models or tests one piece of a Dify app locally; Dify itself was not running. A 360-word document becomes 8 chunks at 50 words with no overlap, 9 chunks (440 stored words) with a 10-word overlap, and 4 chunks at 120 words with a 20-word overlap. Overlap trades storage for continuity. Dify counts characters or tokens rather than words, but the trade-off is the same.

doc = " ".join(f"Sentence {i} about the refund and return policy details." for i in range(1, 41))
def chunk(text, size, overlap):
    words = text.split(); step = size - overlap; out = []
    for start in range(0, len(words), step):
        out.append(" ".join(words[start:start + size]))
        if start + size >= len(words): break
    return out
print("document words:", len(doc.split()))
for size, overlap in [(50, 0), (50, 10), (120, 20)]:
    c = chunk(doc, size, overlap)
    print(f"chunk size {size:>3} words, overlap {overlap:>2}: {len(c):>2} chunks, "
          f"stored words {sum(len(x.split()) for x in c)}")

Output:

document words: 360
chunk size  50 words, overlap  0:  8 chunks, stored words 360
chunk size  50 words, overlap 10:  9 chunks, stored words 440
chunk size 120 words, overlap 20:  4 chunks, stored words 420

Clean documents first

Remove headers, footers, navigation menus and duplicated boilerplate before upload; they pollute retrieval.

त्वरित जाँच: What does chunk overlap help with?

  • Reducing the chance that a fact is split across two chunks
  • Making the index smaller
  • Removing the need for embeddings
  • Speeding up the LLM
Answer

Reducing the chance that a fact is split across two chunks — Overlap keeps boundary context.