# Removing Near-Duplicates — Prompt & Context Engineering

Source: https://www.skillbyai.com/en/prompt-context-engineering/s-dedupe

> Repeated content wastes budget and over-weights one source.

## Same fact, many copies

Knowledge bases are full of near-duplicates: FAQ copies, versioned pages, email quotes. Sending several copies wastes tokens, pushes out other useful content, and can make the model treat a point as more important than it is. Before packing, remove near-duplicates using a cheap similarity measure: **shingle overlap (Jaccard)** on word sequences works for copies; **embedding similarity** catches paraphrases. Keep the most authoritative or newest copy, not simply the first one returned.

## Jaccard de-duplication on word shingles, run

I ran this with plain Python 3 (standard library only); the data is made-up example data. Two refund sentences share most 3-word shingles (similarity 0.62), so the second is dropped; the unrelated shipping sentence is kept. The 0.6 threshold is a tunable choice.

```python
def shingles(t, k=3):
    w = t.lower().split()
    return {" ".join(w[i:i + k]) for i in range(len(w) - k + 1)}
def jaccard(a, b):
    return len(a & b) / len(a | b)
docs = {
    "d1": "Refunds are issued within 14 days of receiving the returned item.",
    "d2": "Refunds are issued within 14 days of receiving the returned item in good condition.",
    "d3": "Shipping to Pune takes three to five working days.",
}
kept = []
for name, text in docs.items():
    sims = [round(jaccard(shingles(text), shingles(docs[k])), 2) for k in kept]
    if any(s >= 0.6 for s in sims):
        print(name, "dropped as near-duplicate, similarity", max(sims))
    else:
        kept.append(name)
print("kept:", kept)
```

Output:

```
d2 dropped as near-duplicate, similarity 0.62
kept: ['d1', 'd3']
```

## Choose which copy survives

When two chunks are duplicates, keep the one with the newer date or higher-authority source, and log what was dropped.

**Quiz:** What is a cost of sending duplicate chunks?

- [ ] They reduce latency
- [ ] They improve citations automatically
- [x] They waste budget and can over-weight one source
- [ ] They shrink the model

*Answer:* They waste budget and can over-weight one source. De-duplicate before packing.
