Lesson 16 / 27
Train, Validation and Test Splits Without Leakage
Keep related rows together so the test score is honest.
Split by group, not by row
You need data the model never trained on to measure it: a validation set to choose settings and stop early, and a test set touched once at the end. Leakage happens when near-duplicates or the same customer, document or conversation appear on both sides, so the score flatters the model. Split by the unit that must stay unseen (customer, document, thread), de-duplicate first, and keep the test set free of anything used for prompt writing or example selection. If the test set is small, say so and report uncertainty.
Group split versus random split, run
I ran this with plain Python 3 (standard library only), using example numbers. A group-aware split of 40 rows keeps customers separate (no customer on both sides). A naive random split of the same rows puts 6 of 10 customers on both sides.
import hashlib, random
def split_by_group(examples, key, test_frac=0.2):
"""Keep every example from the same customer/document in ONE split, so near-duplicates cannot leak across."""
def bucket(k): return int(hashlib.md5(k.encode()).hexdigest(), 16) % 100 / 100
train, test = [], []
for ex in examples: (test if bucket(ex[key]) < test_frac else train).append(ex)
return train, test
rows = [{"customer": f"c{i % 10}", "text": f"ticket {i}"} for i in range(40)]
train, test = split_by_group(rows, "customer")
shared = {r["customer"] for r in train} & {r["customer"] for r in test}
print(f"train {len(train)} rows, test {len(test)} rows; customers appearing in both splits: {sorted(shared)}")
rnd = random.Random(0); rnd.shuffle(rows); naive_train, naive_test = rows[:32], rows[32:]
shared2 = {r["customer"] for r in naive_train} & {r["customer"] for r in naive_test}
print(f"naive random split: customers in both splits: {len(shared2)} of 10 <- leakage risk")
Output:
train 32 rows, test 8 rows; customers appearing in both splits: [] naive random split: customers in both splits: 6 of 10 <- leakage risk
Freeze the test set
Store the test set separately, look at it only for final reporting, and never tune on it.
Quick check: What is data leakage?
- Using too few epochs
- Losing the dataset file
- Related or duplicate rows on both sides of the split, inflating the score
- Encrypting the data
Answer
Related or duplicate rows on both sides of the split, inflating the score — Split by the unit that must remain unseen.