# Train, Validation and Test Splits Without Leakage — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/x-split

> Keep related rows together so the test score is honest.

## Split by group, not by row

You need data the model never trained on to measure it: a **validation set** to choose settings and stop early, and a **test set** touched once at the end. **Leakage** happens when near-duplicates or the same customer, document or conversation appear on both sides, so the score flatters the model. Split by the unit that must stay unseen (customer, document, thread), de-duplicate first, and keep the test set free of anything used for prompt writing or example selection. If the test set is small, say so and report uncertainty.

## Group split versus random split, run

I ran this with plain Python 3 (standard library only), using example numbers. A group-aware split of 40 rows keeps customers separate (no customer on both sides). A naive random split of the same rows puts 6 of 10 customers on both sides.

```python
import hashlib, random

def split_by_group(examples, key, test_frac=0.2):
    """Keep every example from the same customer/document in ONE split, so near-duplicates cannot leak across."""
    def bucket(k): return int(hashlib.md5(k.encode()).hexdigest(), 16) % 100 / 100
    train, test = [], []
    for ex in examples: (test if bucket(ex[key]) < test_frac else train).append(ex)
    return train, test

rows = [{"customer": f"c{i % 10}", "text": f"ticket {i}"} for i in range(40)]
train, test = split_by_group(rows, "customer")
shared = {r["customer"] for r in train} & {r["customer"] for r in test}
print(f"train {len(train)} rows, test {len(test)} rows; customers appearing in both splits: {sorted(shared)}")
rnd = random.Random(0); rnd.shuffle(rows); naive_train, naive_test = rows[:32], rows[32:]
shared2 = {r["customer"] for r in naive_train} & {r["customer"] for r in naive_test}
print(f"naive random split: customers in both splits: {len(shared2)} of 10  <- leakage risk")

```

Output:

```
train 32 rows, test 8 rows; customers appearing in both splits: []
naive random split: customers in both splits: 6 of 10  <- leakage risk
```

## Freeze the test set

Store the test set separately, look at it only for final reporting, and never tune on it.

**Quiz:** What is data leakage?

- [ ] Using too few epochs
- [ ] Losing the dataset file
- [x] Related or duplicate rows on both sides of the split, inflating the score
- [ ] Encrypting the data

*Answer:* Related or duplicate rows on both sides of the split, inflating the score. Split by the unit that must remain unseen.
