# Dataset Format and Validation (JSONL) — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/x-format

> Check every row before you pay for a training run.

## One JSON object per line

Hosted fine-tuning services and open-source trainers usually take **JSONL**: one JSON object per line, for chat models typically a `messages` list of role/content turns ending with the assistant answer. The exact schema differs by provider, so follow the current documentation. Whatever the schema, validate locally first: every line parses, roles are valid, the last turn is a non-empty assistant answer, token lengths fit the limit, and no row contains secrets or personal data. A failed run caught by a 20-line script is far cheaper than one caught by a bill.

## A row validator, run

I ran this with plain Python 3 (standard library only), using example numbers. The checker accepts a valid row and reports a missing assistant turn, an empty final answer and a JSON syntax error in the other three sample lines. The schema shown is a typical chat shape, not any one provider's exact format.

```python
import json

def check(line):
    try: ex = json.loads(line)
    except json.JSONDecodeError as e: return f"bad JSON: {e.msg}"
    m = ex.get("messages")
    if not isinstance(m, list) or len(m) < 2: return "needs a messages list with at least 2 turns"
    if any(t.get("role") not in {"system", "user", "assistant"} for t in m): return "unknown role"
    if m[-1]["role"] != "assistant" or not m[-1].get("content", "").strip(): return "last turn must be a non-empty assistant answer"
    return "ok"

lines = ['{"messages": [{"role": "user", "content": "Where is my order?"}, {"role": "assistant", "content": "Please share your order number."}]}',
         '{"messages": [{"role": "user", "content": "Hi"}]}',
         '{"messages": [{"role": "user", "content": "Hi"}, {"role": "assistant", "content": ""}]}', '{oops}']
seen = set()
for i, l in enumerate(lines, 1):
    dup = l in seen; seen.add(l)
    print(i, check(l), "(duplicate)" if dup else "")

```

Output:

```
1 ok 
2 needs a messages list with at least 2 turns 
3 last turn must be a non-empty assistant answer 
4 bad JSON: Expecting property name enclosed in double quotes
```

## Dry-run with 20 rows

Run a tiny job first to confirm format, cost and the output style before launching the full dataset.

**Quiz:** Why validate the dataset locally first?

- [x] Format errors are caught cheaply before a paid training run
- [ ] Validation makes the model larger
- [ ] JSONL files cannot be edited
- [ ] It replaces the evaluation set

*Answer:* Format errors are caught cheaply before a paid training run. Cheap local checks avoid wasted runs.
