Lesson 15 / 27
Dataset Format and Validation (JSONL)
Check every row before you pay for a training run.
One JSON object per line
Hosted fine-tuning services and open-source trainers usually take JSONL: one JSON object per line, for chat models typically a messages list of role/content turns ending with the assistant answer. The exact schema differs by provider, so follow the current documentation. Whatever the schema, validate locally first: every line parses, roles are valid, the last turn is a non-empty assistant answer, token lengths fit the limit, and no row contains secrets or personal data. A failed run caught by a 20-line script is far cheaper than one caught by a bill.
A row validator, run
I ran this with plain Python 3 (standard library only), using example numbers. The checker accepts a valid row and reports a missing assistant turn, an empty final answer and a JSON syntax error in the other three sample lines. The schema shown is a typical chat shape, not any one provider's exact format.
import json
def check(line):
try: ex = json.loads(line)
except json.JSONDecodeError as e: return f"bad JSON: {e.msg}"
m = ex.get("messages")
if not isinstance(m, list) or len(m) < 2: return "needs a messages list with at least 2 turns"
if any(t.get("role") not in {"system", "user", "assistant"} for t in m): return "unknown role"
if m[-1]["role"] != "assistant" or not m[-1].get("content", "").strip(): return "last turn must be a non-empty assistant answer"
return "ok"
lines = ['{"messages": [{"role": "user", "content": "Where is my order?"}, {"role": "assistant", "content": "Please share your order number."}]}',
'{"messages": [{"role": "user", "content": "Hi"}]}',
'{"messages": [{"role": "user", "content": "Hi"}, {"role": "assistant", "content": ""}]}', '{oops}']
seen = set()
for i, l in enumerate(lines, 1):
dup = l in seen; seen.add(l)
print(i, check(l), "(duplicate)" if dup else "")
Output:
1 ok 2 needs a messages list with at least 2 turns 3 last turn must be a non-empty assistant answer 4 bad JSON: Expecting property name enclosed in double quotes
Dry-run with 20 rows
Run a tiny job first to confirm format, cost and the output style before launching the full dataset.
Quick check: Why validate the dataset locally first?
- Format errors are caught cheaply before a paid training run
- Validation makes the model larger
- JSONL files cannot be edited
- It replaces the evaluation set
Answer
Format errors are caught cheaply before a paid training run — Cheap local checks avoid wasted runs.