Lesson 19 / 27

Safety, Privacy and Licensing of Training Data

What goes into training can come back out.

Treat the dataset as a risk surface

Fine-tuning can weaken safety behaviour the base model had, even on harmless-looking data, so re-run your safety and refusal tests afterwards. Anything in the training set can be memorised and regurgitated: remove secrets, credentials and personal data, apply the same privacy rules as for any copy of customer data, and keep a record of where the data came from and whether you may use it (licences, customer terms, provider terms). Poisoned or user-supplied examples can plant bad behaviour, so review data that arrives from outside your team.

A pre-training data checklist

Run it before every dataset leaves your machine.

[ ] secrets / API keys / passwords scanned and removed
[ ] personal data removed or masked, per our privacy policy
[ ] source and licence recorded for every data source
[ ] provider terms allow training on these inputs/outputs
[ ] external or user-submitted rows reviewed for poisoning
[ ] safety / refusal tests ready to re-run after training
[ ] dataset version and hash saved

Scan, do not trust

Run an automated secret and personal-data scan over every row; manual review alone misses the one API key at row 8,412.

Quick check: Why re-run safety tests after fine-tuning?

  • They are optional marketing
  • Safety tests train the model
  • Tuning always improves safety
  • Tuning can weaken the base model's safety behaviour
Answer

Tuning can weaken the base model's safety behaviour — Check safety after every change to the weights.