Lesson 19 / 27
Safety, Privacy and Licensing of Training Data
What goes into training can come back out.
Treat the dataset as a risk surface
Fine-tuning can weaken safety behaviour the base model had, even on harmless-looking data, so re-run your safety and refusal tests afterwards. Anything in the training set can be memorised and regurgitated: remove secrets, credentials and personal data, apply the same privacy rules as for any copy of customer data, and keep a record of where the data came from and whether you may use it (licences, customer terms, provider terms). Poisoned or user-supplied examples can plant bad behaviour, so review data that arrives from outside your team.
A pre-training data checklist
Run it before every dataset leaves your machine.
[ ] secrets / API keys / passwords scanned and removed
[ ] personal data removed or masked, per our privacy policy
[ ] source and licence recorded for every data source
[ ] provider terms allow training on these inputs/outputs
[ ] external or user-submitted rows reviewed for poisoning
[ ] safety / refusal tests ready to re-run after training
[ ] dataset version and hash savedScan, do not trust
Run an automated secret and personal-data scan over every row; manual review alone misses the one API key at row 8,412.
Quick check: Why re-run safety tests after fine-tuning?
- They are optional marketing
- Safety tests train the model
- Tuning always improves safety
- Tuning can weaken the base model's safety behaviour
Answer
Tuning can weaken the base model's safety behaviour — Check safety after every change to the weights.