# Safety, Privacy and Licensing of Training Data — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/r-safety

> What goes into training can come back out.

## Treat the dataset as a risk surface

Fine-tuning can **weaken safety behaviour** the base model had, even on harmless-looking data, so re-run your safety and refusal tests afterwards. Anything in the training set can be **memorised and regurgitated**: remove secrets, credentials and personal data, apply the same privacy rules as for any copy of customer data, and keep a record of where the data came from and whether you may use it (licences, customer terms, provider terms). Poisoned or user-supplied examples can plant bad behaviour, so review data that arrives from outside your team.

## A pre-training data checklist

Run it before every dataset leaves your machine.

```text
[ ] secrets / API keys / passwords scanned and removed
[ ] personal data removed or masked, per our privacy policy
[ ] source and licence recorded for every data source
[ ] provider terms allow training on these inputs/outputs
[ ] external or user-submitted rows reviewed for poisoning
[ ] safety / refusal tests ready to re-run after training
[ ] dataset version and hash saved
```

## Scan, do not trust

Run an automated secret and personal-data scan over every row; manual review alone misses the one API key at row 8,412.

**Quiz:** Why re-run safety tests after fine-tuning?

- [ ] They are optional marketing
- [ ] Safety tests train the model
- [ ] Tuning always improves safety
- [x] Tuning can weaken the base model's safety behaviour

*Answer:* Tuning can weaken the base model's safety behaviour. Check safety after every change to the weights.
