AI Safety, Evaluation and Cost Control
Build LLM applications that are safe, measurable and affordable: safeguards, evaluation metrics, monitoring and cost control, with runnable Python examples.
Syllabus
Safety Foundations
Input and Output Safeguards
- Moderation and Thresholds
- Handling Personal Data
- Grounding and Hallucination Control
- Refusal Calibration
Misuse and Abuse
Fairness, Privacy and Accountability
- Checking for Bias with Slices
- Privacy and Data Governance
- Documentation, Transparency and Human Oversight
Evaluation
- Building an Evaluation Set
- Metrics: Accuracy, Precision, Recall, F1
- How Sure Is Your Score?
- LLM Judges and Human Agreement
- A/B Tests and Regression Checks
Monitoring in Production
Cost Control
- Token Economics
- Model Routing and Cascades
- Caching, Batching and Trimming
- Budgets, Quotas and FinOps for AI