LLM Engineering Foundations
Turn LLM demos into dependable products: evaluation and statistics, prompt and model versioning, structured output, cost and latency engineering, reliability and observability, serving, governance and team practice, with every experiment run.
Syllabus
From Demo to Product
- What LLM Engineering Is
- A Reference Architecture for an LLM Feature
- Defining Success: Requirements and Metrics
- From Prototype to Production: A Staged Path
Evaluation That You Can Trust
- An Evaluation Harness
- Noise and Confidence Intervals
- Comparing Two Versions Fairly
- LLM-as-Judge: Useful, Biased, Must Be Calibrated
- Test-Set Hygiene: Leakage and Contamination
Engineering for Quality
- Prompts, Models and Settings as Versioned Artefacts
- Release Gates and CI for LLM Changes
- Structured Output, Validation and Repair Loops
- Prompting, RAG or Fine-Tuning: A Decision Process
Latency, Cost and Capacity
- Caching: Exact, Normalised and Semantic
- Model Routing and Cascades
- Tail Latency, Timeouts and Hedged Requests
- Capacity Planning With Little's Law
- Unit Economics: Cost per Resolved Task
Reliability and Observability
Serving, Deployment and Release
Safety, Governance and Team Practice
- Safety Essentials Every LLM Feature Needs
- Privacy, Compliance and Data Handling
- Documentation, Roles and Team Practice