Lesson 24 / 25
How LLMOps Differs
Same principles, different artefacts.
Prompts, providers, evaluation and cost
Systems built on large language models follow the same MLOps principles with different emphasis. The versioned artefacts are often prompts, retrieval settings and model choices rather than trained weights; the provider may update models underneath you. Evaluation relies on test sets of questions, rubric-based or model-based grading, and human review. Monitoring adds token cost, latency, refusal and hallucination signals, safety filters and user feedback. Release practices (gates, canaries, rollback) apply unchanged.
Classic MLOps versus LLMOps
Where the emphasis moves.
classic ML LLM applications
main artefact trained model weights prompts, retrieval config, model id
training you train regularly mostly provider-trained; fine-tune sometimes
evaluation labelled metrics eval sets + rubric/LLM graders + humans
monitoring drift, accuracy cost, latency, quality samples, safety, feedback
change risk new data / retraining prompt edits, provider model updatesPin model versions
Use dated model identifiers where providers offer them, and re-run evaluations before switching.
Quick check: What is often the main versioned artefact in LLM applications?
- The training dataset of the provider
- Only GPU drivers
- Prompts, retrieval settings and the chosen model id
- The office floor plan
Answer
Prompts, retrieval settings and the chosen model id — Version what you control.