Lesson 24 / 25

How LLMOps Differs

Same principles, different artefacts.

Prompts, providers, evaluation and cost

Systems built on large language models follow the same MLOps principles with different emphasis. The versioned artefacts are often prompts, retrieval settings and model choices rather than trained weights; the provider may update models underneath you. Evaluation relies on test sets of questions, rubric-based or model-based grading, and human review. Monitoring adds token cost, latency, refusal and hallucination signals, safety filters and user feedback. Release practices (gates, canaries, rollback) apply unchanged.

Classic MLOps versus LLMOps

Where the emphasis moves.

                 classic ML                         LLM applications
main artefact    trained model weights              prompts, retrieval config, model id
training         you train regularly                mostly provider-trained; fine-tune sometimes
evaluation       labelled metrics                   eval sets + rubric/LLM graders + humans
monitoring       drift, accuracy                    cost, latency, quality samples, safety, feedback
change risk      new data / retraining              prompt edits, provider model updates

Pin model versions

Use dated model identifiers where providers offer them, and re-run evaluations before switching.

Quick check: What is often the main versioned artefact in LLM applications?

  • The training dataset of the provider
  • Only GPU drivers
  • Prompts, retrieval settings and the chosen model id
  • The office floor plan
Answer

Prompts, retrieval settings and the chosen model id — Version what you control.