Lesson 17 / 25
What to Monitor After Launch
Quality, safety, cost and experience.
Four groups of signals
Monitor reliability (error and timeout rates, fallback rate, latency percentiles), quality (thumbs-down rate, regenerate and edit rates, escalations, sampled human ratings), safety (check blocks, safety flags, reported incidents) and cost (tokens and spend per request and per day). Break them down by config version, cohort and slice, and compare the AI group with the control. Dashboards should make it obvious within minutes whether a ramp step made things worse.
Watch, decide, act
Live dashboards, automatic triggers and human review keep a launched feature safe.
A launch dashboard layout
One screen per AI feature.
row 1 reliability error % | timeout % | fallback % | p50 / p95 latency
row 2 quality thumbs-down % | regenerate % | escalation % | sampled rating
row 3 safety input blocks | output blocks | safety flags per 1k | incidents
row 4 cost tokens / request | $ / request | $ today vs cap
filters: config version | cohort | language | AI vs controlSplit by config version
Most regressions follow a prompt or model change; a per-version breakdown finds them immediately.
Quick check: Why break monitoring down by config version?
- Versions do not affect behaviour
- To hide problems
- To link regressions to specific prompt or model changes
- It reduces token usage
Answer
To link regressions to specific prompt or model changes — Change tracking makes diagnosis fast.