पाठ 17 / 25

What to Monitor After Launch

Quality, safety, cost and experience.

Four groups of signals

Monitor reliability (error and timeout rates, fallback rate, latency percentiles), quality (thumbs-down rate, regenerate and edit rates, escalations, sampled human ratings), safety (check blocks, safety flags, reported incidents) and cost (tokens and spend per request and per day). Break them down by config version, cohort and slice, and compare the AI group with the control. Dashboards should make it obvious within minutes whether a ramp step made things worse.

Watch, decide, act

Live dashboards, automatic triggers and human review keep a launched feature safe.

Three ideas: signals, rollback triggers, sampled review.
Figure 6.1 — Signals, triggers and review.

A launch dashboard layout

One screen per AI feature.

row 1  reliability   error %  | timeout %  | fallback %  | p50 / p95 latency
row 2  quality       thumbs-down % | regenerate % | escalation % | sampled rating
row 3  safety        input blocks | output blocks | safety flags per 1k | incidents
row 4  cost          tokens / request | $ / request | $ today vs cap
filters: config version | cohort | language | AI vs control

Split by config version

Most regressions follow a prompt or model change; a per-version breakdown finds them immediately.

त्वरित जाँच: Why break monitoring down by config version?

  • Versions do not affect behaviour
  • To hide problems
  • To link regressions to specific prompt or model changes
  • It reduces token usage
Answer

To link regressions to specific prompt or model changes — Change tracking makes diagnosis fast.