Lesson 1 / 25

How AI Features Differ From Ordinary Features

Quality is a distribution, not a pass/fail.

Five differences

A normal feature either works or throws an error. An AI feature (summaries, chat, classification, drafting) produces varied outputs whose quality spreads across a range; most are fine, some are wrong, a few are harmful. It can change behaviour when the provider updates a model or when you edit a prompt. Its cost scales with usage and output length. Users may over-trust fluent but wrong answers. And misuse (prompt injection, jailbreaks) is part of the threat model. A rollout plan must handle all five.

Probabilistic features, real consequences

AI features fail in fuzzy, varied ways, so launches need explicit metrics, bars and exit routes.

Three ideas: what is different, the plan, metrics.
Figure 1.1 — Differences, plan and metrics.

Ordinary versus AI feature launches

What changes in the plan.

aspect            ordinary feature            AI feature
correctness       pass/fail tests             quality distribution, sampled review
changes           only when you deploy        prompt edits, provider model updates
cost              mostly fixed                per request, grows with usage and length
failure visibility errors, crashes            plausible wrong answers, silent
abuse             standard security           + injection, jailbreaks, data leakage

Plan for the bad tail

Assume a small share of outputs will be wrong or unsafe and design how you detect, limit and recover from them.

Quick check: Why is "it passed QA" not enough for an AI feature?

  • QA cannot run AI code
  • AI features never fail
  • Outputs vary, so quality must be measured as a rate across many cases
  • Tests make models worse
Answer

Outputs vary, so quality must be measured as a rate across many cases — Measure distributions, not single passes.