पाठ 13 / 25

Batch, Online and Streaming Serving

When and how predictions are produced.

Pick by freshness and latency needs

Batch scoring runs on a schedule (nightly churn scores written to a table): simple, cheap, easy to validate, but predictions are hours old. Online serving answers requests in real time behind an API (fraud check at payment): fresh, but needs low latency, high availability and online features. Streaming scores events as they flow through a queue (Kafka): near real time without a request/response API. Many teams start with batch and move only the use cases that truly need freshness to online.

Serve, shadow, canary

Choose the serving pattern, then expose new models gradually with evidence at each step.

Four ideas: serving patterns, latency, shadow, canary.
Figure 5.1 — Serving, latency, shadow and canary.

Choosing a serving pattern

A starting guide.

pattern     freshness      latency need     complexity   example
batch       hours          none             low          nightly churn scores in the CRM
online      per request    < 100 ms         high         fraud check during checkout
streaming   seconds        moderate         medium       anomaly alerts on sensor events
edge        per request    very low, offline medium     on-device keyboard suggestions

Prefer batch when you can

If a prediction used tomorrow can be computed tonight, batch scoring avoids a whole class of operational problems.

त्वरित जाँच: Which serving pattern fits nightly churn scores loaded into a CRM?

  • Online API with 10 ms latency
  • Batch scoring
  • On-device inference
  • Reinforcement learning
Answer

Batch scoring — Freshness of a day is enough here.