पाठ 13 / 25
Batch, Online and Streaming Serving
When and how predictions are produced.
Pick by freshness and latency needs
Batch scoring runs on a schedule (nightly churn scores written to a table): simple, cheap, easy to validate, but predictions are hours old. Online serving answers requests in real time behind an API (fraud check at payment): fresh, but needs low latency, high availability and online features. Streaming scores events as they flow through a queue (Kafka): near real time without a request/response API. Many teams start with batch and move only the use cases that truly need freshness to online.
Serve, shadow, canary
Choose the serving pattern, then expose new models gradually with evidence at each step.
Choosing a serving pattern
A starting guide.
pattern freshness latency need complexity example
batch hours none low nightly churn scores in the CRM
online per request < 100 ms high fraud check during checkout
streaming seconds moderate medium anomaly alerts on sensor events
edge per request very low, offline medium on-device keyboard suggestionsPrefer batch when you can
If a prediction used tomorrow can be computed tonight, batch scoring avoids a whole class of operational problems.
त्वरित जाँच: Which serving pattern fits nightly churn scores loaded into a CRM?
- Online API with 10 ms latency
- Batch scoring
- On-device inference
- Reinforcement learning
Answer
Batch scoring — Freshness of a day is enough here.