# Batch, Online and Streaming Serving — MLOps

Source: https://www.skillbyai.com/en/mlops/p-patterns

> When and how predictions are produced.

## Pick by freshness and latency needs

**Batch** scoring runs on a schedule (nightly churn scores written to a table): simple, cheap, easy to validate, but predictions are hours old. **Online** serving answers requests in real time behind an API (fraud check at payment): fresh, but needs low latency, high availability and online features. **Streaming** scores events as they flow through a queue (Kafka): near real time without a request/response API. Many teams start with batch and move only the use cases that truly need freshness to online.

## Serve, shadow, canary

Choose the serving pattern, then expose new models gradually with evidence at each step.

![Four ideas: serving patterns, latency, shadow, canary.](assets/figures/mlops/section-5-map.svg) — Figure 5.1 — Serving, latency, shadow and canary.

## Choosing a serving pattern

A starting guide.

```text
pattern     freshness      latency need     complexity   example
batch       hours          none             low          nightly churn scores in the CRM
online      per request    < 100 ms         high         fraud check during checkout
streaming   seconds        moderate         medium       anomaly alerts on sensor events
edge        per request    very low, offline medium     on-device keyboard suggestions
```

## Prefer batch when you can

If a prediction used tomorrow can be computed tonight, batch scoring avoids a whole class of operational problems.

**Quiz:** Which serving pattern fits nightly churn scores loaded into a CRM?

- [ ] Online API with 10 ms latency
- [x] Batch scoring
- [ ] On-device inference
- [ ] Reinforcement learning

*Answer:* Batch scoring. Freshness of a day is enough here.
