Lesson 14 / 25
Serving Performance: Latency and Batching
Measure tail latency, batch where possible.
p95 and p99, not averages
For online serving, measure latency percentiles (p50, p95, p99) under realistic load, not just the average; users feel the slow tail. Per-call overhead often dominates for small models, so batching requests (inside the server or by the client) can raise throughput dramatically. Other levers: smaller or distilled models, fewer features requiring slow lookups, caching repeated predictions, and right-sized hardware. Set latency budgets per endpoint and test them in CI.
Single-row calls versus one batch, run
I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. Predicting 200 rows one at a time with a 200-tree random forest takes more than five times as long as predicting the same 200 rows in one batch call. Exact timings depend on the machine, so only the comparison is shown.
import time, numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
X, y = load_breast_cancer(return_X_y=True)
model = RandomForestClassifier(n_estimators=200, random_state=0, n_jobs=1).fit(X, y)
single = []
for i in range(200):
t = time.perf_counter(); model.predict(X[i:i + 1]); single.append((time.perf_counter() - t) * 1000)
t = time.perf_counter(); model.predict(X[:200]); batch = (time.perf_counter() - t) * 1000
print(f"200 single calls take {'longer' if sum(single) > batch else 'less'} than one batch of 200: "
f"ratio {'>5x' if sum(single) > 5 * batch else '<=5x'}")
Output:
200 single calls take longer than one batch of 200: ratio >5x
Load test before launch
Replay realistic traffic at expected peak plus margin and check p99 latency and error rates before going live.
Quick check: Why report p95 or p99 latency instead of the average?
- p99 measures accuracy
- Averages are always higher
- Percentiles are faster to compute
- Users experience the slow tail, which averages hide
Answer
Users experience the slow tail, which averages hide — Tail latency defines user experience.