Lesson 23 / 26

Deploying Vision Models

Cloud, edge and the latency budget.

Where inference runs

Vision models run in the cloud (large models, easy updates, needs network), on edge devices (cameras, phones, gateways: low latency, privacy, offline) or both. Edge deployment often uses smaller backbones, quantisation (8-bit weights), pruning and runtimes such as ONNX Runtime, TensorRT, Core ML or TFLite. Measure end-to-end latency including image decoding and preprocessing, batch frames when possible, and monitor accuracy per camera and site after launch.

Deployment trade-offs

Choose by latency, privacy, connectivity and cost.

target    latency        model size     update ease   privacy
cloud GPU low-medium     any            easy          images leave device
edge box  low            small-medium   moderate      stays on site
phone     very low       small          app updates   stays on device
MCU       very low       tiny           hard          stays on device

common tools: ONNX export, int8 quantisation, TensorRT / Core ML / TFLite runtimes

Measure on the target hardware

Benchmarks on a laptop GPU say little about a camera's chip; profile on the real device early.

Quick check: Why deploy a vision model on an edge device?

  • Lower latency, offline operation and keeping images on site
  • To make models larger
  • Because cloud GPUs cannot run CNNs
  • To avoid any testing
Answer

Lower latency, offline operation and keeping images on site — Edge trades capacity for latency and privacy.