Lesson 13 / 25

Design a Ride-Hailing Service

Location updates, matching and geo-indexing.

Requirements and the hard parts

Functional: riders request a ride, nearby drivers are matched, both see live location, trips are priced and paid. Non-functional: low matching latency, high write rate of driver locations, availability. The hard parts are (1) ingesting location updates: assume drivers send a position every few seconds, which at large driver counts is a heavy write stream best kept in memory rather than a relational table; (2) a geo-index to find drivers near a point, using a grid such as geohash, S2 or H3 cells, where each cell maps to the set of available drivers; and (3) matching: query the rider's cell and neighbours, rank candidates by estimated arrival time, and offer the trip to one driver at a time with a timeout, using a lock or conditional update so a driver is never assigned two trips.

Real-time, transactional and media systems

Three systems that stress different things: location updates, consistency of money and stock, and large media delivery.

Three problems: ride-hailing, e-commerce checkout and inventory, and video streaming.
Figure 5.1 — Ride-hailing, checkout and streaming architectures.

Components and data flow

All numbers are illustrative assumptions.

Assumptions: 1M online drivers, 1 update / 4 s  ->  ~250k location writes/s

Driver app --(WebSocket/HTTP, every ~4s)--> Location Service
    Location Service: update in-memory geo-index  cell_id -> {driver_id}
                      (sharded by region/cell), publish to stream for analytics

Rider app --POST /rides--> Ride Service --> Matching Service
    Matching: cells(rider, radius) -> candidate drivers
              -> rank by ETA (routing service) -> offer to best driver (timeout ~10s)
              -> on accept: conditional update driver.state AVAILABLE -> ON_TRIP

Trip DB (strongly consistent): trips, state transitions, fares
Push/notification service: offers and trip updates to apps

Shard by geography

Matching is local, so partitioning the geo-index by region or cell keeps most queries on one shard; handle cell boundaries by also querying neighbouring cells.

Quick check: Why keep live driver locations in an in-memory geo-index?

  • Locations must be kept forever
  • Updates are frequent and short-lived, and nearby lookups must be fast
  • Relational joins are needed for every update
  • Drivers rarely move
Answer

Updates are frequent and short-lived, and nearby lookups must be fast — High-rate, ephemeral data with spatial queries.