SkillByAIOpen interactive version →

Lesson 19 / 25

Load Balancing gRPC

Per-request, not per-connection.

Why L4 balancing is not enough

Because HTTP/2 multiplexes many calls over a long-lived connection, a layer-4 (TCP) load balancer assigns all of a client's calls to one backend, causing uneven load. Options: an L7 proxy that understands HTTP/2 and balances per request (Envoy, NGINX, cloud application load balancers, service meshes), or client-side load balancing, where the client resolves all backend addresses (for example via a Kubernetes headless Service and DNS) and spreads calls with a policy such as round robin. Also set maximum connection age on servers so clients rebalance after scaling.

Running gRPC among other systems

Long-lived connections need special load balancing, browsers need a bridge, and tools make services discoverable.

Figure 7.1 — Load balancing, gateways and tooling.

Client-side round robin with DNS

Go client using a headless Kubernetes Service (a sketch).

conn, err := grpc.NewClient(
    "dns:///orders-headless.shop.svc.cluster.local:50051",   // resolves to every pod IP
    grpc.WithTransportCredentials(creds),
    grpc.WithDefaultServiceConfig(`{"loadBalancingConfig": [{"round_robin":{}}]}`),
)

Watch per-pod request rates

If one pod handles most traffic after a scale-up, connection-level balancing is the likely cause.

Quick check: Why can a TCP load balancer distribute gRPC traffic unevenly?

  • TLS prevents balancing
  • gRPC does not use TCP
  • Protobuf messages are too small
  • Many calls share one long-lived HTTP/2 connection to a single backend
Answer

Many calls share one long-lived HTTP/2 connection to a single backend — Balance per request at L7 or client-side.