Lesson 19 / 25
Load Balancing gRPC
Per-request, not per-connection.
Why L4 balancing is not enough
Because HTTP/2 multiplexes many calls over a long-lived connection, a layer-4 (TCP) load balancer assigns all of a client's calls to one backend, causing uneven load. Options: an L7 proxy that understands HTTP/2 and balances per request (Envoy, NGINX, cloud application load balancers, service meshes), or client-side load balancing, where the client resolves all backend addresses (for example via a Kubernetes headless Service and DNS) and spreads calls with a policy such as round robin. Also set maximum connection age on servers so clients rebalance after scaling.
Running gRPC among other systems
Long-lived connections need special load balancing, browsers need a bridge, and tools make services discoverable.
Client-side round robin with DNS
Go client using a headless Kubernetes Service (a sketch).
conn, err := grpc.NewClient(
"dns:///orders-headless.shop.svc.cluster.local:50051", // resolves to every pod IP
grpc.WithTransportCredentials(creds),
grpc.WithDefaultServiceConfig(`{"loadBalancingConfig": [{"round_robin":{}}]}`),
)Watch per-pod request rates
If one pod handles most traffic after a scale-up, connection-level balancing is the likely cause.
Quick check: Why can a TCP load balancer distribute gRPC traffic unevenly?
- TLS prevents balancing
- gRPC does not use TCP
- Protobuf messages are too small
- Many calls share one long-lived HTTP/2 connection to a single backend
Answer
Many calls share one long-lived HTTP/2 connection to a single backend — Balance per request at L7 or client-side.