# Load Balancing gRPC — gRPC

Source: https://www.skillbyai.com/en/grpc/e-lb

> Per-request, not per-connection.

## Why L4 balancing is not enough

Because HTTP/2 multiplexes many calls over a **long-lived connection**, a layer-4 (TCP) load balancer assigns all of a client's calls to one backend, causing uneven load. Options: an **L7 proxy** that understands HTTP/2 and balances per request (Envoy, NGINX, cloud application load balancers, service meshes), or **client-side load balancing**, where the client resolves all backend addresses (for example via a Kubernetes headless Service and DNS) and spreads calls with a policy such as round robin. Also set maximum connection age on servers so clients rebalance after scaling.

## Running gRPC among other systems

Long-lived connections need special load balancing, browsers need a bridge, and tools make services discoverable.

![Three ideas: load balancing, browsers and gateways, health and tooling.](assets/figures/grpc/section-7-map.svg) — Figure 7.1 — Load balancing, gateways and tooling.

## Client-side round robin with DNS

Go client using a headless Kubernetes Service (a sketch).

```go
conn, err := grpc.NewClient(
    "dns:///orders-headless.shop.svc.cluster.local:50051",   // resolves to every pod IP
    grpc.WithTransportCredentials(creds),
    grpc.WithDefaultServiceConfig(`{"loadBalancingConfig": [{"round_robin":{}}]}`),
)
```

## Watch per-pod request rates

If one pod handles most traffic after a scale-up, connection-level balancing is the likely cause.

**Quiz:** Why can a TCP load balancer distribute gRPC traffic unevenly?

- [ ] TLS prevents balancing
- [ ] gRPC does not use TCP
- [ ] Protobuf messages are too small
- [x] Many calls share one long-lived HTTP/2 connection to a single backend

*Answer:* Many calls share one long-lived HTTP/2 connection to a single backend. Balance per request at L7 or client-side.
