पाठ 22 / 25

Observability

Metrics, traces and logs per method.

Instrument by method and status

Track request rate, error rate (by status code) and latency for every method, on both client and server. OpenTelemetry provides gRPC instrumentation that creates spans per RPC and propagates trace context in metadata, so calls appear in distributed traces. Log the method, status code, duration and request ID for each call, but never message contents containing personal data. Stream metrics also need message counts and stream durations.

Observable, consistent, tested

Observability, API conventions, testing and a checklist keep gRPC services dependable.

Four ideas: observability, API design conventions, testing, checklist.
Figure 8.1 — Observability, conventions, testing and checklist.

Useful per-method signals

What to chart and alert on.

rate        rpc calls per second by service/method
errors      calls by grpc status code (alert on INTERNAL, UNAVAILABLE, DEADLINE_EXCEEDED spikes)
latency     p50/p95/p99 duration by method
streams     active streams, messages sent/received, stream duration
clients     retries, hedged attempts, connection churn

Separate client and server views

Client-side latency includes network and queueing; comparing it with server-side latency locates delays.

त्वरित जाँच: Which dimension should gRPC metrics be broken down by?

  • Random IDs per request
  • Message field values
  • User passwords
  • Service method and status code
Answer

Service method and status code — Bounded, meaningful labels.