Lesson 25 / 25

An OpenTelemetry Checklist

Review before relying on it.

Questions to ask

Does every service set service.name, service.version and environment? Do traces cross every service boundary, including queues? Are span names low-cardinality and attributes following semantic conventions? Are logs correlated with trace IDs? Is a Collector in place with memory_limiter and batch processors? Is the sampling strategy deliberate and documented? Are sensitive values excluded and redacted? Are Collector health and dropped data monitored?

The checklist

Use it in reviews.

[ ] service.name, service.version, deployment.environment.name on every service
[ ] W3C trace context propagated over HTTP, gRPC and message queues
[ ] zero-code instrumentation + manual spans on key business paths
[ ] span names low-cardinality; semantic convention attribute names
[ ] logs carry trace_id/span_id; metrics have bounded attributes
[ ] Collector pipelines: memory_limiter first, batch last
[ ] sampling strategy documented (parent-based head and/or tail)
[ ] PII and secrets excluded; redaction in the Collector; TLS + auth
[ ] Collector self-metrics monitored (refused / dropped data)
[ ] backend retention and cost reviewed

Revisit the checklist after incidents

Each incident reveals a missing span, attribute or log; feed it back into instrumentation.

Quick check: Which belongs on an OpenTelemetry checklist?

  • Disable context propagation for speed
  • Put user emails in span names
  • Every service sets service.name and service.version
  • Export without any batching
Answer

Every service sets service.name and service.version — Consistent identity enables everything else.