SkillByAIOpen interactive version →

Lesson 22 / 25

Overhead, Volume and Cost

Telemetry is not free.

Control volume at every stage

Instrumentation adds some CPU, memory and network overhead, usually small with batching but measurable on hot paths. Backend costs scale with spans, series and log volume. Control them with sampling, dropping noisy spans (health checks, static assets), limiting attribute cardinality, setting span attribute and event limits in the SDK, and choosing retention per signal. Monitor the Collector itself: it exposes its own metrics about received, refused and dropped data.

Reliable, affordable, safe

Overhead, cost, privacy and a gradual rollout decide whether observability succeeds.

Figure 8.1 — Cost, privacy, rollout and checklist.

Levers for reducing volume

Where to cut, from cheapest to most targeted.

SDK         sampler ratio; disable unneeded instrumentations; attribute count/length limits
Collector   filter health checks and static assets; drop unused metrics; tail sampling
backend     retention per signal; downsampling of old metrics
design      bounded attributes; IDs as span attributes (not metric attributes); fewer, better spans

Watch dropped-data metrics

Refused or dropped counters on the Collector reveal back-pressure before data silently disappears.

Quick check: Which is a good way to cut trace volume without losing error visibility?

  • Sampling 0% of everything
  • Disabling tracing for failing services
  • Removing trace IDs from logs
  • Tail sampling that keeps all error traces and a fraction of the rest
Answer

Tail sampling that keeps all error traces and a fraction of the rest — Keep the valuable traces.