Lesson 22 / 25
Overhead, Volume and Cost
Telemetry is not free.
Control volume at every stage
Instrumentation adds some CPU, memory and network overhead, usually small with batching but measurable on hot paths. Backend costs scale with spans, series and log volume. Control them with sampling, dropping noisy spans (health checks, static assets), limiting attribute cardinality, setting span attribute and event limits in the SDK, and choosing retention per signal. Monitor the Collector itself: it exposes its own metrics about received, refused and dropped data.
Reliable, affordable, safe
Overhead, cost, privacy and a gradual rollout decide whether observability succeeds.
Levers for reducing volume
Where to cut, from cheapest to most targeted.
SDK sampler ratio; disable unneeded instrumentations; attribute count/length limits
Collector filter health checks and static assets; drop unused metrics; tail sampling
backend retention per signal; downsampling of old metrics
design bounded attributes; IDs as span attributes (not metric attributes); fewer, better spansWatch dropped-data metrics
Refused or dropped counters on the Collector reveal back-pressure before data silently disappears.
Quick check: Which is a good way to cut trace volume without losing error visibility?
- Sampling 0% of everything
- Disabling tracing for failing services
- Removing trace IDs from logs
- Tail sampling that keeps all error traces and a fraction of the rest
Answer
Tail sampling that keeps all error traces and a fraction of the rest — Keep the valuable traces.