पाठ 17 / 25

Tail Sampling

Decide after the trace completes.

Keep errors and slow traces

Tail sampling buffers spans until a trace is (likely) complete, then decides using its content: keep all traces with errors, all traces slower than a threshold, and a small percentage of the rest. In OpenTelemetry it is implemented by the Collector's tail_sampling processor (contrib). It needs memory to buffer spans for a decision window and all spans of a trace must reach the same Collector instance, so it runs in a gateway tier.

A tail sampling policy

Errors and slow traces kept, 5% of the rest.

processors:
  tail_sampling:
    decision_wait: 10s
    num_traces: 50000
    policies:
      - name: keep-errors
        type: status_code
        status_code:
          status_codes: [ERROR]
      - name: keep-slow
        type: latency
        latency:
          threshold_ms: 1000
      - name: sample-rest
        type: probabilistic
        probabilistic:
          sampling_percentage: 5

Combine head and tail sampling carefully

Heavy head sampling before a tail sampler removes the errors the tail sampler would have kept.

त्वरित जाँच: Why must tail sampling run where all spans of a trace arrive together?

  • Spans are encrypted otherwise
  • It only works on metrics
  • The decision needs the complete trace
  • Because exporters require it
Answer

The decision needs the complete trace — Route by trace ID to one instance.