Lesson 20 / 25

Debugging Latency and Errors With Traces

A practical method.

From symptom to cause

Start from a symptom (an alert on latency or errors), then search traces for the affected service and time window, filtering for errors or long durations. Open a slow trace and look for the critical path: the longest chain of sequential spans. Common findings are a slow downstream dependency, many sequential calls that could run in parallel (N+1 patterns), retries hidden inside a span, or time spent waiting in queues. Compare with a fast trace for the same endpoint, then check the related logs via the trace ID.

Patterns to look for in a waterfall

What common problems look like.

slow dependency      one long CLIENT span; child SERVER span equally long -> the callee is slow
network / queueing   CLIENT span much longer than the matching SERVER span -> time lost between them
N+1 calls            dozens of short identical DB spans in sequence -> batch or join
sequential fan-out   calls to independent services one after another -> parallelise
retries              repeated CLIENT spans to the same endpoint with errors -> check timeouts / backoff
missing spans        gaps with no children -> uninstrumented code or lost context

Compare slow and fast traces

The difference between a typical and a slow trace of the same endpoint usually points straight at the cause.

Quick check: What does a CLIENT span that is much longer than its matching SERVER span suggest?

  • The server is slow
  • Time is lost between services, such as network, queueing or connection waits
  • The trace is complete and healthy
  • Sampling is disabled
Answer

Time is lost between services, such as network, queueing or connection waits — Compare both sides of a call.