पाठ 18 / 25
Observability Across Services
Trace requests across services and centralise logs and metrics.
One request, many services
Debugging a monolith often means reading one log file. In microservices, one user request may touch ten services, each with many instances. You need three things working together. Distributed tracing: each request carries a trace ID (the W3C traceparent header), every service records spans, and a tracing backend (Jaeger, Grafana Tempo, Zipkin, or a commercial tool) shows the whole path with timings; OpenTelemetry provides vendor-neutral SDKs and auto-instrumentation. Centralised, structured logs that include the trace ID and service name, so you can jump from a slow span to its logs. Metrics per service using RED (rate, errors, duration) plus saturation, with SLOs and alerts owned by the service's team. Add service catalogues (such as Backstage) listing owners, dashboards, runbooks and dependencies so on-call engineers know who to call.
OpenTelemetry auto-instrumentation for a Python service
Spans for incoming and outgoing HTTP calls are created and linked automatically.
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install # adds instrumentations for detected libraries
export OTEL_SERVICE_NAME=checkout
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector.observability:4317
export OTEL_RESOURCE_ATTRIBUTES=deployment.environment=prod,service.version=1.8.2
opentelemetry-instrument gunicorn app:app --bind 0.0.0.0:8080Put the trace ID in every log line
When a customer reports an error, a trace ID in the error page or response header lets support find every related log line across all services in seconds.
त्वरित जाँच: What lets you see the full path and timing of one request across many services?
- Distributed tracing with propagated trace context
- A bigger log file
- More CPU
- A shared database
Answer
Distributed tracing with propagated trace context — Trace context propagated between services links spans into one end-to-end trace.