Distributed Tracing
Follow one request across services with spans and context propagation; Micrometer Tracing over OpenTelemetry, exporting to Tempo or Zipkin.
On this page
Metrics say checkout latency's p99 doubled; they can't say where the time went. Once BookVault is several services - a gateway, a catalog service, a loans service - a single user action fans out across all of them, and you need to follow one request through the whole chain. That's distributed tracing: spans, a shared trace ID, and context propagation, wired in Spring with Micrometer Tracing over OpenTelemetry.
Traces and spans
A trace represents one end-to-end request. It's a tree of spans, where each span is one timed unit of work with a name, start/end time, and attributes:
Trace abc123 (GET /api/checkout, total 540 ms)
ββ span: gateway.route [ 8 ms]
ββ span: loans.checkout [520 ms]
β ββ span: db.query loans [340 ms] β the culprit
β ββ span: http catalog-service [175 ms]
β ββ span: db.query books [150 ms]
ββ span: response.serialize [12 ms]Every span shares the trace ID (abc123) and records its parent span ID, so the collector reassembles
the tree. Reading it, the 340 ms database query jumps out - tracing turns "checkout is slow" into "this
query is slow," which is the entire point.
Context propagation: the hard part
The magic that makes it distributed is context propagation: when the loans service calls the catalog
service, it must pass the trace ID along so the downstream spans join the same trace instead of starting a new
one. The standard is the W3C traceparent HTTP header:
traceparent: 00-abc123...-def456...-01
β β β β flags (sampled?)
β β β parent span id
β β trace id (shared across the whole request)
β versionMicrometer Tracing propagates this automatically across Spring's instrumented clients - RestClient,
WebClient, RestTemplate, and messaging - so a call chain stays one trace with no manual header-passing. The
same context also flows into your logs (next lesson), which is what lets you jump from a slow span straight to
its log lines.
Wiring it in Spring
Micrometer Tracing is a facade (like Micrometer is for metrics) with a pluggable tracer - today, OpenTelemetry (OTel), the industry standard. Add the bridge and an exporter:
implementation 'io.micrometer:micrometer-tracing-bridge-otel'
implementation 'io.opentelemetry:opentelemetry-exporter-otlp' // send spans to a collectormanagement:
tracing:
sampling:
probability: 0.1 # sample 10% of traces (see below)
otlp:
tracing:
endpoint: http://otel-collector:4318/v1/tracesSpans are exported (via OTLP) to a backend like Grafana Tempo, Jaeger, or Zipkin, where you search
and visualize them. @Observation/@Observed or the ObservationRegistry let you add custom spans around
business operations the framework doesn't auto-instrument.
Sample - you can't trace everything
Recording and storing a span for every request at high volume is expensive and mostly redundant. Sampling keeps a fraction (e.g. 10%) - enough to see patterns and catch slow requests, at a fraction of the cost. Use tail- based sampling (decide after seeing the trace) if you must keep all errors/slow traces; head-based (a fixed probability up front) is simpler and the common default. Whatever the rate, propagation must still pass the context on every hop so sampled traces stay complete.
A trace ID is a tracking number stamped on a parcel. Each facility it passes - origin depot, air hub, customs, local courier - logs a scan (a span) with a timestamp, all filed under the same tracking number, so you reconstruct the parcel's whole journey and see it sat in customs for two days (the slow span). Context propagation is each carrier writing the tracking number on the label it hands to the next carrier - drop it at any handoff and the parcel 'disappears' into an untracked leg. Sampling is only deep-scanning one parcel in ten to keep the system affordable while still spotting the bottleneck.
BookVault's traces look complete within the loans service, but every call to the catalog service shows up as a brand-new trace with no parent - so you can never see the full checkout path in one view. What's almost certainly wrong, and where would you look to fix it?
What makes tracing 'distributed' - i.e. lets spans from different services join one trace?
Key takeaways
- A trace is one end-to-end request; it's a tree of spans, each a timed unit of work sharing the trace ID and recording its parent.
- Reading a span tree localizes latency/errors to a specific hop - turning 'checkout is slow' into 'this DB query is slow'.
- Context propagation (the W3C traceparent header) carries the trace ID downstream so spans across services join one trace - the essence of 'distributed'.
- Micrometer Tracing over OpenTelemetry auto-propagates context through Spring's instrumented clients and exports spans (OTLP) to Tempo/Jaeger/Zipkin.
- Sample a fraction of traces to control cost, and know that one un-instrumented outbound call severs the trace by dropping the context.