Start Learning
Javaneer
Back to stage
Module 13·Advanced Observability

Structured Logs & Correlation

JSON logs a machine can query, and the trace/span IDs in the MDC that stitch a log line to its trace - so metrics, traces, and logs point at the same request.

13 min readIntermediate
On this page

Logs are the pillar with the most detail and, done badly, the least usable. A human-readable line like Loan 8842 rejected for member over limit is fine for one developer tailing a file - and useless when you have fifty pods emitting millions of lines an hour. Two upgrades make logs a first-class observability signal: structured logging (so machines can query them) and correlation (so a log line links to its trace).

Structured logging

Traditional logs are unstructured strings. Structured logs are events with named fields - typically emitted as JSON - so a log platform can filter and aggregate them like a database:

# unstructured - only greppable
2026-07-22 14:03:11 WARN  Loan 8842 rejected for member 42: over limit

# structured (JSON) - queryable
{"timestamp":"2026-07-22T14:03:11Z","level":"WARN","message":"loan rejected",
 "loanId":8842,"memberId":42,"reason":"over_limit","service":"loans"}

Now you can ask your log platform "all loan rejected events where reason=over_limit in the last hour, grouped by memberId" - impossible with grep over prose. Spring Boot 3.4+ can emit structured JSON logs with a single property, or you add an encoder like Logstash's:

logging:
  structured:
    format:
      console: ecs        # Elastic Common Schema JSON to stdout

Log to stdout, let the platform ship it

In containers, don't manage log files inside the app - write structured lines to stdout and let the platform (Loki, Elasticsearch, CloudWatch) collect and index them. This is the twelve-factor approach: the app treats logs as an event stream, and rotation/shipping/retention are the platform's job, not code's.

Correlation: linking logs to traces

Here's the upgrade that ties the three pillars together. When Micrometer Tracing is active, it puts the current trace ID and span ID into the logging MDC (Mapped Diagnostic Context) - a per-thread bag of key-value pairs the log encoder automatically includes on every line:

{"message":"loan rejected","loanId":8842,"traceId":"abc123","spanId":"def456", ...}

Every log line emitted while handling a request now carries that request's traceId. The payoff is huge:

  • From a slow span in Tempo, copy its trace ID → search logs for traceId=abc123 → read every log line from that exact request across all services, in order.
  • From an error log, copy its trace ID → open the full trace to see where in the request it happened.

That bidirectional jump - trace ↔ logs by shared ID - is what turns three separate tools into one investigation. Metrics point at the trace, the trace's ID pivots to the logs, the logs explain it.

Getting the context right

Two practical rules keep correlation working:

  • Propagate the MDC across threads. The trace context lives on the request thread. If you hand work to an @Async method, an executor, or a reactive publishOn, the MDC can be lost unless you use a context- propagating executor (Micrometer's ContextSnapshot/ContextPropagatingTaskDecorator). A log line with no traceId is an orphan you can't correlate.
  • Put identifiers in fields, not the message. Write member rejected with a memberId field, not "member 42 rejected" baked into the string. Fields are queryable and keep the message low-cardinality (good for grouping); IDs embedded in the message text are neither.
Case files stamped with the same case number

Unstructured logs are a pile of loose handwritten notes - full of detail, impossible to sort at volume. Structured logs are the same notes on standardized forms with labeled boxes you can file and query. Correlation is stamping every form touched during one investigation with the same case number (the trace ID): pull the case number from the surveillance timeline (the trace) and you instantly gather every form from every department for that one case - instead of searching the whole archive by hand. The case number turns scattered paperwork into a single coherent story.

Restore the missing trace IDs

BookVault logs are structured JSON and correlate perfectly - except log lines written inside an @Async notification method always have an empty traceId, so you can never tie a notification back to the request that triggered it. Why does this happen, and how do you fix it without threading the trace ID through by hand?

How do log lines get linked to the distributed trace they belong to?

Key takeaways

  • Structured logs are queryable events (typically JSON with named fields), not prose strings - so a log platform can filter and aggregate them.
  • In containers, log structured lines to stdout and let the platform collect, index, and retain them (twelve-factor).
  • Micrometer Tracing puts traceId/spanId in the MDC, so every log line during a request is stamped with them - the key to correlation.
  • The shared traceId lets you pivot both ways: from a slow span to all its logs, and from an error log to the full trace.
  • Propagate the MDC/tracing context across async and reactive thread hops, and put identifiers in fields (queryable) rather than baked into the message text.
Was this lesson helpful?
Edit this page on GitHub