Metrics, Logs & Traces
Monitoring answers 'is it up?'; observability answers 'why is it slow?'. The three pillars, how they complement each other, and what each is for.
On this page
The Production module wired up Actuator and a few Micrometer metrics - enough to know BookVault is up. This module is about knowing why it's slow at 2 a.m. when nothing is technically down. That shift - from monitoring (is it up?) to observability (why is it behaving this way?) - rests on three complementary signals: metrics, logs, and traces. Knowing what each is for is the foundation everything else builds on.
Monitoring vs. observability
Monitoring watches known failure modes: CPU high, disk full, endpoint returning 500s. You set it up because you anticipated those questions. It answers "is a thing I predicted happening?"
Observability is the property that lets you ask questions you didn't predict - "why are checkouts from German members slow only on the catalog endpoint after the last deploy?" - without shipping new code. A system is observable when its outputs (the three pillars) are rich enough to reconstruct its internal state after the fact.
The distinction matters because production failures are rarely the ones you predicted. Monitoring tells you that something's wrong; observability lets you find out what.
The three pillars
Each pillar answers a different question, and their power is in the overlap:
Metrics - aggregated numbers over time. Request rate, error rate, p99 latency, active DB connections. Cheap to store (they're just numbers), great for dashboards and alerts, and they tell you that something changed - "error rate jumped at 14:03." But an aggregate can't tell you which requests failed or why.
Traces - the path of one request across services and components. A trace shows that a single catalog request spent 12 ms in the controller, 340 ms in a database query, and 190 ms calling the search service. Traces tell you where the time or error is - which hop, which span.
Logs - discrete, timestamped events with detail. "Loan 8842 rejected: member over limit." Logs tell you what exactly happened at a point in time, with the specifics metrics and traces omit.
The pillars are a workflow, not a menu
The signals shine when you move between them: a metric alert fires (error rate up) → you open traces to find the slow/failing span (the DB call to one shard) → you jump to the logs for that trace to read the exact exception. Metrics find it, traces localize it, logs explain it. The next lessons wire all three so you can navigate between them by a shared trace ID.
Why all three
Any one pillar alone leaves a blind spot:
- Metrics only - you know error rate doubled, but not which requests or why. Blind to specifics.
- Logs only - infinite detail, but at scale you're grepping millions of lines with no idea where to look, and aggregates (p99 latency) are painful to compute.
- Traces only - you see one request's path, but not whether it's representative or a one-off, and not the overall trend.
Together they cover that (metrics), where (traces), and what/why (logs) - the three questions every incident asks.
Metrics are the ward's vital-signs monitors - aggregate heart rate and temperature trending on screens, instantly showing something is wrong across patients. A trace is one patient's journey through admissions, X-ray, surgery, recovery - showing exactly where the delay or complication occurred in their path. Logs are the detailed case notes at each station - the specific readings, drugs given, the surgeon's remark. A doctor needs all three: the monitors to notice trouble, the journey to localize it, the notes to understand it. Watch only the monitors and you know a patient is deteriorating but not why.
During an incident you ask three questions in order: (1) 'Did something get worse, and when?' (2) 'Which part of the request is slow - our code, the database, or a downstream call?' (3) 'What was the exact error message and the member ID that triggered it?' Match each question to the pillar that answers it best, and explain why that pillar and not the others.
What distinguishes observability from monitoring?
Key takeaways
- Monitoring answers predicted questions ('is it up?'); observability lets you investigate unanticipated ones ('why is this slice slow?') without shipping new code.
- The three pillars: metrics (aggregated numbers over time), traces (one request's path across components), logs (discrete detailed events).
- Metrics tell you THAT something changed, traces tell you WHERE, logs tell you WHAT/WHY - each covers the others' blind spot.
- Use them as a workflow: a metric alert fires, traces localize the slow/failing span, logs explain the exact cause.
- No single pillar is enough - real diagnosis moves between all three, ideally linked by a shared trace ID.