SLOs, Dashboards & Alerts
Aus Signalen Betrieb machen: SLIs, SLOs und Error-Budgets, die RED- und USE-Methode für Dashboards und Alerting auf Symptome statt auf Rauschen.
Deutsche Übersetzung in Arbeit
Diese Lektion ist noch nicht ins Deutsche übersetzt und wird daher auf Englisch angezeigt. Der Rest der Seite ist vollständig lokalisiert.
Auf dieser Seite
Instrumentation produces signals; operations is turning those signals into decisions - what "healthy" means, what to put on a dashboard, and when to wake someone up. This final lesson covers the vocabulary of reliability - SLIs, SLOs, and error budgets - and the discipline of dashboards and alerts that make BookVault something you can run, not just watch.
SLI, SLO, error budget
Three terms, often confused, that give reliability a number instead of a vibe:
- SLI (Service Level Indicator) - a measured quantity about the service's behavior. E.g. "the proportion of checkout requests that succeed in under 300 ms." An SLI is a metric you actually collect.
- SLO (Service Level Objective) - the target for an SLI over a window. E.g. "99.5% of checkouts succeed under 300 ms over 30 days." It's the line between acceptable and not.
- Error budget - the allowed shortfall:
100% − SLO. A 99.5% SLO grants a 0.5% error budget - about 3.6 hours of failure per 30 days. Spend it however: a bad deploy, a slow dependency, a spike.
The error budget is the powerful idea. It turns "don't break things" into a quantity: while budget remains, you can ship boldly; when it's exhausted, you freeze features and fix reliability. It aligns dev and ops around one number instead of arguing about "stable enough."
An SLO isn't 100%
Chasing 100% is a trap - each extra nine costs exponentially more and users rarely notice past a point. The right SLO is 'reliable enough that users are happy,' which leaves an error budget to spend on shipping features. A team with budget to spare is being too cautious; a team always over budget is shipping too recklessly. The budget is the dial.
Dashboards: RED and USE
A dashboard should answer a question fast, not display every metric you have. Two proven recipes:
- RED (for request-driven services like BookVault's API) - Rate (requests/sec), Errors (failures/sec or %), Duration (latency distribution, p50/p95/p99). Three panels per endpoint tell you if users are being served well.
- USE (for resources like the DB, thread pools, hosts) - Utilization, Saturation, Errors. Tells you if a resource is running out of headroom.
Most of http.server.requests (auto-instrumented last-but-one lesson) already gives you RED for free. Build the
dashboard around the user's experience (RED on key endpoints) first, and keep resource (USE) panels for
diagnosis, not the front page.
Alerting: on symptoms, not causes
The final discipline: alert on what users feel, not every internal wobble.
- Alert on SLO burn - "we're spending error budget fast enough to blow the month's SLO." This is symptom- based: it fires when checkouts are actually failing users, whatever the cause. Burn-rate alerts (budget consumed over a short and a long window) catch both sudden outages and slow leaks without flapping.
- Avoid cause-based alert spam - paging on "CPU 85%" or "one pod restarted" wakes people for things users never noticed, and trains them to ignore the pager (alert fatigue). If high CPU isn't hurting the SLO, it's a dashboard line, not a 3 a.m. page.
The test for every alert: would a human need to act right now to protect users? If not, it's a metric to watch, not an alarm.
Bad alerting is every monitor in the ward beeping at once - a slightly-off reading here, a sensor glitch there - until the nurses tune it all out and miss the patient who's actually crashing (alert fatigue). Good alerting is triage tied to the patient's condition: you're paged when vital signs cross the line that means this person needs help now (SLO burn), regardless of which machine noticed. The error budget is the chart showing how much stability you have banked this month; the RED dashboard is the bank of vitals you glance at to see if patients are being cared for. You act on the patient's outcome, not on every device's noise.
BookVault's checkout endpoint currently succeeds under 300 ms about 99.7% of the time. Product accepts a 99.5% target over 30 days. Compute the error budget in time. Then decide: should a 'database CPU at 90%' alert page someone at 3 a.m.? What alert should?
Why should production alerts fire on SLO burn (symptoms) rather than on causes like 'CPU at 85%'?
Key takeaways
- An SLI is a measured indicator (e.g. % of fast successful checkouts); an SLO is the target for it over a window; the error budget is 100% − SLO, the allowed failure.
- The error budget turns reliability into a spendable quantity: ship boldly while budget remains, freeze and fix when it's exhausted; the right SLO is not 100%.
- Build dashboards from proven recipes: RED (Rate, Errors, Duration) for request services, USE (Utilization, Saturation, Errors) for resources - front-page the user experience.
- Alert on symptoms (SLO burn-rate), not causes: page only when users are actually being harmed, whatever the root cause.
- Cause-based alerts (CPU 85%, a pod restart) cause alert fatigue when users feel nothing - keep those on dashboards, not the pager.