Burn Rate Alerts: Catch SLO Breaches Before They Land

Learn how multi-window, multi-burn-rate alerting turns your error budget into pages that matter: fewer false alarms and faster detection of real outages.

← Back to Blog

The Problem with Alerting on Raw Error Rates

Most teams start their monitoring journey with a threshold: page someone when the error rate goes above 1%. It feels reasonable, and as the Google SRE Book chapter on monitoring distributed systems argues, it fails in two directions at once. A brief spike that self-heals in ninety seconds wakes an engineer for nothing, while a slow, steady 0.9% error rate quietly drains a month of reliability without ever tripping the alert. Neither outcome tells the on-call engineer what they actually need to know: is this failure fast enough, and large enough, to matter?

Burn rate alerting answers that question directly, because it measures failures against the budget the service is allowed to spend rather than against an arbitrary percentage.

What a Burn Rate Actually Measures

A burn rate is the speed at which a service consumes its error budget, expressed as a multiple of the sustainable rate. A burn rate of 1 means the budget will be exactly exhausted at the end of the compliance window. A burn rate of 2 means it will be gone in half the time. A burn rate of 14.4 against a 30-day window means the entire month of allowed failure disappears in roughly 50 hours.

This reframing matters because it makes alert severity proportional to user harm. A 99.9% availability target permits about 43 minutes of downtime per 30 days. If ten of those minutes vanish in a single hour, that is worth waking someone. If they trickle away over three weeks, that is a ticket, not a page. To see how a target translates into concrete minutes for a given service, work through the error budget calculator.

Why Two Windows Beat One

A single measurement window forces an unpleasant trade-off. Short windows detect incidents quickly but fire on transient noise. Long windows are stable but slow, and worse, they keep firing long after an incident is resolved because the bad data is still inside the window.

The multi-window, multi-burn-rate pattern described in the Google SRE Workbook solves both problems by requiring two conditions to be true simultaneously: a long window confirms the burn is sustained, and a short window (typically one twelfth of the long one) confirms it is still happening right now. When the incident stops, the short window clears quickly and the alert resolves.

A widely used starting configuration for a 30-day window uses three tiers:

  • Burn rate 14.4 over 1 hour, which corresponds to consuming 2% of the budget. Page immediately.
  • Burn rate 6 over 6 hours, which corresponds to 5% of the budget. Page.
  • Burn rate 1 over 3 days, which corresponds to 10% of the budget. Open a ticket instead of paging.

Putting It Into Practice

Evaluating several windows on every scrape is expensive, so pre-compute the error ratios with Prometheus recording rules and let the alert expressions reference those series. Start with the paging tiers only, run them in parallel with your existing alerts for a few weeks, and compare what each would have caught.

Expect to tune. A low-traffic service can hit an extreme burn rate from a handful of requests, so add a minimum-volume condition before it pages. Above all, the alert is only as meaningful as the indicator underneath it, so make sure the SLI reflects a real user journey before wiring it to a pager. If that groundwork is still ahead of you, start with mapping critical user journeys to SLIs and SLOs.

The Payoff

Teams that adopt burn rate alerting usually end up with far fewer alerts and far more trust in the ones that remain. That is the real goal: an on-call rotation where a page means something is genuinely wrong, and silence means users are being served.

This article was generated with the help of AI.