The Benefits of SLOs: What They Change During an Incident

Better incident response, clearer prioritisation and honest customer updates all come from the same place — a number everyone agreed on before anything broke.

← Back to Blog

It is 3am and something is wrong. The dashboard is amber in four places. Two alerts have fired, one of which fires most weeks and has never mattered.

The question the on-call engineer actually needs answered is not "what is broken". It is how much does this matter, and does it matter enough to wake anyone else. Without an agreed number, that judgement gets made by a tired person under pressure, and it gets made differently each time.

Detection stops being about thresholds

Most alerting starts life as a set of thresholds on machine state — CPU above 80%, memory climbing, a queue getting long. These are useful for diagnosis and poor for waking people, because they answer the wrong question. A saturated CPU that nobody notices is not an incident. A healthy CPU while checkout fails is.

An SLO moves the trigger to the thing users experience. You are no longer asking whether a machine is unhappy. You are asking how fast the budget is being spent, which is called the burn rate.

Burn rate decides the responseFour burn rates and the response each warrants: 14.4 times for one hour and 6 times for six hours both page someone, 3 times for a day raises a ticket, and 1 times sustained is spending exactly as planned.14.4x for 1 hour2% of the month gone — wake someonepage6x for 6 hours5% of the month gone — still urgentpage3x for 1 day10% gone — handle it in hours, not minutesticket1x for 3 daysspending exactly as plannedno action
Burn rate is how fast you are spending the budget. 1x means you end the month exactly on target.

Burn rate is simply pace. At 1x you are consuming the budget exactly as fast as the month allows, and will finish precisely on target. At 14.4x you will burn 2% of a month in an hour, which is worth someone's sleep. At 3x you have a real problem that can wait until morning. The multi-window approach combines a fast and a slow window so that brief spikes do not page and slow decay does not go unnoticed.

The practical effect is fewer pages that turn out not to matter, which is the single biggest determinant of whether people trust their alerts at all.

Severity stops being a debate

Most severity scales are prose — "significant customer impact", "degraded experience" — and prose gets argued about at exactly the moment nobody has time. Budget spent gives you a number instead: an incident that consumed a fifth of the month is not the same event as one that consumed a twentieth, whatever either felt like in the room.

That matters most for the incidents in the middle. The catastrophic ones declare themselves. The genuinely ambiguous ones are where consistent classification saves the most time, and where a shared number does the most work. Our incident management guide goes further into severity and escalation.

Customer updates get concrete

"We are investigating elevated error rates" tells a customer nothing they can act on. It is also, quietly, a promise you have not defined — elevated compared to what?

A team with a published objective can be specific: what the commitment is, where the service currently sits against it, and what that means for the customer today. The value is not transparency for its own sake. It is that you have stated in advance what normal looks like, so an incident update is a comparison rather than an adjective. And because the objective also states what you are not promising, a bad hour inside budget does not have to be reported as a crisis.

The review argues about the number

The most durable benefit arrives after everything is fixed.

A post-incident review without an SLO tends to circle the same questions: was this bad enough to act on, and who should have caught it. With one, the first question already has an answer — the incident cost a measurable share of the budget — so the conversation moves to what to change, and the ranking against other work is evidence rather than advocacy.

An SLO through the life of an incidentFour stages of an incident with the objective involved at each: detection alerts on user impact, triage uses budget spent to set severity, communication states status against the commitment, and the review measures what the incident cost.Detectthe alert fires on user impact, not on CPUTriagebudget spent so far sets the severityCommunicatestatus stated against the commitmentReviewwhat it cost, and what to fix first
The same number carries through all four. That continuity is most of the benefit.

That continuity is most of the benefit. The same number detects the problem, sizes it, describes it to customers and ranks the fix — so four groups who would otherwise reason in four vocabularies are looking at one thing.

None of this requires a large programme. It requires one journey, one indicator and one agreed number. If you have not set that yet, start with the introduction to SLIs and SLOs, then use the error budget calculator to see what a target commits you to before you commit to it.

This article was generated with the help of AI.