SLIs and SLOs: What Good Looks Like, in Numbers

An introduction to service level indicators and objectives — how to turn a vague promise about reliability into a measurement, a target and a budget your team can spend.

← Back to Blog

Ask a team whether their service is reliable and you will usually get a confident yes. Ask what number that yes refers to, and the room goes quiet.

Closing that gap is the whole job of service level indicators and objectives. They are not complicated, and you do not need to adopt a framework to start using them. You need two things: something you can measure, and a number you are willing to defend.

An SLI is a measurement

A service level indicator is one specific measurement of how your service behaves, taken from the outside. Not CPU usage. Not queue depth. Something a user would notice.

The most useful SLIs are ratios — the proportion of events that went well, out of the events that counted.

SLI  =  good events  ÷  valid events

For a web API that might be successful responses over total responses. For a checkout flow, completed purchases over attempted ones. For a nightly data job, runs that finished with fresh data over runs that were supposed to.

The word valid is doing real work there. If a user's browser gave up before the request ever reached you, that is not your failure and it does not belong in the denominator. Deciding what counts is half the work of defining a good SLI, and it is worth arguing about once rather than every incident.

An SLO is a target for that measurement

A service level objective takes an SLI and adds two things: a target, and a window.

"99.9% of requests succeed" is only half a sentence. Over what period? Over a day, one bad hour breaches it. Over a quarter, the same bad hour disappears into the average. The window is not a detail — it decides how the target behaves.

A complete SLO reads: 99.9% of valid requests return successfully, measured over a rolling 28 days. That is a claim somebody can check, disagree with, and hold you to.

From a user journey to an error budgetFour steps: choose a user journey, define a service level indicator as good events over valid events, set an objective with a target and a window, and the difference between the target and one hundred percent becomes your error budget.Pick one user journeysigning in, checking out, loading a reportDefine the SLIgood events ÷ valid eventsSet the SLOa target plus a rolling windowYou now have an error budgetthe failure you have agreed to accept
Each step depends on the one before it. Skipping the first is why SLO programmes stall.

What the target actually costs

Targets stay abstract until you convert them into time. Every nine you add divides the allowance by ten.

What each target allows in downtimeAvailability targets converted to allowed downtime over thirty days: ninety-nine percent allows seven hours twelve minutes, 99.9 percent allows forty-three minutes, 99.95 percent allows twenty-one and a half minutes, and 99.99 percent allows four minutes nineteen seconds.99%7 hours 12 minutesper 30 days99.9%43 minutesper 30 days99.95%21 minutes 36 secondsper 30 days99.99%4 minutes 19 secondsper 30 days
Every extra nine divides the allowance by ten, and usually multiplies the engineering cost by more.

This is why "let's just do four nines" tends to end badly. Four minutes and nineteen seconds a month is less than one bad deploy. If your release process cannot reliably beat that, the number is a wish rather than an objective — and the first time you miss it, people stop believing any of your targets.

The pair gives you a budget

Here is the part that changes how a team behaves.

If your SLO says 99.9%, you have explicitly accepted 0.1% failure. Over a rolling 28 days that is about 40 minutes. Those 40 minutes are your error budget, and a budget is something you are allowed to spend — on a risky migration, a faster release cadence, an experiment you are not certain about.

Reliability stops being a vague pressure never to break anything, and becomes a quantity you can plan against. Budget remaining, ship. Budget spent, stabilise. Without the SLO, "are we safe to ship this?" is settled by nerve and seniority. With it, it is arithmetic, and the most junior person in the room can check the answer.

Where to start

Pick one journey your users would actually notice losing — signing in, checking out, loading a report. Write one SLI for it as a ratio. Then measure it for a few weeks and set no target at all, so you learn what your service already does.

Only then set the target, just above what you observed. A target you are already close to meeting is one the team can defend; you can tighten it later, once you know what tightening costs. Starting from an aspiration instead is how programmes end up with a dashboard nobody trusts.

The error budget calculator converts any target into minutes, so you can see what you would be committing to before you commit to it. When you are ready to decide which journeys deserve an SLO at all, the critical user journey walkthrough covers that step.

This article was generated with the help of AI.