Ask a team whether their service is reliable and you will usually get a confident yes. Ask what number that yes refers to, and the room goes quiet.
Closing that gap is the whole job of service level indicators and objectives. They are not complicated, and you do not need to adopt a framework to start using them. You need two things: something you can measure, and a number you are willing to defend.
An SLI is a measurement
A service level indicator is one specific measurement of how your service behaves, taken from the outside. Not CPU usage. Not queue depth. Something a user would notice.
The most useful SLIs are ratios — the proportion of events that went well, out of the events that counted.
SLI = good events ÷ valid events
For a web API that might be successful responses over total responses. For a checkout flow, completed purchases over attempted ones. For a nightly data job, runs that finished with fresh data over runs that were supposed to.
The word valid is doing real work there. If a user's browser gave up before the request ever reached you, that is not your failure and it does not belong in the denominator. Deciding what counts is half the work of defining a good SLI, and it is worth arguing about once rather than every incident.
An SLO is a target for that measurement
A service level objective takes an SLI and adds two things: a target, and a window.
"99.9% of requests succeed" is only half a sentence. Over what period? Over a day, one bad hour breaches it. Over a quarter, the same bad hour disappears into the average. The window is not a detail — it decides how the target behaves.
A complete SLO reads: 99.9% of valid requests return successfully, measured over a rolling 28 days. That is a claim somebody can check, disagree with, and hold you to.
What the target actually costs
Targets stay abstract until you convert them into time. Every nine you add divides the allowance by ten.
This is why "let's just do four nines" tends to end badly. Four minutes and nineteen seconds a month is less than one bad deploy. If your release process cannot reliably beat that, the number is a wish rather than an objective — and the first time you miss it, people stop believing any of your targets.
The pair gives you a budget
Here is the part that changes how a team behaves.
If your SLO says 99.9%, you have explicitly accepted 0.1% failure. Over a rolling 28 days that is about 40 minutes. Those 40 minutes are your error budget, and a budget is something you are allowed to spend — on a risky migration, a faster release cadence, an experiment you are not certain about.
Reliability stops being a vague pressure never to break anything, and becomes a quantity you can plan against. Budget remaining, ship. Budget spent, stabilise. Without the SLO, "are we safe to ship this?" is settled by nerve and seniority. With it, it is arithmetic, and the most junior person in the room can check the answer.
Where to start
Pick one journey your users would actually notice losing — signing in, checking out, loading a report. Write one SLI for it as a ratio. Then measure it for a few weeks and set no target at all, so you learn what your service already does.
Only then set the target, just above what you observed. A target you are already close to meeting is one the team can defend; you can tighten it later, once you know what tightening costs. Starting from an aspiration instead is how programmes end up with a dashboard nobody trusts.
The error budget calculator converts any target into minutes, so you can see what you would be committing to before you commit to it. When you are ready to decide which journeys deserve an SLO at all, the critical user journey walkthrough covers that step.