Why SLOs Matter: What Changes When Reliability Has a Number

SLOs turn reliability from an argument into a trade. How a single agreed number changes ship decisions, prioritisation and the conversation between engineering, product and customers.

← Back to Blog

Every engineering team has a version of the same stalled conversation. Someone wants to ship. Someone else thinks the system has been fragile this week. Both are arguing from the same absence of evidence, and whoever is more senior, or simply more tired, tends to win.

A service level objective does not make that conversation friendlier. It makes it shorter, because there is finally something to point at.

They turn a standoff into a trade

Reliability and speed get framed as opposing camps, which is exactly why the argument never resolves — nobody can be against either one in the abstract.

An SLO reframes it. If you have agreed 99.9%, you have also agreed to tolerate 0.1% failure. That remainder is an error budget, and a budget turns a disagreement about values into a decision about a resource. Is there budget left? Then the risky migration is affordable this week. Is it gone? Then it is not, and the next work is stabilising rather than shipping.

Nobody has to win the argument that reliability matters. It has a price, and the price is visible to everyone.

The same decision, with and without an objectiveTwo columns contrasting how a ship-or-hold decision is made. Without an objective the inputs are impressions and the loudest voice decides. With one, the inputs are budget figures and the number decides.Without an SLO• “It feels fragile”• “It's been fine for me”• Loudest voice decides• Re-argued every releaseWith an SLO• 68% of budget left• Spent mostly on checkout• The number decides• Re-checked, not re-argued
The argument does not get friendlier. It gets shorter, because there is something to point at.

They replace impressions with evidence

"The service has felt flaky lately" is a real observation and a nearly useless input. It cannot be checked, compared against last month, or acted on proportionately.

The same concern measured against an SLO becomes: we have spent 60% of this month's budget in nine days, almost all of it in the checkout journey. That sentence says there is a problem, roughly how large, where it lives, and how urgent it is — before anyone has opened a dashboard.

It works in the other direction too, which teams consistently underrate. An objective that is met comfortably for two quarters running is also telling you something: either the target is too loose to be informative, or you are buying reliability nobody asked for and could spend that effort somewhere it would be noticed.

Both readings are only available because a number exists. Without one, a good quarter and a lucky quarter look identical.

They give everyone the same sentence

The quiet benefit is organisational. Engineering, product, support and customers usually describe service quality in four different vocabularies, and most reliability arguments are really translation failures between them.

An SLO is a single statement each group can read for its own purposes.

One number, four audiencesA service level objective read by four different groups: engineers see how much risk the week can carry, product sees when to spend a cycle on reliability, support sees whether today is unusual, and customers see what the service commits to.Engineershow much risk this week's work can carrybudget remainingProductwhen to spend a cycle on reliability instead of featuresburn rateSupportwhether today is unusual or business as usualcurrent statusCustomerswhat the service commits to, without reading a dashboardthe target
The same objective answers a different question for each group — which is why it ends arguments between them.

This is also what makes an objective safer than a promise. A team that says "we take reliability seriously" has committed to nothing and can still be accused of anything. A team that says "99.9% of checkout requests succeed over 28 days, and here is the current figure" has committed to something specific — and, crucially, has also stated what it is not promising. Both halves protect you.

What actually changes

Three things, in roughly this order.

Escalation gets calmer. A single bad hour stops being an emergency by default. You check what it cost against the budget, and most of the time the honest answer is "not much" — which is a far better conversation than arguing about how bad it felt.

Prioritisation gets easier. Reliability work competes with feature work on the same evidence rather than on advocacy. "We are burning budget in this journey" is a roadmap argument that product can actually evaluate.

Targets start getting revised. This is the sign the practice has taken hold. A team that has never changed an SLO is not using it; a team that tightens one because the old target stopped being informative, or loosens one because it was costing more than it returned, is.

None of this requires adopting a methodology. It requires one journey, one indicator, one number, and the willingness to look at it when the argument starts. If you have not set the first one yet, the introduction to SLIs and SLOs covers the mechanics, and the error budget calculator shows what a given target commits you to in minutes.

This article was generated with the help of AI.