Reliability Engineering

Learn how reliable systems are actually measured

Service level objectives, error budgets and the alerting that depends on them — explained with worked examples, interactive tools and the trade-offs you will meet in production.

Reliability is a measurement problem before it is an engineering one

Most teams know their service should be reliable. Far fewer can say what that means in numbers, or how much unreliability they can afford before it costs them users. This hub covers the measurement side: choosing indicators that reflect real user experience, setting targets you can defend, and spending the budget those targets create.

Everything here is built around worked examples rather than definitions. Read the fundamentals, then use the calculators to test the numbers against your own service.

What an SLO actually is

Three ideas do most of the work. Understand these and the rest of the practice follows.

Service Level Objectives

A Service Level Indicator (SLI) measures a specific aspect of your service — like request success rate or response time. A Service Level Objective (SLO) sets the target for that measure — for example, "99.9% of requests succeed". Together they define what good looks like for your users.

Read the introduction →

Level IntroductoryRead 3 min

Why SLOs Matter

SLOs help teams balance reliability with innovation, make data-driven decisions, and establish clear expectations with stakeholders about service quality.

Read the introduction →

Level IntroductoryRead 3 min

Key Benefits

Better incident response, improved prioritization, clearer communication with customers, and a shared understanding of acceptable reliability levels.

Read the introduction →

Level IntroductoryRead 3 min

The same idea, in two languages

An SLO describes a promise you make to users. Each card states that promise in everyday terms, then again as an indicator and target you could put in a dashboard. Select a card to turn it over.

Availability
99.9%
Coffee shop

A coffee shop opens at 7am every day. If it's locked or broken for more than ~43 minutes a month, regulars stop relying on it and go elsewhere.

Click to see the technical definition →
Availability
99.9%
In technology terms
SLI
% of HTTP requests returning a successful response (2xx/3xx)
Target
99.9% of requests succeed over a 30-day rolling window
Means
At most ~43 minutes of downtime per month before users lose trust
← Click to go back
Latency
< 200ms
Airport security

An airport promises that 95% of passengers clear security in under 10 minutes. A few complex cases take longer, but most travellers get through quickly and make their gate.

Click to see the technical definition →
Latency
< 200ms
In technology terms
SLI
HTTP response time distribution, measured at the 95th percentile
Target
p95 latency < 200ms, measured over a 1-hour window
Means
95% of users get a response in under 200ms — only 1 in 20 requests may be slower
← Click to go back
Error Rate
< 0.1%
Bank ATM network

A bank aims to complete at least 999 out of every 1,000 cash withdrawals successfully. One failed transaction in a thousand is tolerable — more than that and customers start queuing at branches.

Click to see the technical definition →
Error Rate
< 0.1%
In technology terms
SLI
Ratio of failed requests (5xx responses) to total requests
Target
Error rate stays below 0.1% per day
Means
No more than 1 error per 1,000 requests — leaving a small, predictable error budget
← Click to go back

Three steps, in order

Each step depends on the one before it. Skipping the first is the most common reason an SLO programme stalls.

1

Learn the Basics

Understand what SLIs, SLOs and error budgets are, and how the three depend on each other. Start with the fundamentals above.

2

Define Your SLOs

Map the critical user journeys your service supports, then define indicators that measure what those users actually experience.

3

Monitor and Iterate

Instrument the indicator, track the budget it produces, and revise the target once real traffic tells you whether it was right.