Reliability Engineering
Learn how reliable systems are actually measured
Service level objectives, error budgets and the alerting that depends on them — explained with worked examples, interactive tools and the trade-offs you will meet in production.
Why this exists
Reliability is a measurement problem before it is an engineering one
Most teams know their service should be reliable. Far fewer can say what that means in numbers, or how much unreliability they can afford before it costs them users. This hub covers the measurement side: choosing indicators that reflect real user experience, setting targets you can defend, and spending the budget those targets create.
Everything here is built around worked examples rather than definitions. Read the fundamentals, then use the calculators to test the numbers against your own service.
01 — Foundations
What an SLO actually is
Three ideas do most of the work. Understand these and the rest of the practice follows.
Service Level Objectives
A Service Level Indicator (SLI) measures a specific aspect of your service — like request success rate or response time. A Service Level Objective (SLO) sets the target for that measure — for example, "99.9% of requests succeed". Together they define what good looks like for your users.
Read the introduction →Why SLOs Matter
SLOs help teams balance reliability with innovation, make data-driven decisions, and establish clear expectations with stakeholders about service quality.
Read the introduction →Key Benefits
Better incident response, improved prioritization, clearer communication with customers, and a shared understanding of acceptable reliability levels.
Read the introduction →02 — Worked examples
The same idea, in two languages
An SLO describes a promise you make to users. Each card states that promise in everyday terms, then again as an indicator and target you could put in a dashboard. Select a card to turn it over.
A coffee shop opens at 7am every day. If it's locked or broken for more than ~43 minutes a month, regulars stop relying on it and go elsewhere.
- SLI
- % of HTTP requests returning a successful response (2xx/3xx)
- Target
- 99.9% of requests succeed over a 30-day rolling window
- Means
- At most ~43 minutes of downtime per month before users lose trust
An airport promises that 95% of passengers clear security in under 10 minutes. A few complex cases take longer, but most travellers get through quickly and make their gate.
- SLI
- HTTP response time distribution, measured at the 95th percentile
- Target
- p95 latency < 200ms, measured over a 1-hour window
- Means
- 95% of users get a response in under 200ms — only 1 in 20 requests may be slower
A bank aims to complete at least 999 out of every 1,000 cash withdrawals successfully. One failed transaction in a thousand is tolerable — more than that and customers start queuing at branches.
- SLI
- Ratio of failed requests (5xx responses) to total requests
- Target
- Error rate stays below 0.1% per day
- Means
- No more than 1 error per 1,000 requests — leaving a small, predictable error budget
03 — Apply it
Three steps, in order
Each step depends on the one before it. Skipping the first is the most common reason an SLO programme stalls.
Learn the Basics
Understand what SLIs, SLOs and error budgets are, and how the three depend on each other. Start with the fundamentals above.
Define Your SLOs
Map the critical user journeys your service supports, then define indicators that measure what those users actually experience.
Monitor and Iterate
Instrument the indicator, track the budget it produces, and revise the target once real traffic tells you whether it was right.