The SLO That Lived in a Dashboard
Most SLO programmes start the same way. Someone builds a dashboard, adds a panel with a 99.9% line, and writes the recording rule by hand. Six months later the target reads 99.5%, nobody remembers who changed it or why, the recording rule in a sister team's repository computes the same ratio slightly differently, and the burn rate alert was tuned once during an incident and never revisited. The SLO exists, but not in any form that can be reviewed, tested or trusted.
Treating SLOs as code fixes that by giving reliability targets the same lifecycle as the software they describe: a file in a repository, a pull request to change it, a history of who agreed to what, and tooling that turns the declaration into rules and dashboards nobody has to hand-craft.
What "as Code" Actually Buys You
- Review: a target change is a diff that a product owner and an engineer can both approve, which is how the Implementing SLOs chapter of the SRE Workbook says reliability decisions should be made.
- Consistency: every team's availability SLO is computed the same way, from the same template, so 99.9% in one service means the same thing as in another.
- Generated machinery: recording rules, multi-window burn rate alerts and dashboards come from the definition rather than being maintained beside it.
- Portability: a definition that is not tied to one vendor's clicks can move when the monitoring stack does.
OpenSLO: A Vendor-Neutral Way to Write It Down
OpenSLO is an Apache 2.0 licensed YAML specification for describing services, SLIs, SLOs and alert policies independently of any tool. Version 1 of the specification is stable and a second version is being developed in the open. Nobl9 was a founding member and its platform imports OpenSLO documents directly, as documented in its SLOs as code guide, while open source generators can consume the same files. The value is less about any one tool and more about having a shared language that a service catalogue, a CI check and a monitoring backend can all read.
Sloth and Pyrra for Prometheus Shops
If Prometheus is your source of truth, two open source projects do the heavy lifting. Sloth takes a compact YAML definition and generates the recording rules, the page and ticket alerts using the multi-window, multi-burn-rate method, and a Grafana dashboard. It runs as a CLI in a pipeline or as a Kubernetes operator with custom resources, and it also accepts OpenSLO input. A definition looks like this:
version: "prometheus/v1"
service: "checkout"
labels:
owner: "payments-team"
slos:
- name: "requests-availability"
objective: 99.9
description: "Checkout API availability as seen at the load balancer."
sli:
events:
error_query: sum(rate(http_requests_total{job="checkout",code=~"(5..|429)"}[{{.window}}]))
total_query: sum(rate(http_requests_total{job="checkout"}[{{.window}}]))
alerting:
name: CheckoutAvailabilityBurn
page_alert:
labels:
severity: page
ticket_alert:
labels:
severity: ticket
Twenty lines of intent become a full set of correct PromQL rules, and every service in the organisation gets the alert behaviour described in burn rate alerting without anyone re-deriving the thresholds.
Pyrra approaches the same problem with a Kubernetes-native ServiceLevelObjective custom resource. Its operator watches those objects, generates recording and alerting rules, and serves a UI that shows each SLO's error budget and burn rate over time. Teams who already manage everything through Kubernetes manifests and GitOps tend to find it the most natural fit.
Managed Platforms Speak Terraform Too
Commercial backends have their own declarative paths. Datadog exposes SLOs through the datadog_service_level_objective Terraform resource. Grafana SLO generates a dashboard, recording rules and alert rules for each SLO it manages and recreates them if they are deleted. Google Cloud's service monitoring models SLIs, SLOs and compliance periods as API resources with alerting policies on error budget burn. The pattern is identical: the definition lives in version control, and the platform renders it.
A Workflow That Sticks
Keep SLO definitions in the service's own repository, under a directory the platform team provides a template for. Have CI validate the files and generate the rules on every pull request, so a broken query fails the build rather than silently producing an SLO that always reads 100%. Require three things in the pull request description: the user journey the SLI measures, the reason for the target, and a link to the error budget policy that will govern it. The error budget calculator helps reviewers sanity-check what a proposed target means in minutes. Treat a change to the objective the way you would treat a change to a database schema: it is small, it is easy, and it is exactly the kind of change that deserves a second pair of eyes.