Fresh, Complete, Correct: Setting SLOs for Data Pipelines

Availability SLOs make little sense for a nightly batch job. Learn how freshness, coverage and correctness SLIs bring SRE discipline to data pipelines.

← Back to Blog

The Job Succeeded and the Numbers Were Wrong

A nightly pipeline finishes with a green tick. The orchestrator is happy, the on-call engineer sleeps, and at 9 a.m. the finance team discovers that yesterday's revenue dashboard is showing the day before. Nothing failed in the sense a monitoring system understands. The data was simply stale, and by the time anyone noticed, decisions had already been made on it.

Request-driven services have a comfortable vocabulary for reliability: availability, latency, error rate. Pipelines are a poor fit for that vocabulary. A batch job is unavailable by design for 23 hours a day, and a streaming job can be up and processing while silently dropping a partition. Data teams need indicators built for what they actually promise, which is that data arrives on time, in full, and right.

Three SLIs Built for Pipelines

The Data Processing Pipelines chapter of the SRE Workbook lays out the indicator types that fit this world:

  • Freshness: how old is the newest data a consumer can see? Common shapes are that X% of records are processed within Y minutes, that the oldest unprocessed record is no older than Y, or that a scheduled run completes within Y of its trigger.
  • Correctness: does the output match what it should be? One approach the workbook describes is injecting inputs with known outputs and counting how often the pipeline gets them right.
  • Coverage: of the records that should have been processed, how many were? The pipeline exports both numbers and the ratio becomes the SLI.

Each one maps to a question a consumer would ask, which is the same test any good SLI has to pass. If your team has not yet run that exercise, the guide to user journeys, SLIs and SLOs applies to data products exactly as it does to APIs; the user is just an analyst or a downstream model instead of a browser.

Freshness Is the Easiest Place to Start

Most stacks already have the raw material. dbt's source freshness checks compare the newest loaded record against warn_after and error_after thresholds, which is a freshness SLI with two alert tiers already attached. Apache Airflow removed its old SLA feature in version 3.0 and replaced it in 3.1 with Deadline Alerts, which fire a callback when a run passes a deadline measured from a reference point such as the scheduled start, and unlike the old mechanism they fire even when the task never finishes. For streaming, consumer lag is the freshness signal, and LinkedIn's Burrow evaluates Kafka consumer lag as a service without static thresholds.

The step most teams skip is exporting these checks as time series rather than leaving them as pass or fail events in a scheduler log. Once freshness is a metric, it can carry a target, a window and an error budget like any other SLI.

Correctness Needs Tests, Not Vibes

Great Expectations treats data quality checks as unit tests for data: row counts within a range, no nulls in a key column, values drawn from an allowed set, distributions that have not drifted. dbt's built-in data tests cover the same ground closer to the transformation layer. A correctness SLI is then the proportion of validation runs, or of validated rows, that pass. Pair the automated checks with a small set of canary records whose correct output you know in advance; they catch the class of bug where every check passes and the join was still wrong.

Coverage and Knowing What Broke Downstream

Coverage sounds simple until a pipeline fans out across a dozen jobs and three teams. OpenLineage, a graduated project of the LF AI and Data Foundation, records datasets, jobs and runs in a standard model with integrations for Airflow, Spark, Flink and dbt. With lineage in place, a stale upstream table is not just a failed check; it is a list of every report and model that inherited the problem, which is the blast radius you need when deciding who to tell.

Error Budgets When the Unit Is a Run

Batch pipelines break the usual arithmetic because a 30-day window contains 30 runs, not 30 million requests. A 99.9% target would permit zero failures, which makes it a hope rather than a budget. Better to state the objective in the pipeline's own units: 28 of 30 daily runs land within two hours of schedule, or total lateness across the month stays under six hours. The error budget calculator is still the right tool for near-real-time pipelines where volume is high enough for a percentage to be meaningful. Whichever form you choose, attach a policy: when the freshness budget is gone, schema changes and new sources wait until it recovers. The original SRE Book chapter on pipelines observed that periodic jobs start out stable and are then quietly stressed by organic growth until they break, and a budget is how you notice the stress before the break.

Data pipelines increasingly feed AI systems too, where a stale retrieval index turns into a confidently wrong answer. The same freshness discipline is the foundation for the SLOs described for RAG pipelines.

This article was generated with the help of AI.