Reliability for RAG: Setting SLOs for AI Retrieval Pipelines

Latency SLOs alone will not keep a RAG pipeline reliable. Learn to measure retrieval quality, generation quality and cost as SLIs your team can act on.

← Back to Blog

A Green Dashboard and an Unhappy User

Retrieval-augmented generation (RAG) pipelines break a comfortable assumption. Traditional services fail loudly: a request returns a 500, latency spikes, a queue backs up. A RAG pipeline can return HTTP 200 in 400 milliseconds with a fluent, confident, completely wrong answer. Every classic signal stays green while the product quietly fails.

That gap is why availability and latency alone are insufficient targets for AI systems. Reliability here has to include whether the answer was any good.

Measure Three Layers, Not One

A RAG request passes through distinct stages, and each fails in its own way. Treating the pipeline as one opaque box makes incidents almost impossible to diagnose, so define indicators at each layer:

  • Retrieval: did the search return documents that actually contain the answer? Offline evaluation sets scored for recall and precision give you a measurable target here.
  • Generation: did the model use those documents faithfully? The failure mode to watch is a claim that is not supported by any retrieved passage.
  • End-to-end: latency, availability and cost for the full request, including time to first token if the interface streams.

Splitting the layers pays off during an incident. When answer quality drops, the first question is whether retrieval stopped finding the right documents (often an indexing or embedding problem) or the model stopped using them well (often a prompt or model version change). Those have entirely different fixes.

Turning Quality Into an SLI

Quality feels subjective, but it can be sampled and scored like any other signal. The common production pattern is LLM-as-judge: route a sample of live completions to a separate evaluator model that scores them against explicit criteria such as factual support, relevance and instruction following. Sampling 10 to 20% of traffic is generally enough for a stable trend without a prohibitive bill.

Two cautions are worth building in from the start. Judges drift, so keep a small human-labelled set to periodically check the judge itself, and treat a judge model upgrade as a change that can move your metric independently of the system it measures. Publish the sampled score as a metric and it becomes a genuine SLI, with all the usual machinery available: a target, an error budget, and a burn rate.

Instrument With Open Standards

Bespoke logging for AI systems ages badly. The OpenTelemetry GenAI semantic conventions define shared names for model calls, token usage and operation duration, so telemetry means the same thing regardless of which provider or framework produced it. Capture the retrieval step and the generation step as separate spans on one trace, and a slow or wrong answer becomes something you can follow end to end. The tracing concepts documentation is a good starting point if distributed tracing is new to your team, and the OpenTelemetry project's walkthrough of GenAI observability shows what a fully instrumented model call looks like in practice.

Budget for Wrongness, and Watch the Bill

No AI system is perfectly accurate, so decide explicitly how much wrongness the product can absorb. A 97% faithfulness target is a real objective with a real budget: 3 wrong answers in 100 is either acceptable for an internal research assistant or unacceptable for medical guidance, and naming the number turns an anxious debate into a design decision. The AI SLO and error budget guide walks through applying this to AI workloads, and the error budget calculator converts a target into concrete allowances.

Finally, track cost per successful answer as a first-class signal. Token spend that suddenly climbs usually means retrieval is returning more context than it should, or the model is retrying. In AI systems the bill is often the earliest indicator that reliability is about to follow.

This article was generated with the help of AI.