Error Budgets for Autonomous Delivery: Governing Agent Speed Without Losing Control

Autonomy should be earned, not assumed. Learn how an error budget policy decides when an AI-DLC run may skip approval gates and when the guardrails come back automatically.

← Back to Blog

Most teams decide how much to trust an agent by temperament. The cautious lead keeps every approval gate switched on forever and wonders why the factory is slower than doing the work by hand. The bold one turns them all off after a good week, and finds out in production what the gates were for.

Neither is measuring anything. Both are guessing — and the guess gets re-litigated every sprint, usually by the same two people, usually with the same two anecdotes.

This is a solved problem. Not solved for AI specifically, but solved: SRE settled the equivalent argument about production change more than a decade ago, and the mechanism transfers with almost no modification.

Why: The Method Names the Tension But Not the Number

The AI-DLC method definition paper is unusually candid that this is a genuine trade-off rather than a free lunch. One of its stated principles is "minimise stages, maximise flow" — cut handoffs and transitions so work moves continuously, which is the entire promise of iterations measured in hours rather than weeks. A few lines later, the same principle insists that human validation stays critical so that AI-generated code "does not become rigid ('quick-cement') but stays adaptable for future iterations", describing those validation points as acting "like a loss function", identifying and pruning wasteful downstream effort before it occurs.

The trade-off an error budget resolvesTwo columns. Removing gates buys flow and shorter iterations but risks rigid code. Keeping gates prunes errors early and keeps code adaptable but costs velocity.Every gate you remove• Faster flow• Shorter Bolts• Less waiting on humans• Risk: 'quick-cement'Every gate you keep• Earlier error pruning• Adaptable code• Traceable decisions• Risk: lost velocity
The method asks for oversight that is "minimal but sufficient". A budget is what makes sufficient a number.

Both halves are right, which is precisely what makes the trade-off real rather than rhetorical. Strip out the gates and you get flow — and you also get code that nobody understood well enough to change three months later. Keep every gate and you get adaptable code that arrives too late to matter, produced by a factory costing more than the team it was supposed to accelerate.

The paper resolves the tension with a phrase rather than a number: oversight should be "minimal but sufficient". Sufficient for what, judged how, on whose evidence, is left to the practitioner — and that gap is where most AI delivery programmes quietly come apart. Not because the principle is wrong, but because sufficient, without a measurement attached, is just the more persuasive person's opinion.

It is worth being clear about why this is not a change advisory board with better branding. A CAB is a fixed tax: every change pays the same review cost regardless of evidence, and because nothing measures whether the reviews catch anything, the cost never comes down. An error budget inverts that. The tax floats, it is set by the factory's own recent track record, and a team that demonstrably produces clean work pays less of it. That is the difference between governance that decays into theatre and governance that stays honest.

What: Autonomy as a Governed Setting

The dials already exist. They are simply not connected to anything.

An AI-DLC run declares a workflow profile that sets its stage route, artefact depth, test strategy and ceremony. The published profiles run from express at 10 of 33 stages with minimal depth and minimal tests, through feature at all 33 stages at standard depth, up to enterprise at all 33 at comprehensive depth — which deliberately retains the compliance, security, observability and incident-response work "because the cost of an undocumented decision is higher than the cost of the additional ceremony".

Inside Construction there is a second dial. After the first approved stage, the engine fires a one-time ladder prompt asking whether the remainder of the run should be autonomous or gated. Autonomous skips the remaining per-stage gates; gated keeps them. The answer is recorded in state and honoured when the session resumes.

So a run's autonomy is already a setting. On most teams today it is chosen by whoever happens to be at the keyboard, based on how the last one went. The proposal is only that you choose it from evidence instead.

Four SLIs are enough to start, and you can collect all of them from a typed audit trail plus your existing version control:

  • Gate first-pass rate — approvals without revision, over all gate outcomes. Quality as judged by a human who had to live with the result.
  • Sensor pass rate — deterministic checks passed, over checks fired. Correctness with no model in the judging seat.
  • Rework ratio — stage revisions over stage completions, always broken down per stage.
  • Escaped defect rate — changes needing a fix within 14 days of merge. The only one sourced from production rather than the factory, and therefore the only one that cannot be gamed from inside it.

Give each a target and a rolling 28-day budget — rolling rather than calendar-month, so a bad fortnight is not absolved by a date change. Then write the policy as an operations manual rather than an aspiration.

An error budget policy for autonomous deliveryFour budget states from healthy to exhausted, each granting a different level of agent autonomy: full speed above fifty percent remaining, narrowing through the middle bands, and a full stop when the budget is spent.Above 50% remainingautonomous construction · parallel batches · express profilefull speed25% to 50%autonomous for bugfix and refactor only · features return to gatednarrowedBelow 25%every profile gated · parallelism off · one unit at a timeguardedExhaustedno new autonomous work until the factory's own reliability is fixedstop
Autonomy is earned and withdrawn automatically. Nobody has to win an argument to make the factory safer.

The thresholds are yours, and you should expect to tune them at least twice in the first quarter. What matters is that the decision now has an owner, a number and a trigger. Nobody has to win an argument to make the factory safer on a bad week, and nobody has to ask permission to go faster on a good one.

One structural note. The platform team should own the policy; each delivery team should own its own budget. A single organisation-wide budget produces the worst outcome available — one team's bad fortnight throttles everybody, and no team ever feels the consequences of its own work.

How: Close the Loop, and Alert on the Burn

A monthly budget report tells you about a problem three weeks after it started. The policy above only does anything if budget state actually reaches the place where runs are configured, and reaches it quickly.

The governance loop for an AI delivery factoryA four-step cycle: measure the factory SLIs, compute the error budget state over a rolling 28-day window, set the next run's profile, autonomy and parallelism from that state, and let the run emit the events that feed the next measurement.Measuregate pass · sensor pass · rework · escaped defectsCompute budget staterolling 28-day window, per teamSet run configurationprofile · autonomy · parallelismNext run executesand emits the events you measureevery run feeds the next
The policy only works if budget state actually reaches the run configuration. Close this loop or it is a wall poster.

Then apply multi-window burn rate alerting to the factory SLOs exactly as you would to a service: a fast window to catch an acute regression, a slow window to catch the drift that no single run would ever flag.

The acute case is not hypothetical. A provider updates a model, or somebody edits a rule that ships in every prompt, and the sensor pass rate drops fifteen points within hours. A 14.4x burn rate over a one-hour window catches that the same afternoon, while the change responsible is still the most recent thing anyone touched. Without it, you find out a fortnight later, when a reviewer gets tired of rejecting things and mentions it in a retro.

The slow case is subtler and considerably more common. A codebase grows past the point where the agent can hold enough of it in context, and the first-pass rate declines by a point a week. No single run looks wrong; every individual rejection has a plausible local explanation. A six-hour window evaluated over three days is what surfaces it.

Budgets govern the aggregate, so you still need a fast local stop for the individual failure. AI-DLC's halt-and-ask is the right pattern: when a unit's code generation fails, Construction stops immediately — even in autonomous mode — and offers retry, skip or abort. If one unit in a parallel batch fails, the healthy units finish and keep their artefacts, and only the failed one raises a question. That is a circuit breaker with a human as the fallback path, and it is what makes the autonomous setting defensible at all. Autonomy that cannot stop itself is not autonomy; it is an unattended loop.

Finally, decide in advance what exhaustion means, because that promise is what makes the rest of the policy credible. When the budget is spent, the factory's own reliability becomes the work: no new autonomous runs until somebody has read the audit trail for the failures, found the pattern, and fixed the rule, the sensor or the prompt that produced it. It is the equivalent of a production freeze, and it will be unpopular exactly once.

Why This Makes You Faster, Not Slower

The objection arrives immediately: this is bureaucracy with extra steps. It is the opposite, for the same reason error budgets make product teams faster in production.

Without a budget, the only defensible posture is permanent caution, because no evidence would ever justify relaxing. Every gate you add is therefore forever, and the factory ratchets steadily back toward the ceremony it was built to remove. With a budget, caution is priced. A team that consistently produces clean runs accumulates headroom and spends it on speed — wider parallel batches, fewer gates, lighter profiles — and everyone can see exactly why they were allowed to. That is the flexibility half of the bargain. The resilience half is that the same mechanism claws the speed back automatically, without anyone having to be the villain.

It also makes failure legible. A budget that is never spent does not mean the factory is excellent; it usually means the targets are too loose, or that the factory is running well below the speed it could sustain. A budget consumed steadily and deliberately is a factory operating at the edge of its designed envelope, which is where you want it.

Start with one SLI. Take four weeks of gate first-pass data, set the target just above observed performance, and write two lines of policy: above target, Construction may run autonomous; below it, the gates come back. Run it for a month. You will spend a great deal less time arguing about whether the agents can be trusted, because you will be able to look it up.

And the same discipline applies to the AI features you ship, not only to the ones that build them — our guide to AI SLOs and error budgets covers the production side of that coin.

This article was generated with the help of AI.