SLIs for the Dark Factory: Turning AI-DLC Stage Events into Service Level Indicators

A lights-out software factory still needs SLIs. Learn how to turn AI-DLC gate, sensor and bolt events into good-over-valid ratios your team can measure and defend.

← Back to Blog

Manufacturing has a name for a plant that runs without people in the building: the dark factory. The lights stay off because nothing on the floor needs to see. Humans set the tolerances, watch the instruments from somewhere else, and walk in only when something needs a decision a machine should not make.

That is the shape a lot of software organisations are now building toward. Agents pick up an intent, elaborate the requirements, decompose the work, write the code, run the tests and open the pull request. The humans are still accountable — they are just not in the loop for every step. They are watching the instruments.

Which is fine, right up until you ask what the instruments are actually showing.

Why: A Dark Factory With No Instruments Is Just a Dark Room

Before you take your hands off the controls, "working" has to mean something in numbers. And the numbers most teams have are the wrong ones.

Velocity measures how much effort a team expended. It says nothing about whether the factory is producing changes anyone can trust, and it becomes actively misleading when the marginal cost of producing a story point collapses. The AI-DLC method definition paper raises exactly this doubt, asking whether velocity stays relevant at all when AI erases the boundary between simple, medium and hard tasks, and floating business value as the replacement.

DORA metrics get closer, and if you already have them, keep them. Deployment frequency and lead time for change measure the pipeline honestly. But they were designed to measure a pipeline operated by humans, and two of the four assume a human author in ways that stop holding. Change failure rate is the closest cousin to what you need — it is the one metric in the set that a factory cannot game from inside itself — while mean time to restore says nothing about whether the thing you restored should have been built that way. DORA tells you the pipeline is fast. It does not tell you the agent is trustworthy, and those are now different questions.

The replacement is the discipline we already teach for user-facing services: start from the critical user journey, not from the metrics you happen to have lying around. The journey for a delivery factory is short to state and awkward to measure.

A developer describes an intent. Some time later, a reviewed, tested change is merged — and the developer did not have to rescue it along the way.

"Some time later" is your latency SLI. "Reviewed and tested" is your quality SLI. "Did not have to rescue it" is your autonomy SLI, and it is the one most teams skip, because it is the one that might return an uncomfortable answer.

What: The Method Already Asks for This

You are not inventing from scratch. AI-DLC puts a ritual called Mob Elaboration at the heart of its Inception phase — product owner, developers, QA and stakeholders in one room with a shared screen, AI proposing the breakdown of an Intent into user stories and Units while the humans correct whatever is over- or under-engineered. The paper claims this ritual condenses weeks of sequential work into a few hours, and lists what it must produce: a PRFAQ, user stories, non-functional requirement definitions, a description of risks matching the organisation's risk register, the suggested Bolts, and "Measurement Criteria that traces to the business intent".

The AI-DLC artefact hierarchyA vertical flow: an Intent decomposes into Units, each Unit is built across one or more Bolts, and the result is a tested Deployment Unit. A dashed arrow returns from the end to the start, labelled measurement criteria.Intenta high-level statement of purposeUnitsloosely coupled, independently deployableBoltsthe smallest iteration — hours, not weeksDeployment Unittested, operations-readymeasurement criteria
The factory's critical user journey, in the method's own vocabulary. Measure at every level.

That last item is an SLI in all but name, and it is demanded at exactly the right moment — while the humans are still in the room agreeing what the work is for, and before anyone has written code that would be awkward to instrument. Retrofitting measurement is expensive precisely because the decisions that make a system measurable are design decisions, and by the time the code exists they have all been made.

So the failure mode is not that the method forgets to ask. It is what comes back when it does.

Prose versus a measurable criterionTwo columns contrasting vague measurement criteria written as prose with the same intent expressed as service level indicators that have thresholds, windows and denominators.Measurement criteria as prose• “Fast and reliable”• “Good recommendation quality”• “Minimal errors”• “Scales well”Measurement criteria as an SLI• p90 under 400ms, 28-day window• 99.5% return a result• CTR within 2pts of baseline• a ratio with a denominator
The method asks for measurement criteria. The left column is what teams actually write.

Everything on the left is a real thing a real team has written into a real requirements document, and none of it can be computed. The test is mechanical: can someone who was not in the room write a query that returns this number? If not, what you have is a shared feeling, and shared feelings do not survive the quarter.

How: Five Ratios You Can Collect Today

Good SLIs are ratios of good events over valid events, as the SRE Workbook puts it. If your harness writes a typed audit trail, you are not instrumenting anything — you are choosing which rows to count.

Five SLIs for an AI delivery factoryFive service level indicators, each expressed as a ratio of good events over valid events, labelled by the dimension it measures: quality, correctness, throughput or latency.Gate first-pass rateapproved without revision ÷ all gate outcomesqualityRework ratioSTAGE_REVISING ÷ STAGE_COMPLETEDqualitySensor pass rateSENSOR_PASSED ÷ SENSOR_FIREDcorrectnessUnit convergenceconverged ÷ (converged + failed)throughputIntent lead timeWORKFLOW_STARTED → _COMPLETED, p50 and p90latency
Every one is a ratio of good events over valid events. None of them needs a new agent to collect.

Gate first-pass rate is the closest thing to a quality SLI the harness produces on its own, because the judge is a human who had to live with the result. When it falls, something upstream has degraded — a prompt change, a model change, an unfamiliar codebase — and you will know within a day rather than at the next retrospective. It also carries the sharpest measurement trap in the set. If you publish it as a team metric, humans learn to approve and then quietly fix, and your number improves while your factory does not. Pair it with the rework ratio, which counts the fixing, and watch the two together: a first-pass rate rising while rework rises is not an improvement, it is a reporting artefact.

Rework ratio is only useful broken down per stage. A factory-wide ratio of 0.2 sounds tolerable right up until you discover it is 0.9 concentrated entirely in domain design — at which point you have a specific, fixable problem instead of a vague sense that the agents are bad at architecture.

Sensor pass rate is the one number on the list with no model and no human in the judging seat, which makes it the most trustworthy and the least flattering. It moves when the code genuinely stops compiling or type-checking, and it is the first thing to look at after a provider ships a model update.

Unit convergence tells you whether parallelism is producing throughput or just expensive churn. Below about 0.9, the decomposition is wrong more often than the code is — which is a problem in Inception, not Construction, and no amount of retrying units will fix it.

Intent lead time should be split into machine time and human-wait time before you optimise anything. Most teams discover the harness is fast and the queue in front of a reviewer is not.

One SLI is missing from that list on purpose, because it cannot come from the harness. Escaped defect rate — changes needing a fix within fourteen days of merge — is sourced from production rather than from the factory, which makes it the only measure in the set that the factory cannot flatter itself with. Every other number here is the system grading its own homework, honestly but from inside. Add this one as soon as you can attribute a fix back to the change that caused it, and treat disagreement between it and your internal numbers as the most interesting signal you have: a factory with excellent gate first-pass rates and a rising escaped defect rate is one where the reviews have stopped reviewing.

Spend some care on the denominator. "Valid events" should exclude runs a human abandoned for reasons that have nothing to do with output quality — a changed priority, a meeting, an intent that turned out to be a duplicate. Leave those in and your SLIs will drift with your calendar. This is the same discipline as excluding synthetic traffic from an availability SLI, and it is the step most often skipped.

Then set targets from evidence, not ambition. Run the factory for a few weeks collecting the SLIs and changing nothing, and set each objective just above observed performance, so the first month's error budget is tight but not already spent on day three. The error budget calculator makes it concrete: a 95% gate first-pass SLO over 200 gates a month gives you ten rejections to spend before you have breached — a number a team can reason about, unlike "the AI should mostly get it right".

Finally, write the definitions where they can be reviewed and argued with. The case for SLOs as code applies here with more force than it does to production services, because factory SLIs are newer, less standardised and far more contested. A definition living in one person's dashboard will drift the moment that person changes teams, and nobody will notice until the number stops meaning what everybody assumed.

What This Buys You

Instrumentation turns arguments into measurements. "The agents are getting worse" becomes a first-pass rate with a date on it. "We should add another approval gate" becomes a question about which stage actually carries the rework, answerable in an afternoon.

And it produces the raw material for the harder decision waiting behind all of this: not whether the factory works, but how much autonomy it has earned — which is a question you can only answer with a budget, and a budget you can only build on numbers like these.

This article was generated with the help of AI.