An AI factory that ends at "the tests pass" has not finished the job. It has moved the unfinished part somewhere less convenient.
Somebody still has to answer a page about the service the agent built, and that somebody will be a human with no memory of writing it, reading code they have never seen, at an hour when nobody reads carefully. Everything that would normally make that survivable — the author is in the building, someone remembers the trade-off, the design is still in somebody's head — has quietly been removed.
Why: Generation Outran Operational Readiness
This is the failure that will define the next couple of years of AI-assisted delivery. Capacity to produce code has raced ahead; capacity to operate it has not moved at all. The teams who get through it will be the ones who treat observability as a deliverable of the build rather than a follow-up ticket that never reaches the top of a backlog — and the ones who do not will accumulate services nobody can explain, faster than they ever could before.
The AI-DLC method definition paper gets the architecture right and the vocabulary wrong, in a way worth naming precisely. Its Operations phase has AI analyse metrics, logs and traces to detect patterns, identify anomalies and "predict potential SLA violations", integrate with predefined incident runbooks, propose actions such as resource scaling, performance tuning or fault isolation, and execute those resolutions once developers approve. Developers "serve as validators, ensuring AI-generated insights and proposed actions align with SLAs and compliance requirements".
Every structural decision there is correct. The noun is not.
An SLA is a contract with a penalty attached — a lagging, legal artefact. By the time you are predicting a violation of one, you are already in a conversation about service credits, and your options have narrowed to the expensive ones. An SLO is the internal target you steer by, set deliberately tighter, and the error budget is the headroom between them.
That distinction changes what an agent can usefully do, which is why it is worth more than a terminology quibble. "Predict an SLA violation" is a binary alarm that fires late and offers nothing to trade off — the answer is always the same, and it is always urgent. "This Unit has burned 60% of its 28-day budget in nine days, and at the current rate it exhausts on the 21st" is a continuous signal with proportionality built in. It supports a small response now or a large one later, and it distinguishes between them.
If you are handing an agent the authority to propose — let alone execute — scaling and fault isolation, that proportionality is the whole safety argument. An agent reasoning about a cliff edge has exactly one move and will make it every time. An agent reasoning about burn rate can tell the difference between a blip worth noting and a trend worth acting on, which is the difference between an assistant and an incident generator. Our CUJ to SLI to SLO walkthrough covers the ladder from journey to target if that distinction is new to your team.
What: Operations Belongs Inside the Lifecycle
The structural answer is to make operations a phase, not an afterthought. AI-DLC does this explicitly: the Operation phase owns the deployment pipeline and rollback runbook, environment provisioning, deployment execution with smoke tests and health checks, observability setup — dashboards, alarms and SLO config, incident response runbooks and an escalation matrix, performance validation against the NFRs, and finally an SLO report and cost analysis.
Note that SLO config is listed as an artefact of the run, sitting beside the dashboards and alarms — a thing the lifecycle is expected to produce, not a thing a team gets around to. And it does not start at the end. During Construction, each Unit gathers non-functional requirements covering performance, security, scalability, reliability and observability, so the instrumentation requirement is captured while the design is still being made and while changing it is still cheap.
There is a catch worth naming, because it is where this breaks in practice. Every Operation stage is conditional, and the whole phase is skipped for the MVP and proof-of-concept profiles. That is a defensible default — a PoC genuinely does not need an escalation matrix, and forcing one would guarantee the profile goes unused. But it is also precisely how unobserved services reach production, because a PoC that survives contact with real users gets promoted by somebody in a hurry, on a good day, with a demo behind them and no time to re-run anything.
Tooling will not stop them. The organisational rule has to: a profile that skips Operation may not serve production traffic until it has been re-run through one that does not. Write it down before you need it, because the argument is unwinnable once there is a customer on the other end.
How: Make "No SLO, No Merge" a Real Gate
Requiring observability as an artefact only means something if something checks. The practical version is a deterministic sensor, or simply a required CI job, that fails a change when the Unit it belongs to has no reliability definition attached.
Read that against what the method already collects and most of it is a promotion rather than an addition. Mob Elaboration is supposed to produce "Measurement Criteria that traces to the business intent" for every Unit at the very start of the lifecycle. The checklist above is simply that artefact made executable — the same criterion, expressed as a query that runs, a target committed to the repository, and an alert with a name on it. If your Inception ritual produces measurement criteria that cannot become those three things by the end of Construction, what it produced was prose.
The SLI line matters most, and it is the one most often fudged. An SLO written against a metric the service does not emit is a wish, and it will read as 100% forever — which is worse than having no SLO at all, because it actively reassures. Require that the instrumentation ships in the same change as the objective, so a reviewer sees the target and the query that feeds it side by side.
Agents are unusually good at producing the rest of this. The cost of a runbook was always tedium rather than difficulty, and tedium is exactly what they remove; the same goes for dashboard definitions, alert rules and the first draft of an escalation matrix. What they cannot supply is the judgement about what the target should be and who gets woken up. Both are business decisions wearing technical clothing — the first prices reliability against velocity, the second commits a named human's nights — and neither should be inferred from a codebase. Keep those two at the approval gate and let the machine do the typing.
Incident Response When the Author Is an Agent
Incident response leans on context that is quietly disappearing: somewhere in the building is a person who remembers why the code is like that. When the author was an agent, provenance has to replace them — and it can, provided you kept it. An audit trail recording which Intent and which Unit produced a change lets an investigation walk back from the failing deployment to the requirements, the design decisions and the gates behind it. That is a better trail than most human-authored code leaves; the qualifier is that it only exists if somebody decided to retain it before they needed it.
Your escalation matrix then needs an honest row for AI-generated components: who is the accountable human owner, and what is the fallback when nobody understands the code well enough to fix it safely under time pressure? Usually that fallback is the rollback runbook from the deployment pipeline stage — which is a strong argument for making sure that artefact is real and rehearsed, rather than a heading in a document generated at the end of a run nobody read.
Then run the post-incident review as you always would, blamelessly, with one addition: ask whether the lifecycle should have caught it. A missing NFR at the Construction stage, a sensor that should have been blocking rather than advisory, a rule that would have prevented the pattern entirely. In an AI factory the highest-leverage action item is rarely the code fix — it is the change to the loop that prevents the whole class of problem from being generated again next week, at speed, across every repository.
Close the Loop
The final stage produces an SLO report and a cost analysis, and its output feeds back to the start of the lifecycle as input to the next Intent. That arrow is the entire argument in one line: production reliability data is not the end of the pipeline, it is the beginning of the next one.
A factory that runs in the dark and never reads its own instruments will drift until something expensive happens. One that feeds SLO reports and error budget burn back into what it builds next is doing what SRE has always been for — using evidence about how the system behaves in the real world to decide what to change. That the builder is now a machine does not change the discipline. It only raises the cost of not having it.