For twenty years the software factory was boring in the best possible way. A commit landed, a runner picked it up, a fixed sequence of steps ran, and the answer was pass or fail. When something broke you read the log, and the log told you which step failed, because the steps never changed.
An AI harness is not that. It is long-running, stateful and non-deterministic. It reads your repository, holds a conversation, dispatches sub-agents, writes files, runs commands and occasionally stops to ask permission. It has a working memory that fills up. It is, in every way that matters to an SRE, a distributed system that has been quietly promoted into the critical path of your delivery — and most teams are running it with no telemetry at all.
Why: The Old Feedback Loops Cannot Keep Up
Raja SP's AI-DLC method definition paper makes this case better than any monitoring vendor could. Traditional methods assumed iterations measured in weeks and months, and that assumption is exactly what made rituals like the standup and the retrospective viable — a human-paced feedback loop was fast enough to catch a human-paced process drifting. Properly applied AI, the paper argues, produces cycles measured in hours or days, which "needs continuous, real-time validation and feedback mechanisms, rendering many of the traditional rituals less relevant."
That is an observability argument wearing methodology clothing. When the delivery loop turns over hourly, no ceremony staffed by humans reading transcripts will keep pace. The paper goes further and asks whether velocity survives as a useful metric at all. It is a fair question, and the SRE answer is a specific one: you replace metrics that count effort with metrics that are measured continuously and tied to outcomes.
If that still sounds abstract, here are three questions teams running agents at scale ask within the first quarter, none of which a terminal transcript can answer. Why did every run on Tuesday afternoon stall at the same stage? Which of our forty repositories produces the most rework, and is it the codebase or the prompt? The provider shipped a model update on the seventeenth — did anything get worse, and can we prove it rather than argue about it?
There is a fourth question, and it arrives with the invoice. Token spend is now a production cost that varies with how well the system is working: a run that thrashes costs several times what a clean run costs, so cost per merged change is a reliability signal wearing an accountant's coat. Teams that cannot attribute spend to a workflow discover this the month somebody enables parallel unit execution across forty repositories.
What: Four Layers Worth Instrumenting
Observability still answers the question it always answered — can you ask something new about the system's behaviour without shipping code to find out? For an agent harness that means four layers, and they differ enormously in volume and in what they cost to keep.
Start at the stage layer. It is where the equivalent of a "request" lives, it is the cheapest to collect, and it answers the question that is probably costing you the most today. A stage has a name, a lead, inputs, outputs, and an outcome of completed, revised or skipped. That is a span with a status, and almost every factory SLI worth having is a ratio over those outcomes.
The session layer is thin but worth having early, because it carries the events that explain otherwise baffling behaviour — resumptions, and the compaction that marks the moment a conversation outgrew its window. A stage that succeeded before compaction and failed after it is telling you something specific.
The tool-call layer is where you eventually see the difference between a model that is thinking and one that is thrashing: the same file read eleven times, an edit reverted and reapplied, a test suite run against unchanged code. It is also high-cardinality and expensive, so sample it. Keep every tool call for runs that failed or were rejected at a gate, and sample the successes at five or ten percent. You are keeping the evidence for the investigations you will actually run.
The artifact and cost layer closes the loop between engineering and money, and it needs no sampling at all because there is one record per deliverable.
One caution that will save you an awkward conversation. Agent telemetry contains prompts, and prompts contain source code, customer identifiers and occasionally credentials somebody pasted in. Decide before you start what is a metric, what is an attribute and what is a payload. Durations, outcomes and identifiers are safe to keep indefinitely. Prompt and completion bodies are a different asset class, and they belong behind the same controls as the repository they were drawn from, with a retention window measured in days rather than years.
How: Treat the Audit Trail as a Span Stream
You may not have to build the event stream at all. AI-DLC, the AWS Labs implementation of that methodology, already writes one. Every intent carries an append-only audit trail, and the engine emits 99 typed event types into it:
STAGE_STARTED STAGE_REVISING STAGE_COMPLETED
GATE_APPROVED GATE_REJECTED HUMAN_TURN
SENSOR_FIRED SENSOR_PASSED SENSOR_FAILED
BOLT_STARTED BOLT_COMPLETED BOLT_FAILED
SWARM_STARTED SWARM_UNIT_FAILED RULE_LEARNED
That is not a log file, it is a schema. The names are stable, the transitions are paired, failures have their own types rather than being buried in prose, and every row carries an ISO timestamp. Which makes the translation mostly mechanical.
Send those spans to the backend that already holds your service traces, rather than buying a separate "AI observability" product on day one. The payoff is correlation, and it is worth more than any purpose-built dashboard: when a change ships and an SLO starts burning, the trace of the agent run that produced it sits one click away — the same move that exemplars make between a metric spike and the span that caused it. Emit the event payloads as structured logs rather than prose, and keep the field names identical to the audit event types so nobody maintains a translation table.
Resist one temptation while you are wiring this up: do not start by building a bespoke "agent analytics" pipeline with its own storage and its own query language. Every team that does it ends up maintaining two observability stacks and correlating between them by hand at exactly the moment correlation matters most. The harness is a distributed system, you already run a backend for distributed systems, and the unglamorous choice is the one that still works in a year.
Then restate the golden signals in the harness's own terms.
Two of those have traps in them. Latency must separate machine time from time spent waiting at an approval gate, or your percentiles will measure your team's calendar rather than your system — and the fix for each is completely different. A slow stage is a prompt or context problem; a slow gate is a staffing problem, and no amount of model tuning will touch it.
Errors must include GATE_REJECTED, which is not a crash and will not appear in any error rate you inherit from a service. It is a human looking at the output and sending it back — the single most informative failure signal the system produces, because the judge had to live with the result.
Saturation deserves a note too, because it does not look like saturation. A saturated web server returns 503s. A saturated agent session simply forgets what it was doing: quality degrades smoothly and nothing anywhere reports an error. Context-window headroom before compaction is the closest thing to a queue depth you have, and it belongs on the dashboard next to the ones that look more familiar.
Start With One Question
You do not need all four layers this week, and a team that tries to build the complete picture before shipping anything usually ships nothing. Pick the question costing you the most — for most teams it is "how often does a run need human rework before it is accepted?" — and instrument only what answers it.
That single ratio is the first SLI of your delivery factory. Once it is on a wall, the arguments change character: "the agents are getting worse" acquires a date and a number, and the case for the next layer of instrumentation tends to make itself.
A sequence that works: stage events first, because they are cheap and they carry the SLIs. Then cost per merged change, because it is the number that buys you the budget for everything after it. Then sampled tool calls, once you have a specific investigation that needs them. Then prompt and completion capture, last and behind access controls, because by then you will know precisely which runs are worth keeping and for how long.
The destination is not a wall of dashboards nobody reads. It is the ability to answer a question you have not thought of yet, about a system that is now writing a meaningful share of your software — and to answer it the same afternoon, rather than at a retrospective three weeks after it stopped mattering.