The Error Detector That Has Never Seen an Error

An agent finishes a document review at three in the afternoon. It has taken forty steps to get there: tool calls, retrievals, a few reasoning passes, some backtracking. The output is wrong. Not obviously wrong, which would be a mercy, but wrong in the way that survives a skim and dies in cross examination.

Now find the step where it broke.

This is the problem four researchers at the University of Wisconsin-Madison and Microsoft Research set out to solve, and their framing of it will be familiar to anyone who has ever reconstructed how a file went sideways. They call it failure attribution. Given a trajectory that ended badly, identify the step or steps that caused it. They note, almost in passing, that doing this by hand takes a human expert hours per trajectory, and that the root cause is often obscured by later steps which partially compensate for the earlier mistake. Every lawyer who has traced an error back through a chain of documents that each half corrected the one before knows exactly what that sentence describes.

The obvious response is to hand the transcript to the best model you own and ask it where things went wrong. The researchers tried that. On their in-domain test set, GPT-5 scored an F1 of 0.181 at identifying the failing steps. GPT-4o scored 0.212. A baseline that picks one step at random, with no reasoning, no context, and no model at all, scored 0.255.

Read that again, because it is the reason the paper exists. On this task, in this setting, the frontier reasoning model performed worse than a die roll.

The move nobody makes

Everyone who has looked at this problem has reached for one of two tools. Either you build a more elaborate prompting pipeline around a frontier model, which is expensive at inference and slow enough that nobody runs it on every matter, or you post-train a model on failure trajectories in which a human has labeled the step that went wrong.

The second option is the one that should give a practitioner pause. Where does that training data come from? Somebody has to sit down with a failed run and annotate, action by action, which one caused the failure. The researchers did exactly this for their own dataset, and they are candid about how it went: trajectories with ambiguous attribution had to be discussed among the annotators until consensus was reached. The question of which step caused the outcome turns out to be a judgment call, made by humans, disputed among humans, and expensive at every scale.

So the paper does something else entirely.

It trains only on trajectories that succeeded.

The insight is almost embarrassing in its simplicity. Nobody annotates their own mistakes, but everyone accumulates work that came out right, automatically, as a byproduct of operating. Successful runs cost nothing to collect and require no labeling at all. You already have them. They are sitting in your archive.

So the model learns what a well shaped path from question to answer looks like. It treats a trajectory as a continuous path through a latent space rather than as a list of discrete events, and it learns the flow of that path across roughly a hundred successful runs. Then, given a failed run, it scores each step by how far the actual path has drifted from where a successful path would have been at that point. High deviation, high suspicion. The detector has never been shown a single error, and it never needed a definition of one.

What it costs to run

The accuracy numbers matter less than the operating numbers, so take them quickly. In domain, the method reached an F1 of 0.435 against 0.181 for GPT-5, which the authors summarize as a twenty point improvement over the prompting baselines. Tested out of distribution, on a different benchmark built from multi-agent systems running on a different base model, it still came out ahead by about seven points, and this is the more interesting of the two results because it suggests that the learned sense of a well shaped trajectory transfers to settings the model was never trained on.

Now the part that changes how you would deploy it. The final model is a three layer network. It needs less than a gigabyte of video memory. It produces zero output tokens, because there is no language model in the loop at inference time. And it returns a verdict in about seven milliseconds, against roughly four seconds for GPT-4o and roughly forty seconds for GPT-5.

That gap of two to three orders of magnitude is what converts quality control from sampling into census. You would not review a suspicious subset of agent runs. You would review all of them, in real time, on hardware already sitting under a desk, without a single token leaving the building.

Four things a practitioner can take from this

  1. Your closed files are the asset, not your incident log. Firms hold almost no structured record of how work went wrong, because nobody documents their own errors at the granularity that would be useful. What firms do hold, in volume, is completed work that came out right. This paper is a demonstration that the second kind of record is enough to build a detector for the first kind of event, which inverts the usual assumption that quality control has to begin with a catalog of failure modes.

  2. The tool points at the origin rather than the symptom. In one of the case studies, an agent failed to retrieve a fact, invented an answer at step seven, and then built further steps on top of the fabrication. The detector flagged the invented step with a high score and the downstream steps with progressively lower ones. That decay is the useful behavior. An inherited error is less anomalous than the step that introduced it, and a reviewer whose attention is drawn to the origin instead of to the visible symptom is being sent to the right place. Anyone who has watched a mistaken assumption in a term sheet propagate through a dozen dependent documents will recognize why the ordering matters more than the raw detection rate.

  3. The sensitivity is a policy dial, not a technical parameter. The method offers two ways to decide how many steps to flag. One selects a fixed number of the most anomalous. The other uses conformal prediction, which sets the threshold from a held out sample of successful runs and carries a formal guarantee that the false positive rate on normal steps stays below a level you choose. That level is a supervision policy expressed as a number. How much noise will you tolerate in order to avoid missing something? That is a question for the partner responsible for the file, not for whoever configures the model.

  4. You do not need access to the model that did the work. The authors tested extracting the step representations with models entirely different from the one that generated the trajectories, and performance degraded only slightly. The practical consequence is that you can monitor an agent running inside a vendor’s closed system using your own local model as the observer. Oversight does not require privileged access to the thing being overseen, which is worth knowing when the thing being overseen belongs to somebody else.

Where it fails, which is where it gets interesting

The authors publish their failure cases, and those are more instructive than the successes.

The method reliably catches the step where a failure becomes explicit. It misses the quiet error at the beginning. In one case the agent misjudged, at step one, which tools were available to it, then spent the following steps pursuing a retrieval strategy that could not work, before announcing at step eleven that it was unable to complete the task. The detector scored step eleven very highly and did not flag step one at all. The explanation the authors give is honest: a reasoning trace saying that no suitable tool appears to be available is not anomalous in itself, because successful runs contain statements like that all the time, just before the agent finds another route.

For legal work, that limitation lands in the worst possible place. The quiet early misjudgment is the expensive one. By the time an error announces itself, the cost of the detour is already sunk. So this is a triage instrument that ranks steps by suspicion, and not an adjudication of cause, and it belongs in the hands of someone who will read the flagged steps rather than act on the flag.

There is a second limitation the paper does not dwell on, but which any firm should. The model’s definition of normal is whatever your successful trajectories happen to look like. Train it on your archive and it learns your habits, including the ones that have worked so far without being good. A genuinely better approach, unfamiliar to the training data, reads as deviation. Anomaly detection built on precedent will always be conservative about precedent, and a profession that already treats precedent as evidence should be alert to a tool that treats it as ground truth.

The authors close their own discussion of impact with a warning worth repeating: attribution tools of this kind could be used to make agents better at evading detection rather than better at the work, and automated attribution should support human oversight rather than replace it. That sentence was written by researchers about their own contribution, which is more than most vendors manage about theirs.

Back to three in the afternoon

The agent has taken forty steps and produced the wrong answer. You still have to find the break.

What has changed is where you would go looking. Not to a bigger model with a longer prompt, which this paper suggests would do worse than guessing. Not to a catalog of known failure modes, which nobody has and nobody is going to build. You would go to the hundred matters that came out right, ask what shape they had, and then look for the place where this one stopped having it.


Paper: Tracing Agentic Failure from the Flow of Success, Yeh, Zhu, Deep and Li (University of Wisconsin-Madison and Microsoft Research)

https://arxiv.org/abs/2607.12747v1

Torna alle news