Dear Counsel: The Agent Will (Almost Always) Tell You Everything Went Fine

A customer writes to her auto insurer because her payment is seven days overdue, and she asks for an extension on $210. The agent does all the preliminary work flawlessly: it pulls her profile from the CRM, verifies her identity against her date of birth, checks the status of the policy, opens a ticket and, before doing anything else, queries the history of payment arrangements. That history, however, reports that she had already been granted two extensions in the previous twelve months, which for her customer tier is exactly the ceiling the policy allows. The agent grants the third one anyway, then closes the ticket as solved and sends her a courteous summary, complete with the new due date and an updated count of three arrangements.

No tool returned an error and the conversation ended in an orderly fashion: the ticket looks worked and the customer is satisfied. The only thing that does not work here is the matter itself, because the extension should never have been granted, and there is now an irregular operation sitting in the billing system that no one has any reason to go looking for.

That conversation is reproduced in full in the appendix to ThinkingBox, the sandbox and benchmark that Microsoft published on August 20, 2026 together with researchers from the University of Pittsburgh, Northwestern University and the University of California, Irvine. It is an actual failure by a frontier model, reported by the authors turn by turn, and for anyone working in a law firm it is more instructive than any leaderboard.

A benchmark that looks at the database, not at the answer

The idea behind ThinkingBox is easy to explain and uncomfortable to accept, namely that almost every evaluation of agents measures whether the model picked the right tool and built the arguments of the call correctly. The authors argue that this tells us nothing about whether the work was actually done, because in an operational flow with persistent state you can call the correct API on the wrong entity, update a record before you have collected the confirmation you needed, or produce an impeccable message while leaving the backend exactly as it was.

The benchmark contains 507 tasks spread across five domains: retail, hotel booking, auto insurance, internal support at a neobank, and IT and HR support at a consulting firm. Each task is an executable world with an initial state, tools exposed over MCP in isolated sessions, a domain policy and a simulated user who releases information only if the agent asks for it. The final verdict compares the terminal state of the database against the expected state and checks for side effects, accepting different paths as long as the world ends up the way it should. For 477 of the 507 tasks the judgment concerns the backend alone; only 30 add a check on the response given to the user.

The number that should concern us as lawyers

The authors ran an analysis that one rarely sees in the benchmark literature, in the sense that they took 79,853 failed trials out of the 121,680 recorded across twelve models and asked how many of those trials would have looked successful to an observer who only examined the surface.

In 84.86% of cases the conversation closed cleanly, with a final response that was not empty, contained no questions and signaled completion. In 80.88% the agent had also carried out at least one genuine state-changing action on the backend, so it had not merely asserted something. In 67.24% the last tool response did not even contain an explicit error. On the opposite side, the comparison against the expected state found a database mismatch in 98.95% of the failed trials, a wrong field value in 77.61%, and effects beyond those required in 43.30%.

If we carry these numbers into a law firm, the control that is normally exercised over an agent coincides almost entirely with its own report: what it says it has done gets read, the tool trace may get skimmed, and if there are no obvious errors the matter is treated as closed. Those three figures tell us that the report and the state of the world are unrelated quantities, and that the band in which they diverge is not marginal. Four failures out of five, in this benchmark, get past exactly the kind of check a lawyer typically performs.

Finding a solution is not the same as knowing how to repeat it

The second interesting result in the paper concerns consistency, since the authors found that the best model reaches 65.36% on the first attempt and, given twenty attempts per task, succeeds at least once in 91.12% of cases. If instead you require all twenty attempts to succeed, the figure drops to 25.25%.

In absolute terms, out of 507 tasks there are 45 the model never completed and 128 it completed every time; the remaining 334 sit in a middle band where the outcome depends on the individual attempt.

That middle band is the real problem, because it is indistinguishable from the outside. An associate who systematically botches the same kind of file can be identified within a week and trained, whereas an associate who botches one out of three identical files, in a different spot each time, requires close and continuous review of all three, and at that point the time we save by delegating has already been consumed by the checking.

Where the wall is highest

Model performance on retail averages around 52%, while on auto insurance it drops to roughly 23%. The authors do not attribute the difference to the number of tasks, which is comparable, but to the policy and interaction burden that each domain imposes.

There is one figure that makes the point well. In the diagnostic category the authors call wrong state update, meaning an action executed without technical errors but wrong on the merits, the insurance domain accounts for 32.9% of its own failures, against 4.5% for travel.

The difference lies in the fact that in insurance the action depends on an eligibility check that has to happen before, not after.

This is the same structure as our own work, where the question that matters concerns whether the conditions are met, while the technical feasibility of the operation is almost always beyond dispute. In the case of the payment extension, the agent had in fact read the very data point that made the request inadmissible, and acted anyway, which means it had the information and did not treat it as a constraint.

The reservations I share, and the one I would shift

The most common criticisms of this kind of work are well founded and worth taking seriously. The workflows are synthetic, as the authors openly state, so ecological validity is limited. The simulated user is the same for every model evaluated and behaves impeccably: it invents no facts, does not change its goal halfway through, stays cooperative even after repeated failures by the agent, answers only what it is asked and speaks at most ten times. Anyone who works with real clients knows how far that description sits from reality. Finally, the verdict remains the product of binary checks against a predefined outcome, so prudent decisions that would be correct in practice risk being counted as failures.

There is, however, one limitation the authors themselves acknowledge, and in my view it is the most relevant one for us as lawyers, because it turns the perspective around. Given that for 477 tasks the verdict looks only at the backend, a trial that performs the correct operation and then describes it badly to the user is still counted as a success. The benchmark therefore measures precisely what we in the firm never see, namely the state of the systems, and it does not measure what we do see every day, namely the quality of the report. The two halves of the problem remain separate, and whoever puts an agent into production has to deal with both.

What changes in practice

On the simpler domains the success rate is substantial, and the traces reported in the paper show matters handled cleanly from start to finish, so the problem has more to do with where we place the control than with the ability of agents to do the work.

As long as verification amounts to reading the summary, we are verifying the wrong variable. What is needed instead is that every action with an effect on the world leaves a readable trace in the destination system, independent of the agent’s own account, and that someone actually looks at that trace. What is also needed is that the human confirmation point sits before the irreversible act rather than after it, because in the extension case the useful moment to step in was the one where the history returned “two out of two,” and downstream of that call the conversation was already formally impeccable.

The opposite case in the paper is worth a look as well, the one where a customer asks for her insurance ID card without remembering either her policy number or her security answer. The agent verifies what it can, establishes that the second factor is missing, opens a ticket, puts it on hold and refuses to issue the document. The task counts as passed precisely because the agent did not act. This is a good thing, because a control system that rewards completion alone would not be able to tell that prudence apart from a matter left undone, and in a law firm a reasoned refusal to act is often the correct outcome.

The lawyer’s desk test

Before handing an agent a flow that touches files, deadlines or accounting systems, four questions are worth asking, in my view.

  1. First gate. Can I read the outcome without going through the agent’s account of it? If the only evidence that the operation took place is the agent’s closing message, what I am holding is a statement by the party being checked. What I need instead is a trace in the destination system that I can open on my own;

  2. Second gate. Does eligibility depend on a condition to be verified before acting? If so, that condition has to be readable by a tool and it has to block the action, rather than merely inform the agent. The third arrangement case shows that reading a constraint and honoring it are two different things;

  3. Third gate. Where does the irreversible act sit, and does the confirmation point come before it? The filing, the certified email, the closing of a position, the payment. Human confirmation placed after the summary arrives once the effect has already been produced, and by then it is too late;

  4. Fourth gate. Am I measuring the success rate or the consistency? A flow that works four times out of five is fine for a draft and not fine for a deadline. The right question is how many times in a row it works on the same matter, because that is the threshold beyond which I can tell how far to trust it.

If even one gate stays shut, the agent can still do almost all of the work, gathering material, preparing the file, drafting the act and stopping one step short, and that is probably where it is worth keeping it for now.


Sources

Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai, Calvin Wang, Mirco Milletari, Tuhin Kundu, Vadim Smolyakov, Kjartan Ólafsson, Tommy Guy, One Success Isn’t Reliability: ThinkingBox, a Sandbox and Benchmark for Agents in Stateful Business Workflows, arXiv:2608.19741, August 20, 2026.

Paper: https://arxiv.org/abs/2608.19741

Code: https://github.com/microsoft/thinkingbox

Back to news