Every Legal Benchmark Has a Lawyer Hidden Inside It

A man loses a large sum from his current account over the course of nine days. He complains to the bank, receives a letter telling him the transactions were properly authenticated, and decides to bring the matter before an out-of-court board that costs him almost nothing and does not require him to instruct a lawyer.

He has a chatbot open beside the form, because he has been using one all year for everything else, and he tells it his story the way he would tell it to his daughter. He starts several years back, when he opened the account. He mentions that the branch manager was very kind. Around the eleventh sentence he notes, in passing, that a man who said he worked in the fraud office called him and that he read out a code from a text message. Then he goes back to the branch manager.

That one clause decides the case. The model has no way of knowing this, because nothing in the way the story was told marked it as the hinge. It writes him a fluent submission about the bank’s duty of care.

Two different skills wearing one name

A paper accepted to the AI4Law workshop at ICML 2026, by Andrew Lou and David Shin of Yale Law School, makes an argument that is simple to state and hard to put back down. Legal AI research keeps justifying itself by appealing to access to justice, while legal AI benchmarks keep measuring something else.

The distinction they draw is between legal reasoning and lawyering.

Legal reasoning is the application of doctrine to a question that has already been put into legal form. Lawyering is the work of putting it into that form, which means extracting the operative facts from a rambling account, discarding the kind branch manager, reordering what remains by legal relevance rather than by chronology, and noticing what the client has failed to say so that you can ask him about it.

Every benchmark the authors survey evaluates the first skill on inputs that have already had the second performed on them. By the time a benchmark prompt reaches a model, a substantial amount of expert cognitive labor has been done for free.

The point sharpens when you consider what one of those benchmarks is made of. LEXam is built from real law school examinations, and a well-written exam hypothetical is the most heavily processed legal input in existence. A professor constructed it so that every fact needed for the answer is present, every distractor is there deliberately, and the procedural posture is stated rather than inferred. It is a laboratory specimen of a legal question. Scoring well on it demonstrates something real, and it demonstrates it under conditions that no client has ever produced.

Why the two numbers can move apart

Lou and Shin describe this as a gap between an upper and a lower bound. The upper bound is model performance when a competent professional does the framing. The lower bound is performance when the framing is done by the person who has the problem. Benchmarks report the first and stay silent about the second.

A silent variable is not necessarily a stable one, and this is the part of the argument worth sitting with. Their claim is that the two bounds can move in different directions, because the training that lifts benchmark scores is the same training that erodes a model’s willingness to admit it is missing something. The abstention research they rely on finds that reasoning fine-tuning makes models measurably worse at declining to answer, including in the domains the reasoning training was aimed at. A system that is better at producing an answer is not automatically better at recognizing that the question it received cannot be answered from what it was given.

So the upper bound rises and gets announced, the lower bound moves quietly or not at all, and the distance between them stays invisible because nobody instruments it. Meanwhile the population sitting at the lower bound is the least equipped to notice when the output is wrong, since fluent legal prose that agrees with your own theory of the case is indistinguishable from good advice if you have never seen either before.

What they actually did

The empirical section is deliberately modest, and the authors say so themselves. They call it a primer, run it on a small sample of multiple choice questions, and decline to draw substantive conclusions from it.

They degraded the inputs along two axes taken from the literature on unrepresented litigants. First came typos, inserted at rising density, in three flavors: single character deletions, transpositions, and slips to an adjacent key. Then came dilution, first by wrapping the question inside blocks of irrelevant filler prose, and then by breaking the question apart and threading filler between its own sentences, with the formatting flattened so that no layout cue could rescue the model.

Accuracy fell, which was expected. Two secondary observations are more interesting than the headline:

  1. The first is that the models did not degrade uniformly, and their ranking changed between the clean and the distorted runs even though all of them came from the same provider. If that holds at scale, then the model that leads on a legal benchmark is not necessarily the model you would want in the hands of someone filing alone, and no leaderboard in circulation today would tell you which is which.

  2. The second is a small human detail. One model, instructed to return nothing but a letter, began several of its answers by remarking that it had probably encountered a typo. It noticed the noise. It simply had no framework for treating that noise as information about the person who produced it.

The objection from the civil law side, and why it does not hold

The natural European reaction is that this is an American problem. Litigation without counsel on the American scale does not exist here, and standing personally before a judge is confined to the smallest claims.

That reaction is right about courtrooms and wrong about everything else, because Europe did not address the representation problem by opening the courtroom. It addressed it by building an adjudicative layer outside the courtroom, where technical defense is optional and frequently absent, running from banking and financial arbitration through consumer conciliation to simplified small claims procedures.

Those forums are where the argument bites hardest, and the reason is structural rather than statistical. An American plaintiff without a lawyer files into a system that has slack in it, since there are hearings, a judge who can ask a question from the bench, an opponent with every incentive to point at the gap, and an appeal. A mistake has several chances to surface before it becomes final.

A written, documentary proceeding has none of that slack. The record closes, the panel reads what was submitted, and it decides. The contraddittorio happens on paper. Nobody in the process ever asks the question that would have surfaced the missing fact, because the design contains no moment at which such a question gets asked.

The continental version of this problem is therefore the concentrated one.

The bench test

Three things follow for anyone building or practicing.

  1. If you build client-facing intake, measure your own lower bound before somebody else does it for you. Take matters you won, reconstruct how the client first described the problem on the initial call, and run your system against that version rather than against the tidy summary that ended up in the file. The distance between that result and your demo is the honest description of your product.

  2. If you take over a matter a client has already started on his own, audit for absence before you audit for error. Assume the submission was shaped by a system that never asked a clarifying question, and that the decisive fact is missing because nobody prompted for it. Reading what is on the page is the easy half of that job.

  3. And if you are choosing models for this kind of work, stop treating benchmark rank as transferable. The question that matters is how gracefully a model degrades when the question arrives dirty, and that is a measurement you will have to run yourself, because nobody is publishing it.

The man in the vignette gets his decision on the papers in a few months. Whether it goes his way turns less on what the panel thinks about his rights than on whether sentence eleven reached the file in a form that somebody could see.


Source: Andrew Lou and David Shin, Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice, arXiv:2606.23716, accepted to the AI4Law Workshop at ICML 2026.

Torna alle news