Your Legal AI Probably Already Wrote the Right Answer
There is a quiet, uncomfortable finding buried in most benchmark tables, and a paper out of Stanford, Berkeley, and NVIDIA just made it impossible to ignore. When you let a strong model attempt a hard task five times, one of those five attempts is usually correct. On the coding benchmark the authors use, an oracle that always picks the best of the pool solves nearly the entire leaderboard, reaching 98.9 percent, while any single attempt lands far lower. The capability is already sitting inside the model. What is missing is the judgment to recognize the good answer when it appears.
That gap has a name in the paper, and it is one that everyone shipping legal AI should sit with. The authors call verification a scaling axis, on par with pretraining and test-time compute, and they argue we have spent almost all our effort on generation and almost none on the thing that decides whether generation is worth anything. If you run a legal research agent, a redline agent, or a memo drafter that samples a few candidate outputs and returns one, your product quality is capped not by how well the model writes but by how well something downstream selects. Most teams have never measured that selector. This paper is a good reason to start.
What the paper actually does
The mechanism is almost embarrassingly simple, which is usually a sign it will spread. A standard LLM judge is asked to score a candidate on a scale, and it emits a single token, say a 4 out of 5. That token throws away everything the model was actually thinking. Underneath, the model held a probability distribution across all the score tokens, a little cloud of belief that this answer is probably a 4, maybe a 5, with a whisper of 3. The judge collapses that cloud into one integer and moves on.
LLM-as-a-Verifier keeps the cloud. Instead of reading the emitted digit, it reads the model’s log-probabilities over every score token and computes the expected value, turning a coarse integer into a smooth, continuous score. Nothing new is trained. The verifier is a general model, in most experiments Gemini 2.5 Flash, prompted to rate a trajectory and then mined for its logits rather than its words.
From that one move, three dials appear. The first is granularity: widen the score range from five levels to twenty and the continuous estimate gets finer, lifting pairwise verification accuracy on the coding benchmark from 73.1 to 77.5 percent. The second is repetition: average several independent passes and the noise falls, so much so that a single-pass continuous verifier already matches a discrete judge that was ensembled sixteen times. The third is decomposition: replace one vague question (”is this trajectory correct?”) with a handful of narrower ones, checking specification, output format, and error signals separately, then average, which pushes accuracy to 78.3 percent.
The result that will get quoted is the tie rate. A discrete judge scoring hard coding trajectories calls it a tie 26.7 percent of the time, shrugging on more than a quarter of the comparisons that matter most. The continuous verifier produces zero ties. The clearest illustration is a single SQL optimization task the authors dissect. Two candidate solutions look almost identical, but one validates its optimized query against a tampered, indexed copy of the database rather than the real one, which is a subtle methodological cheat. A discrete one-to-five judge ties the two candidates in 88 of 100 runs. Taking the expectation over the same scale breaks every tie and ranks the correct one higher 69 times. Widening the scale to twenty pushes that to 77.
Stacked together and wrapped in a budget-aware ranking tournament, the framework sets new marks across coding, robotics, and medical agent tasks, and it does so with no fine-tuning at all. There is a second, quieter contribution that legal teams should care about more than the headline scores. The same verifier score, tracked step by step through a task, rises steadily on runs that are going well and stays flat on runs that are drifting toward failure. That gives you a live progress meter and an early-warning light for a long-running agent, which the authors ship as an extension for Claude Code and Codex.
Where it is genuinely strong
The training-free part is not a footnote, it is the whole commercial argument. Nobody has to curate a labeled dataset, stand up a reward-model training run, or babysit a fine-tune that goes stale the moment the base model updates. You prompt a capable model, read its logits, and you have a verifier that transfers across domains it never saw. For a small legal team without a machine-learning function, that is the difference between a research idea and something you could actually wire into a pipeline this quarter.
The calibration is real, and calibration is what legal buyers are quietly starving for. A continuous score that reliably separates a stronger answer from a weaker one, and that can be turned into an honest preference probability, is far more useful than a judge that keeps declaring five-way ties. And the progress signal opens a door that pure scoring does not, because a monitor that can flag a drifting agent mid-task is exactly the kind of control a supervising lawyer needs before an agent commits a bad edit to a document.
Where it strains, and why lawyers should notice
Now the part the abstract does not advertise. Every benchmark in this paper has a checkable ground truth. The hidden test suite either passes or it does not. The robot either grasps the object or it does not. The patient record either gets retrieved or it does not. The verifier is impressive precisely because there is a fact of the matter it is being measured against, even when that fact is expensive to compute. That single assumption is where the method meets the wall that runs through most of legal work.
There is no hidden test suite for “this is the stronger argument on abstention from a restructuring vote,” and there is no unit test that turns green when a brief is persuasive. A great deal of legal judgment is genuinely contested, and on genuinely contested questions a confident continuous score is not a gift, it is a hazard, because it dresses a matter of judgment in the costume of a measurement. A discrete judge that hedges with a tie is at least being honest about its uncertainty. A verifier that returns a crisp 14.7 over a question with no ground truth invites a lawyer to trust a number that means nothing.
Three more constraints matter before anyone gets excited. The method needs access to the verifier’s token log-probabilities, which several closed frontier APIs, including some of the models legal teams actually deploy, do not expose. The authors offer a two-stage workaround that routes a closed model’s reasoning through an open, logit-accessible verifier and recovers most of the gain, but that adds plumbing and a second model to your bill. The criteria decomposition, which does a lot of the heavy lifting, is hand-designed per domain rather than learned, so someone has to sit down and write the legal rubric. And the whole approach assumes you can cheaply generate several candidates per task, which multiplies your generation cost before the verifier ever runs.
The legal desk test
Strip away the benchmark glamour and ask the only question that matters for a working practice: would this survive contact with a real legal desk? Run any candidate use case through four gates.
Is there a checkable notion of correct? If correctness is contested judgment, the verifier gives you false confidence. If it is a fact, you are in business.
Can you generate several candidates cheaply? Selection needs a pool. One draft, nothing to verify.
Do you have logit access, or tolerance for a two-model workaround? No logits, no continuous score.
Does the task split into verifiable sub-criteria? Decomposition is where most of the accuracy lives.
Score a few real tasks against those gates. Verifying that every authority cited in a memo actually exists and says what the memo claims passes cleanly: correctness is checkable, sub-criteria are obvious, and a selector that ranks the citation-clean draft above the citation-sloppy one is immediately valuable. Choosing the best clause extraction or the safest of several redlines mostly passes, because format and rule compliance are concrete even when style is not. Deciding which of five research memos is strongest fails the first gate outright, because there is no ground truth to calibrate against. Ranking litigation strategies fails for the same reason, and fails hardest, because the confident number is most seductive exactly where it is least earned.
The honest verdict is that LLM-as-a-Verifier is a strong selector and monitor for the checkable slices of legal work, and a category error for the contested ones. The teams that win with it will be the ones disciplined enough to point it only at the parts of the practice where “correct” is a fact rather than an argument. That discipline, and not the logits, is the hard part.
Sources
Kwok et al., LLM-as-a-Verifier: A General-Purpose Verification Framework: https://arxiv.org/abs/2607.05391