Fail Closed

Google Research just published something that has nothing to do with law, which is exactly why every managing partner should read it. The paper describes the Paper Assistant Tool, or PAT, a system that reads a full scientific manuscript before submission and reports back on it. It checks proofs, flags logical errors, tests whether the experiments actually support the claims, and points out where the work is thin. The tool was piloted at two of the most demanding venues in computer science, STOC and ICML, and between them it reviewed more than 4,700 submissions.

I read it as a litigator who spends his days building AI for a law firm, and I recognized the problem on the first page. It is our problem, wearing a lab coat.

man wearing blue long-sleeved shirt standing near wooden table
Photo by Jouwen Wang on Unsplash

The bottleneck is the same shape

Science has a scaling crisis that the paper documents plainly. Across three flagship AI conferences, combined submissions went from roughly 17,000 in 2020 to a projected 74,000 in 2026. Expert review does not scale with that curve, because there are only so many qualified reviewers and each careful read of a dense proof can take days.

A law firm has the identical structure. The senior lawyer is the scarcest resource in the building, and too much of that scarce time goes to the least strategic task there is, which is cleaning up a junior’s draft before the real thinking can begin. The constraint is not the volume of work coming in the door, and it is not the talent of the associates. It is the review capacity of the few people qualified to review. Same shape, different building.

What PAT actually does, and why it is not a chatbot

The interesting part is architectural, and it is worth understanding before drawing any lesson from it. The obvious approach, a single model call on the whole paper, runs into context limits, because verifying dense proofs burns through more thinking tokens than a model can hold at once. The next obvious approach, calling the model many times and pooling the results, raises recall but destroys precision, so that a human ends up sifting through a hundred proposed issues to find the one that is real.

PAT is built to avoid both traps. A segmenter breaks the paper into themed sections, and an adaptive budget spends more compute on the proofs than on the introduction. Specialized review agents then verify each segment while holding the full paper as context, and a synthesis agent deduplicates the critiques and, this is the part to underline, grounds the output against search to strip out hallucinated references before assembling the final report.

That last stage is the whole idea. The system is engineered around the assumption that the model will invent things, and it dedicates an entire step to catching the inventions before they reach a person.

The number everyone will quote, and the number that matters

On a set of papers that were later retracted for genuine mathematical errors, PAT caught 89.7% of them, against 55.2% for a single strong model call. That is a jump of roughly 34 points, and it is the number that will travel through every legal-tech deck for the next year.

But the number that matters is the one the authors were honest enough to publish about their own tool. When they asked the STOC and ICML authors whether the feedback was grounded, only 55.8% and 64.8% respectively said it was mostly or entirely grounded. Read that again. Between a third and nearly half of the people using a state-of-the-art review system, built by Google and powered by its strongest reasoning model, found that a meaningful slice of the critiques were not anchored to anything real.

In science, a hallucinated critique costs an author an hour of chasing down a phantom objection. It is annoying, and it is not fatal.

The asymmetry that changes everything for law

Law does not have that luxury, because the cost of a fabricated output is not symmetric with its benefit. A review tool that invents a helpful improvement saves a little time.

A review tool that invents a controlling precedent, or flags a procedural defect that does not exist, or misreads a clause and confidently tells the junior to rewrite an argument that was already correct, does not cost an hour. It can cost the case.

And courts have already sanctioned lawyers for briefs built on citations that a chatbot invented and nobody verified, which is the same failure PAT was engineered to prevent, transplanted into a courtroom where the stakes are a client’s rights rather than a reviewer’s afternoon.

This is why, in legal AI, the model is not the moat. Any firm can license the same frontier model that I can. The moat is the verification layer, meaning the part of the system that decides what is allowed to reach a human at all, and refuses to surface any critique it cannot anchor to a verified source. It is deterministic checking sitting on top of probabilistic generation.

Engineers have a phrase for this, and legal-tech should borrow it: fail closed. When a secure system loses confidence, it denies access rather than granting it. A legal review agent has to behave the same way. If it cannot ground a citation to a real, retrievable authority, it does not hedge and it does not soften, it says nothing. A missed issue is a cost the firm can absorb, because the senior would have caught it anyway. A fabricated issue is a heavier one, because it inverts the reviewer’s job, shifting the partner’s attention from finding problems in the draft to disproving problems that the tool invented. Fail closed, not open. That single design decision is where legal AI will be won or lost.

The taxonomy, read for a law firm

The paper offers a taxonomy of four roles for AI in review, deliberately modeled on the SAE levels of vehicle autonomy, and it maps almost cleanly onto the adoption path of a firm.

Role 1 is the tool for authors, where the junior runs the agent on the draft before it goes up the chain and stays fully responsible for the work. This is where PAT lives today, and where most firms should begin.

Role 2 is the tool for reviewers, where the senior uses the same agent to accelerate their own read but signs off personally.

Role 3 is the supporting reviewer, where the agent produces a genuine first-pass review and the human shifts from writing the review to deciding on it, the way an area chair decides rather than reads. Call it the synthetic junior, and note that it is technically within reach.

Role 4 is full automation, the paper’s imagined “AIrXiv,” a repository where AI vets everything and the human is optional.

Here the analogy breaks, and the break is the whole point. In science, Role 4 is at least imaginable. In law it is not, and not because the technology is missing. It is structurally off limits, because accountability in legal work cannot be delegated to a system that cannot be admitted to the bar, cannot be sanctioned, and cannot owe a client a duty of loyalty.

A principle is already taking shape in administrative law (in Italy), a reserve of human decision, holding that certain judgments must remain with a person. In our field the human in the loop is not a phase we pass through on the way to automation. It is the terminal state.

The ceiling for legal AI is a very good Role 3 sitting under a human who owns the outcome, and that is not a limitation to apologize for. It is the shape of the product.

Where it is actually won

So the lesson from a Google paper about theoretical computer science turns out to be a legal one. Stop competing on the model, because everyone has the same model. Compete on the gate. Build the verification layer that fails closed, tune it on the one reward signal that only real litigation can provide, which is the accumulated record of which objections actually held and which citations actually controlled, and put it in front of a partner who signs. The generation is a commodity. The verification is the practice.

I am building exactly this layer, and Google’s paper just told me, in its own honest numbers, precisely which part is hard.


Reference

Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, and Vincent Cohen-Addad, Towards Automating Scientific Review with Google’s Paper Assistant Tool, arXiv:2606.28277 (2026). https://arxiv.org/abs/2606.28277

Torna alle news