The Committee, the Courtroom, and the Parrots

Almost every demo I have seen for agentic legal AI rests on the same unspoken promise, which is that more agents produce better answers. Give the model a team instead of a single voice, let those voices debate, and the reasoning is supposed to come out sharper. It is an intuitive idea, and as someone who builds these systems I want it to be true. A new paper accepted at the AIDA2J workshop of the International Conference on Artificial Intelligence and Law, held in Singapore this June, put that promise on a bench and tested it properly. The result is more interesting than a clean win or a clean loss, and it is worth sitting with if you are betting your practice, or your product, on multi-agent legal AI.

The authors, Ludi van Leeuwen and Cor Steging of the University of Groningen together with Tadeusz Zbiegień of the Jagiellonian University, did something that maps almost exactly onto how litigators think. Rather than asking a model to reason in a single pass, they built architectures in which distinct agents take distinct positions and a structure forces those positions to meet. Because this is precisely the design philosophy behind the adversarial systems I work on, I read the paper twice, and I want to walk you through what it actually shows rather than what a headline would want it to show.

What they built

The paper compares a plain single-model baseline against three multi-agent designs, and the three designs are worth naming because each encodes a different theory of how good reasoning emerges.

  1. The first is standard Multi-Agent Deliberation, which the authors treat as a committee. Three agents each produce an answer, then read each other’s answers across two rounds and revise, and a majority vote settles the question. Reaching a single verdict this way costs nine model calls rather than one.

  2. The second design is the one that should interest any litigator, because it is a courtroom. The authors call it 3-Ply, and it assigns one agent to argue yes as the plaintiff, one to argue no as the defendant, and a third to sit as an impartial judge who decides on the merits after an opening argument, a counterargument, and a rebuttal. If you have ever sketched a Proponent, an Adversary, and an Arbiter on a whiteboard, you already understand this architecture, since it is the adversarial method formalized into a pipeline of four model calls.

  3. The third design is the strangest and, to me, the most thought-provoking. Drawing on recent argumentation research, the authors stage a single expert agent named Alex against a chorus of four critical “parrots,” each embodying a distinct stance. A Socratic parrot challenges definitions and assumptions, a Cynical parrot tries to undermine the arguments and test their robustness, an Eclectic parrot offers interpretations everyone else missed, and an Aristotelian parrot audits the logic for fallacies. Alex answers, the parrots push back, and Alex is allowed to continue the exchange for up to three rounds before committing. Because Alex decides when the conversation is over, this framework uses a variable number of calls, averaging about 3.48 per question.

A person adjusts another's judicial robe and wig.
Photo by Dwayne joe on Unsplash

To test these designs the researchers assembled five benchmarks, four of them legal and one purely logical. The legal set spans law school examinations, Japanese civil law bar questions, United States federal tax reasoning, and privacy policy interpretation, and the fifth benchmark tests logical reasoning drawn from a civil service exam. They sampled 250 balanced questions from each, which gives 1,250 binary yes-or-no problems in total, and they ran everything on a smaller commercial model, GPT-5-mini, with two-shot prompting where the examples were selected by a standard retrieval method. The choice of a smaller model was a budget decision, and the authors are candid that larger models might behave differently, which is a caveat I will return to.

The finding that should reshape your priors

Here is the headline, stated plainly, because the size of a claim should match the size of its evidence. Across all five benchmarks, none of the three multi-agent frameworks meaningfully outperformed the single-model baseline. The average F1 scores sat within one and a half points of each other, and no architecture pulled clearly ahead. If your entire thesis for multi-agent legal AI is that it produces higher accuracy, this paper does not support you.

And yet stopping there would miss the real result. Although the frameworks did not score higher, they answered differently in a way that is statistically unmistakable, since every multi-agent design diverged from the baseline’s decision pattern at high significance. Roughly seven to ten percent of questions received a different answer under a multi-agent framework than under the single model, and in about half of those disagreements the multi-agent system was right where the baseline was wrong. Depending on the framework, between forty-three and fifty-two percent of the cases where they diverged were cases the multi-agent design rescued.

The most striking detail hides in the privacy dataset, where the authors found an asymmetry that I have not been able to stop thinking about.

Every question the baseline solved correctly was also solved by at least one multi-agent framework, but the reverse was not true, because many questions the baseline failed were recovered by the agents.

In that slice of the data, the multi-agent approach did not trade one set of errors for another of equal size. It strictly expanded what the system could handle. The authors are careful to note this reflects the combined behavior of several frameworks rather than a single one, so it is a hint rather than a law, but it is exactly the kind of hint that tells you where to dig.

This is the reframing that matters for practitioners. The question is not “does multi-agent beat a single model on average,” because on these benchmarks it does not. The better question is “does multi-agent reach reasoning that a single model cannot,” and here the evidence says yes, at least for a specific and identifiable class of problems.

Where the agents earn their keep, and where they do not

The qualitative analysis is where the paper stops being a scoreboard and starts being useful. The authors show that the multi-agent frameworks pull ahead precisely when a legal clause is ambiguous or admits more than one reading, which is to say in the situations lawyers are actually paid to resolve. Their worked example turns on whether a privacy clause covering “geolocation data” also covers WiFi-derived location that the clause never names explicitly. The single model latched onto the literal absence of the word WiFi and answered wrong, whereas the courtroom and the parrots surfaced the tension between a literal and a purposive reading, argued it out, and arrived at the correct, more lawyerly interpretation. This is the deliberative dividend, and it appears exactly where a single narrative is most likely to suffer tunnel vision.

The uncomfortable findings deserve equal airtime, because a newsletter that only reported the flattering half would be the marketing I am trying to avoid. Hallucination did not go away. In the same example, the committee framework reached the right answer only after one of its agents invented a justification about cell phone data that had nothing to do with the question, which means the structure had to first clean up a mess it created. The authors cite the argument that hallucination is a structural feature of these models rather than a bug we will patch, and their own results are consistent with that sobering view.

Two more results puncture common assumptions. First, adding reflection rounds to the committee barely moved the score, since three agents voting once landed within a point of the same agents deliberating twice, which suggests that on these tasks much of the value came from ensembling rather than from the deliberation itself. Second, all of this costs real money, because the committee burns nine model calls and the courtroom four for every single call the baseline uses, and that spend did not buy better raw accuracy. If you are deploying at scale, you are paying a multiple for a different distribution of errors, not for fewer of them.

What I take from this

I build adversarial and multi-role systems for legal work, so I read this paper as a colleague, not a spectator, and my honest reading is that it validates the architecture while demolishing the marketing. The value of putting a plaintiff, a defendant, and a judge inside the machine is not that it is smarter in aggregate. The value is that it makes the reasoning explicit, exposes competing interpretations that a single pass would bury, and rescues a meaningful set of hard, ambiguous cases that a monolithic model gets confidently wrong. For a domain where the reasoning behind an answer matters as much as the answer, that explicitness is not a side effect. It is the product.

There is a deeper point underneath the numbers, and the paper gestures at it in its discussion. These systems generate arguments, but nobody checks whether those arguments are actually valid, and a model can produce a fluent, well-structured, and logically broken chain of reasoning without noticing. Deliberation makes the reasoning visible, and yet visibility is not verification. The agents argue, but the argument is never audited for soundness. That gap, between a system that produces arguments and a system that can be trusted to check them, is where I think the serious work of the next few years lives, and it is a gap that neither more agents nor more rounds will close on their own.

If you are evaluating agentic legal AI, the practical takeaway is to stop asking vendors whether their multi-agent system scores higher, and start asking which problems it solves that a single model cannot, and at what cost per answer. Those are the questions this paper actually equips you to ask.

I plan to revisit the primary source as more replications come out, and I would rather this be a conversation than a broadcast. The paper is titled Investigating Multi-Agent Deliberation in Law, and it is on arXiv at 2606.30906. If you build or buy these systems, I want to hear where your experience confirms or contradicts what the authors found, so tell me in the comments or reply to this email.

What have you seen multi-agent reasoning do that a single model could not?

Back to news