The Narrowing: What an AI Ideation Study Reveals About Legal AI’s Favorite Move

I spend a lot of my time watching legal AI systems generate first drafts: theories of a case, defenses, settlement angles, ways to frame a novel argument to a court that has never seen it. So when a paper lands that measures, at scale, what kinds of ideas language models actually produce, I read it as if it were about my own tools, because in every way that matters it is.

The paper is called Measuring the Gap Between Human and LLM Research Ideas, from a team at Chicago and Yale. It studies scientific ideation, not legal drafting, and I want to be honest about that distance from the first paragraph. But its central result is not a fact about physics papers. It is a fact about how these models behave when you ask them to find an opening and build something on it. That behavior does not stay politely inside the natural sciences.

Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.

library shelf near black wooden ladder
Photo by enrico bet on Unsplash

What they actually did

The setup is clean, which is why the result is hard to wave away. The authors took roughly 11,700 real research papers, split evenly between machine learning venues and Nature Communications, and for each one reverse-engineered the small set of prior works that plausibly preceded the paper’s core idea. They then handed those same prior works to nine current models, from Claude and Gemini to the GPT, Qwen, and DeepSeek families, and asked each to produce a new idea from the same starting point. The real paper is the human answer. The model output is the machine answer, generated from the same local context a human researcher would have had.

Then they classified every idea along two axes. One axis asks why the work is worth doing: is the gap a contradiction, a missing explanation, a scope problem, an evidence gap, a disconnect between literatures, a failure or risk, or a resource bottleneck. The other axis asks how the contribution is built: synthesis, scope extension, robustification, formal derivation, empirical mapping, artifact building, or optimization. That taxonomy was not invented casually. It was assembled from research guidance published by NSF, NIH, AHRQ, and DARPA, then refined on a held-out set of papers.

The comparison is distributional, and this is the part I find genuinely useful. They are not asking whether a single idea is good. They are asking what the whole population of ideas from a given source looks like, and whether that population is as wide as the human one.

The finding, stated plainly

Model ideas are narrower than human ideas, and they are narrow in a specific, repeatable direction.

Only 12.1 percent of human ideas frame the opportunity as connecting disconnected work, and only 5.1 percent build the contribution through explicit synthesis or unification. Across the main models, those same numbers run from 47.1 to 64.2 percent on the framing axis, and from 22.5 to 38.7 percent on the method axis. Human entropy across the categories sits above 0.92 on both axes, close to a flat spread. The models cluster. Even the closest model to the human distribution still needs roughly a third of its idea mass relocated to match it.

Underneath the taxonomy, the mechanism is almost blunt. When the authors reduced each proposal to a single verb, the dominant model move was “integrate”: it showed up 7,994 times in model outputs against 275 times in human ones. Humans, by contrast, reached far more often for verbs like “replace,” “decouple,” and “formalize.” Replace alone accounts for 9.13 percent of human moves and 0.92 percent of model moves. The human researcher tends to change one specific thing. The model tends to bolt two existing things together.

There is one more result that I keep returning to. Turning on extended reasoning, the “thinking” mode everyone assumes makes models more careful and more original, moved the output distribution further from humans, not closer. For one model, enabling thinking pushed bridge-style framing from 49.7 to 71.1 percent and synthesis from 38.7 to 52.2 percent, while the diversity of ideas dropped. Reasoning sharpened the template. It did not break it.

Why a banking-litigation practitioner should care

Read “research idea” as “legal theory” and the paper starts describing my daily problem.

The synthesis move is the safe move, and it is exactly the move that loses cases. In a non-performing loan dispute, in a guarantee enforcement action, in a derivatives claim turning on jurisdiction and public policy, the winning argument is rarely “combine these two established doctrines into something that sounds interdisciplinary.” It is usually the local intervention: replace a brittle premise the other side is relying on, decouple two issues the court has been treating as one, formalize a structure everyone has left implicit. That “replace, decouple, formalize” trio the paper found on the human side is a fair description of good lawyering. The trio the models overproduce, “integrate and unify,” is a fair description of the memo that reads well and persuades no one.

The reasoning result cuts the same way. There is a comfortable assumption in legal tech that longer chains of thought yield sharper legal analysis. This paper is at least a warning against treating that as automatic. If more deliberation collapses the model toward its favorite template, then a tool that “reasons harder” about a fact pattern may simply produce a more confident version of the most generic argument available. Confidence is not the scarce resource in litigation. Specificity is.

There is also a quieter finding that supports something I have argued for a while: verification is a moat, and specificity is measurable. The study’s annotator scored each proposal for how superficial the combination was, how precisely it named the actual bottleneck, and how boilerplate it read. Most models scored worse than humans on precision and generic phrasing, and one model family was a partial exception, scoring slightly better than the human baseline on specificity while still landing in the wrong overall distribution. That is the whole game in a sentence. An output can be polished, specific, and still be reaching for the wrong kind of idea. If your evaluation only checks whether a single answer is coherent, you will never see it.

Finally, the method itself is portable. Legal AI evaluation today mostly grades outputs one at a time: is this citation real, is this summary accurate, is this clause enforceable. Almost nobody grades the distribution. Ask your drafting tool for fifty theories of the case across fifty matters and look at the shape of what comes back. If forty of them are “harmonize these two lines of authority,” you have a homogenization problem that no single-output review will catch.

Where I would push back

I would not run this paper into a courtroom, and neither should you.

The corpus is entirely STEM. There is not a single legal text in it. Everything I wrote in the previous section is an argued transfer, not a demonstrated one. Legal reasoning has its own structure, and it is plausible that the human-model gap looks different, larger or smaller, once you are working with statutes, precedent, and the peculiar rhetoric of persuasion. The paper cannot tell us, and it does not claim to.

The evaluation is also an LLM grading LLMs. The classifier that assigned every one of those 11,700-plus labels is itself a language model, checked against two human annotators on 150 items. The agreement scores were high, in the 0.81 to 0.93 range, which is real. But there is an unavoidable circularity in using one model to characterize the taste of others, and I would want more human adjudication before treating the exact percentages as settled.

Then there is survivorship. The “human ideas” are drawn from published papers. Published work is the filtered tail of human ideation, the part that cleared peer review precisely because it was not another routine synthesis. Researchers produce enormous quantities of derivative, combine-two-things ideas that never reach a journal. So the human distribution here is not “what humans think.” It is “what humans publish,” which is exactly the sample selected to look less template-bound. The gap is real, but part of its size is an artifact of comparing raw model output against humanity’s edited highlight reel.

And the setting is one-shot and non-interactive. No back-and-forth, no retrieval pipeline, no domain-specific system, no lawyer in the loop pushing the model off its first instinct. The paper says as much in its limitations, and it matters, because the legal tools worth building are agentic and iterative. It is entirely possible that the narrowing the authors measured is most severe precisely in the setting they tested, and that a well-designed harness recovers some of the lost range.

What I am taking from it

Three things go into my practice notes. First, treat the model’s first idea as its most generic idea, and design the workflow to fight the synthesis reflex rather than reward it. Second, do not assume that a reasoning mode buys you originality; test whether it is buying you a more elaborate version of the same template. Third, start measuring the distribution of what your tools produce, not just the quality of each output, because homogenization is invisible one answer at a time.

The paper’s real contribution is to reframe the question. The interesting failure of legal AI will not be the hallucinated citation, which we already know how to catch. It will be the quiet convergence on a single, safe, respectable kind of argument, delivered with enough polish that nobody notices the range has collapsed.

Which raises the question I would put to anyone shipping a legal drafting system: are you evaluating whether each output is good, or whether the set of outputs is still wide enough to win?

Reference

Ziyu Chen, Yilun Zhao, Arman Cohan. Measuring the Gap Between Human and LLM Research Ideas. University of Chicago and Yale University. https://arxiv.org/abs/2607.01233 [cs.CL], July 2026. Code and data: github.com/ziyuuc/TasteGap and github.com/IdeaLand/IdeaSeed.

Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.

Back to news