Law Already Runs a Red Queen Race

Every general counsel I speak with eventually asks a version of the same question. Can a legal AI actually get better on its own, given that the law so rarely offers one provably correct answer? It is a fair question, and until recently the honest answer was a shrug. A new preprint from a team at the University of Cambridge and NVIDIA, working with Flower Labs, MBZUAI, and Inria, gives me a better one. The paper is “The Red Queen Gödel Machine,” and although it never mentions law, it is one of the most useful things I have read this year for anyone building AI that has to reason about it.

Let me give you the practitioner translation.

The assumption that breaks on contact with legal work

Most self-improving agents share a hidden premise. They optimize against a fixed scorer, whether that is a benchmark, a unit test, or a labeled dataset that never moves. The agent proposes changes to itself, keeps the variants that raise the score, and discards the rest. This works beautifully for code, because a test either passes or it does not, and the target holds still while the agent chases it.

That premise quietly fails the moment the task starts to resemble legal work. The standard for a good argument is not fixed. It sharpens as opposing counsel gets sharper and as the bench pushes back.

Advocacy is adversarial by construction, which means you are never optimizing against a static environment. You are optimizing against something that adapts to you.

The authors name their system after Leigh Van Valen’s 1973 Red Queen hypothesis, the idea that a species has to keep evolving simply to hold its ground against competitors who are evolving too. Read that description again and tell me it is not a fair account of litigation.

So the interesting move in this paper is not that the agent improves. It is that the thing judging the agent improves with it, while a much older idea sits underneath. A Gödel Machine, in Schmidhuber’s original sense, rewrites itself whenever it can prove the change is beneficial. This work keeps the self-improvement and swaps the intractable proof for something empirical, then adds the piece everyone else held fixed.

Controlled utility evolution, in plain terms

Here is the mechanism, stripped of its notation. The authors let the evaluator co-evolve alongside the agent it scores, but they clamp it with one hard rule. Search is divided into epochs. Within an epoch, one evaluator is frozen and grades everything, so the target holds still long enough for the agent to make real progress. The evaluator is only allowed to change at an epoch boundary, and it can be replaced only if a challenger statistically beats the incumbent on a fixed ground-truth anchor that neither is ever allowed to drift away from. When a replacement happens, the system performs what they call selective erasure. It discards exactly the records that depended on the old evaluator, and nothing else.

The effect is a progressively stricter judge that stays tethered to reality. The agent faces a harder examiner over time, but the examiner never floats free of the ground truth. The authors show this pays off. On open-ended tasks with no clean benchmark, such as writing scientific papers and proofs, the co-evolved system beat fixed-evaluator baselines while spending fewer tokens. Even on verifiable coding, where a real test suite already exists, adding a cheap evolved code reviewer raised the held-out pass rate to 71.7 percent from the prior state of the art’s 69.9 percent, and it did so using 1.35 to 1.72 times fewer tokens, because the reviewer is queried once where a coding agent needs many turns.

The result that should worry every legal AI buyer

Two numbers stopped me.

The strongest baseline reviewer in their study accepted AI-generated papers at up to 1.91 times the rate it accepted human ones. That is self-preference bias, the well-documented tendency of a language model to favor text that looks like its own output. Then the authors introduced an adversarial objective at an epoch boundary. They took the AI-written papers the old reviewer had waved through, replayed them as a hard pool the next reviewer had to catch, and searched for a judge that was equally tough on machine and human work. The corrected reviewer held roughly 80 percent ground-truth accuracy while treating the two sources alike.

Translate that into our world. If your legal AI’s verification layer is itself a language model with a soft spot for fluent, machine-generated prose, it will rubber-stamp a brief that reads well and cites a case that does not exist. The plausible-looking hallucinated citation is not an edge case. It is the predictable output of a lenient judge that likes its own handwriting. The paper is, among other things, a worked example of how to build a verifier that refuses to do that, and of why you should not trust a verifier that has never been tested against its own blind spots.

Is law deterministic? No. Is it unanchored? Also no.

This is where I part company with the reflexive objection that legal reasoning is too open-ended to admit ground truth. Law is not deterministic. But it is not floating either. We have binding precedent, enacted statutory text, and the outcome that actually held on appeal. That is the anchor. What evolves is the harder, softer question of how a novel argument lands, which reviewer’s standard it must satisfy, and how much weight a given line of authority can bear this year.

The paper’s structure maps onto something practitioners already know but rarely formalize. Legal ground truth is itself non-stationary, yet it moves in discrete, legible steps rather than continuously. A statute is amended. A supreme court sits in plenary session and settles a split. A precedent is overruled. Those are epoch boundaries. Between them, the anchor holds still. The right architecture for legal AI is therefore neither the one that lets its notion of correctness drift with every new draft, nor the one that freezes the law as of its training cutoff. It is the one that updates the anchor only at these legible moments, revalues what depended on the old rule, and preserves everything that did not. Controlled utility evolution with selective erasure turns out to be a surprisingly good description of what a responsible legal knowledge system has to do when a court of last resort changes the law under it.

The moat is the anchor, and the anchor is maintenance

I keep telling anyone who will listen that the verification layer is the durable asset in legal AI, and this paper sharpens why. The generating agent is already commoditizing. Whatever writes the first draft this quarter will be beaten by something cheaper next quarter. The thing that decides whether an output is good, and that refuses to certify a clean-looking citation to a judgment nobody handed down, is the part that does not commoditize.

The paper also lands its own caution squarely on us, and I want to sit with it rather than skate past it. Its guarantees are epoch-local, not global. It can show that the system improved this epoch against this version of the ground truth, but it cannot promise convergence to some globally correct evaluator, and it says so plainly. It concedes, too, that an evaluator is only as good as its anchor. A weak or biased legal ground truth does not produce a cautious agent. It produces a confident, biased one, which is the worse failure because it is harder to catch.

That is the whole argument for owning your anchor rather than renting it. A trustworthy, versioned, maintained legal ground truth is not a feature you ship once. It is a discipline you sustain, and it is the asset from which the rest of the system borrows its credibility.

Local guarantees are not a disappointment here either. For anyone who has to defend a system to a regulator or a court, “we can show this improved against a fixed, auditable ground truth between two dated versions” is a far more honest claim than “it converges to correctness,” and it is the one you can actually stand behind.

I am curious whether others building in this space read the anchor the same way, or whether you think the generating layer still holds more of the value than I am giving it credit for.

The paper is The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators, by Alex Iacob, Andrej Jovanović, and William F. Shen with colleagues at the University of Cambridge and NVIDIA, alongside Flower Labs, MBZUAI, and Inria, and it is on arXiv at arxiv.org/abs/2606.26294. I would rather this be a conversation than a broadcast, so if you build or buy legal AI, tell me in the comments or reply to this email. Which parts of your own verification stack would survive being handed to an evolving evaluator, and which anchor would you never let out of your own hands?

Back to news