The Harness Is the Part You Can Actually Build

I’ve just written on LinkedIn about a survey by the University of Illinois, Meta, and Stanford arguing that the reliability of an AI agent depends less on the model and more on the “harness” around it. That post made the case to buyers. This one is for the people on the other side of the table: the lawyers who are quietly building their own tools, wiring an LLM into a retrieval system, a handful of prompts, and a review step, then hoping the result holds up.

The encouraging part of the paper is that the harness is the piece you actually control. You are not going to out-train a frontier lab, but you can out-engineer a competitor on permissions, verification, and record-keeping, which happens to be where legal work lives. Here are the components worth taking, and the one that will bite you if you copy it without thinking.

Ship an evidence bundle with every output

The most useful idea in the paper is almost thrown away: the authors argue that verification should stop being a single pass/fail flag and become a stack, where each check states plainly what it verifies, what it cannot verify, and how confident it is. Then they take the step most legal tools skip: every accepted action should carry an “evidence bundle” recording the checks that ran, the assumptions it preserved, the parts it left untested, and the risks that remain.

Read that as a legal deliverable and it is close to a work-product cover memo. Picture your citation checker returning not “passed” but a short bundle: citations resolved against the official source, quotes matched to the paragraph, holdings not independently verified, jurisdiction assumed to be the one in the caption, three passages flagged as low confidence. That bundle is the line between an output you can defend and an output you still have to re-check by hand. Build the bundle first and the model can improve later.

Retrieve by legal structure, not by token window

A recurring finding in the retrieval sections is that code tools improved not by pulling in more text but by pulling in text aligned to the structure of the code, using structure-aware chunking and query rewriting rather than fixed-size windows.

Most legal RAG systems get this backwards. They slice a contract or a judgment into 800-token blocks and then wonder why the model cites half a clause. Chunk by the units that carry legal meaning instead: the article, the clause, the recital, the holding, the operative paragraph. Let the retriever rewrite the query as it narrows in, the way a junior does when the first search misses. The paper’s memory work adds a second warning that applies directly here: a curated, governed store of past matters beats a large one, because ungoverned history mostly contributes noise and confident wrong retrievals, so more precedent in the index is not more signal.

Give the tool a permission ladder, and a plan it has to sign

The paper’s control model is a clean three-rung ladder. The bottom rung reads and inspects. The middle rung edits and runs things inside a sandbox. The top rung touches the outside world: sending, filing, paying, deleting. Only the top rung requires a human gate, and the important nuance is that the risk of an action depends not on the tool but on the arguments, the data, and the side effects. The same “send” is trivial in a draft and irreversible in a court filing.

Before any of that, the paper treats the plan itself as a contract. A good plan names the files it will touch, the invariants it must preserve, the checks that will confirm success, and the rollback path if it fails. For a legal tool, that is a scope of work the agent commits to before it acts, and a natural place to require sign-off. If your tool cannot state, in advance, what it is about to change and how to undo it, it is not ready to touch anything that leaves the building.

Stop pasting whole documents into the prompt

One quiet technique runs through the whole paper:

keep the full material outside the model, and pass the model a compact, provenance-preserving summary plus a handle back to the original. A failing report becomes a few key lines and a link, not the entire log.

Legal builders do the opposite by reflex: they paste the whole contract, the full brief, the complete deposition into the context and trust the model to find the needle. The research is blunt about why that fails: attention degrades across long inputs, and the middle gets lost. Offload the documents, retrieve the relevant spans, and keep a pointer to the source for anything the model leans on. Your outputs get more accurate, and, as a side effect, more auditable.

Where this breaks for law

Now the critique, and I make it only because the paper half-makes it itself.

The entire paradigm draws its confidence from execution. A code harness can compile the program, run the tests, and read the crash. The authors are careful to scope “code” to things that are machine-checkable, and they say plainly that human intent and judgment are not code. They even warn that execution feedback can create a false sense of correctness, and that a system can grow overconfident precisely because it has a green test to point at.

Hold that next to legal work and the problem is obviou:

Law has almost no execution oracle.

There is no unit test for whether an argument is persuasive, whether a clause allocates risk the way the client actually intended, or whether a filing meets a standard a judge will apply next year. The harness gives you strong guarantees exactly where law is weakest, and near-silence where law actually decides things. Take the scaffolding, but do not let “verifiable” quietly become “correct,” in your own head or in your marketing.

There is a sharper version of this, and it comes from setting two of the paper’s own recommendations next to each other. The paper says to compact evidence into short summaries before showing it to a human. It also says that high-stakes actions should pause for human approval, logged as an auditable record of who signed off and on what basis. Put those together and you get a system that hands a busy partner a tidy three-line summary, records the click as informed authorization, and produces an audit trail that looks flawless. The trail is real and the judgment behind it may be a rubber stamp. In our world, that is liability laundering with good UX.

So if you build the approval gate, resist the urge to over-compress what the approver sees on the actions that carry real weight. An audit trail of uninformed sign-offs is worse than no trail at all, because it manufactures the appearance of the diligence you skipped.

The short version

If you are building, the checklist is small and unglamorous.

  1. Attach an evidence bundle to every output.

  2. Chunk and retrieve by legal structure, and keep your experience store curated rather than large.

  3. Put the tool on a permission ladder, and make it sign a plan before it acts.

  4. Keep documents out of the prompt and pass handles instead.

  5. And on the actions that matter, make human approval slow, informed, and honestly recorded, not fast, compacted, and defensible-looking.

The model is the part everyone is watching but the harness is the part that will decide whether your tool survives its first bad day.


Source: “Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems,” Ning, Tieu, Fu, et al. (University of Illinois Urbana-Champaign, Meta, Stanford, 2026). https://arxiv.org/abs/2605.18747

Share

Torna alle news