Your AI Agent Just Read the Whole File Cabinet to Change One Word

Picture the most boring task in your day: a colleague asks you to swap one email icon on the firm’s homepage for a different one that already sits, ready to use, three lines above. Ten seconds of work, you do it and move on.

Now watch a capable AI agent do the same thing.

It re-reads the icon library; it re-browses the site directory; it re-analyzes the project architecture it has already seen a dozen times; it re-confirms dependencies nobody asked about. Minutes pass, then it makes the two-line change it could have made at the start, verifies it correctly, and reports success.

The edit is right but the path to it is absurd.

A team at the University of Tennessee, Knoxville built a whole paper around that absurdity, and the number they attach to it should make anyone paying by the token sit up. On the trivial icon task, their model of an over-cautious agent hits over 1000% redundancy: it spends roughly eleven times the effort the job actually needs. Not because it fails, but because it refuses to believe the job is small.

Hold that thought, because the real punchline is worse than “AI is inefficient.”

The part nobody wants on the invoice

Here is the finding that turns a productivity note into a business problem.

The waste is not spread evenly but it is concentrated on the simplest tasks.

Measure the redundancy tier by tier and it climbs as the work gets easier: highest on one-line edits, lower on cross-file changes, lowest on the genuinely hard repository-wide refactors. The agent burns the most effort exactly where effort is least warranted. It is calm and efficient on the hard problem and completely unhinged on the easy one.

Sit with that if your practice runs on AI. The routine, high-volume, low-margin work, the stuff you were counting on to be cheap, is precisely where an untuned agent quietly bleeds the most. The complex matter that justifies a real budget is where it behaves.

The authors are honest that part of this pattern is arithmetic: if the agent always reads everything, its cost stays flat while the “necessary” cost shrinks on easy tasks, so the ratio has to spike at the bottom. Fair but the money is not a ratio. The money is the files and tokens actually consumed, and those are being spent on the tasks that least deserve them. The math being partly mechanical does not make the invoice smaller.

Why the agent does this (and why it isn't dumb)

The instinct behind the behavior is not stupidity but it is fear.

Faced with any uncertainty, the agent defaults to what the paper calls maximum-context-first: gather everything, then eliminate every conceivable risk before acting. On a genuinely hard task, that caution is exactly right. On an easy one, it is a lawyer re-reading the entire contract to confirm the client’s name is spelled correctly on page one.

What is missing, the authors argue, is a skill humans use without thinking: a fast, cheap judgment of how hard this actually is before committing to a plan. An experienced associate glances at a task, sizes it in seconds, sketches the smallest move that could work, and only widens the search if that move fails. The agent never makes that first judgment. It has one gear, and the gear is “audit.”

The fix: guess small, verify, expand only if you're wrong

Their answer is a framework with a deliberately plain name: E3 (Estimate, Execute, Expand).

Estimate the task’s real scope up front, cheaply. Execute the smallest path likely to work. Expand into deeper reading and heavier checking only when verification fails. The estimate is allowed to be optimistic, even wrong, because expansion is the safety net that catches the misses. Guess lean, and pay for depth only when the evidence demands it.

The results, in their controlled benchmark, are the kind that get a paper shared:

  • Same 100% success as the most thorough baseline.

  • ~85% lower cost.

  • ~91% fewer tokens and ~92% fewer files dragged into context.

And this is not a win over a strawman. They built a genuinely strong competitor, an adaptive agent that already scales its effort to the task and solves everything, and E3 still comes out ~16% cheaper at equal accuracy. The gain that matters isn’t “think less” but it is “judge first.” E3 lands as both the leanest and the most reliable policy in the study, which is the combination that usually doesn’t come together.

They also stress-tested the cost model, the obvious place a skeptic would push. Redo the accounting under 4,000 different weightings, including ones openly hostile to their method, and E3 stays the cheapest fully-successful policy in 99.8% of them. That is a paper trying to break its own result and failing. Respect.

Now put it on a lawyer's desk

Every clean result deserves the “desk test”: does this survive contact with the real world, where the “user” is an attorney and every “just to be safe” pass is billable time?

Three things it gets right for that world.

The problem is your problem. Legal AI runs on volume, and volume is made of small tasks: pull a clause, swap a party name, reconcile one figure. If the agent is most wasteful there, the waste compounds across thousands of routine jobs. E3 aims straight at that.

“Verify, then expand” is a workflow lawyers already trust. Start narrow, check, widen only on a real signal. That is how good associates work and how good review protocols read. The shape is familiar, which makes it deployable.

And the honesty is disarming. The authors don’t oversell: they call it a controlled probe, not a verdict, and flag their own weak points before you can.

Now the reasons to keep your hand near the brake.

It lives in a simulator: most of the eye-catching numbers come from a controlled environment where every agent is equally able to make the edit, so only the reading behavior varies. Clean for science but your real agent’s capability wobbles from prompt to prompt, and that variance is switched off here by design.

The estimator is brittle where it counts. Its judgment leans on wording cues. Rephrase the tasks into unfamiliar language and its accuracy drops from about 85% to 67%, with the rate of dangerous under-scoping more than doubling. Litigation language is nothing but unfamiliar phrasing, adversarial drafting, and buried cross-references. The exact conditions that fool this estimator are the native habitat of legal text.

The real-model check is thin. They did run E3 on a live agent editing a real open-source library, graded by actually running the project’s test suite, and the effect held: the over-reading was, in their words, “milder but real”. Encouraging but it’s one model on one codebase. “Milder but real” on a single case is a hypothesis with a nice haircut, not proof it holds on your matter.

So does it survive the desk test? As an idea, yes, and it should reshape how legaltech builders think about agent cost. Stop optimizing only for “can it get the answer” and start measuring “how much did it waste getting there.” As a drop-in for tomorrow’s docket, not yet. The estimator that shrugs off a paraphrased coding task would meet its match in a poorly drafted guarantee with three defined terms pointing at each other.

The deeper lesson outlasts the benchmark. The frontier we obsess over is capability: can the model do the hard thing? This paper points at the quieter frontier that actually shows up on the bill: does the model know when the thing is easy, and have the confidence to act like it? An agent that can’t tell a two-line edit from a code audit will handle your simplest work the most expensively. It’s a judgment gap (nota a capability one), and judgment is the whole job.

Have you actually seen this in your own tools: an agent doing a full audit for a one-line change? Or is agent “overthinking” still more benchmark artifact than real cost?

Leave a comment

Source: Junjie Yin & Xinyu Feng, "Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution," University of Tennessee, Knoxville. arXiv:2607.13034 — https://arxiv.org/abs/2607.13034

Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.

Back to news