Throughput Is Not Value: A Lawyer Reads Microsoft's AI Coding Study
Earlier this year, a Meta employee built an internal dashboard so colleagues could compete to be the company’s top AI token user. In one thirty-day stretch, employees burned through more than 60 trillion tokens, and the single heaviest user averaged 281 billion. On the cheapest Claude Opus tier, that one person alone could have run up a bill north of 1.4 million dollars.
I lead an AI R&D function inside a law firm, so numbers like that are not abstract to me. When the meter runs in the millions, the question every managing partner eventually asks is the one Microsoft’s researchers set out to answer: who actually uses these tools, do they keep using them, and does anything measurable come out the other end that justifies the spend.
A new field study from Microsoft (Murphy-Hill, Butler, and Savelieva) is the most honest attempt I have seen to answer that. It tracks tens of thousands of engineers through the company’s early-2026 rollout of two command-line agents, Anthropic’s Claude Code and GitHub’s Copilot CLI, over roughly four months. What makes it unusual is that the authors did not infer AI use from public signals the way most prior work does. They observed who could adopt and who actually did, using real telemetry. For anyone rolling out legal AI, the design alone is worth studying, and the findings are worth arguing about.
First lesson: adoption is social, not demographic
The strongest predictor of whether an engineer tried Copilot CLI in a given week was not seniority, tenure, or prior tooling. It was whether the people around them were already using it. The effect was largest for what the authors call skip-level peers, meaning the colleagues who share a manager’s manager. Where more than a quarter of that group had adopted, the odds of an engineer trying the tool rose by around 216 percent. Having a direct manager who used it lifted the odds by roughly 82 percent, and the reviewers someone regularly trades code with mattered too.
I read this and thought about every failed legal-tech pilot I have watched. Firms tend to roll out AI the way they roll out a new document management system, which means a vendor demo, a policy memo, and a mandatory training slot. That approach treats adoption as an information problem. This study says it is a social one. If the partners two desks over are visibly working with AI-drafted analysis and talking about it, associates follow. If adoption happens quietly behind closed doors, nothing spreads.
The practical instruction for a firm is uncomfortable but clear: make competent use visible on purpose, and let it travel through the review and reporting lines that already exist.
Second lesson: retention is behavioral, and the “first tool” advantage is real
Trying a tool and sticking with it turn out to be governed by different forces, which is the finding I keep coming back to. Engineers who had leaned heavily on AI inside their IDE were more likely to try the new command-line agent, yet they were slightly less likely to stay with it. Every retention marker for prior heavy IDE users came out negative.
The authors offer a reading I find persuasive. People who already trust an AI tool have a comfortable fallback, so they experiment with the new one and then drift back to what they know. Engineers for whom the command-line agent was their first serious AI tool had no such fallback, and when they stayed, they stayed firmly.
For legal AI this reframes an assumption I hear constantly, namely that the associates already using ChatGPT will be your best adopters of a purpose-built legal tool. Maybe not. They may kick the tires and then return to the generic tool they already trust. The lawyer who builds a genuine habit around your verified, firm-specific system may be the one for whom it is the first tool that ever earned their confidence on legal work. Retention is a question of habit formation rather than enthusiasm, and habits form around whichever tool becomes the path of least resistance.
Third lesson: the output moved, and it did not fade
Now the number everyone quotes. Using a synthetic-control method, the authors estimate that adopters merged about 24 percent more pull requests than they otherwise would have, and crucially the lift did not decay across the window. A comparable open-source study of the Cursor editor had found an early bump that vanished by the third month. Here, the gain in February and the gain in late April were statistically indistinguishable. A within-person analysis pointed the same way, since the more days an engineer used the tools in a week, the more they shipped, rising from roughly 15 percent more output at three days a week to around 50 percent at five or more.
If you are the one signing the invoice, that is the evidence you wanted. A real output metric moved, and it stayed moved.
The question a lawyer cannot skip: throughput is not value
Here is where I part company with the triumphant reading, and where the authors, to their credit, part company with it too. Their proxy for output is the merged pull request. They say plainly that a merged pull request is not the same as the value it delivers, and they close the paper by naming the open question directly, which is whether all this extra throughput actually produces better software. They do not claim it does. They say the field does not yet have the measures to know.
That distinction is the whole game for lawyers. Our version of the merged pull request is the drafted motion, the produced memo, the reviewed contract. It is entirely possible to produce more of those, faster, while producing less value, because a legal work product that is 24 percent faster and 5 percent wrong is not 24 percent better. It is a liability with a shorter turnaround.
Speed that outruns verification does not compound into value, it compounds into risk, and in our field that risk lands on a client and on a professional license.
This is precisely why the work I spend my days on sits on the verification side rather than the generation side. Getting a model to produce a confident-sounding brief is easy and mostly solved. Knowing whether each citation is real, whether each holding says what the draft claims it says, and whether the reasoning is grounded rather than merely plausible is the hard part, and it is the part that turns throughput into value instead of exposure. The Microsoft study is a clean demonstration that generation-side gains are real and measurable. It is also a clean demonstration that the profession still has to build the quality measure that sits underneath them.
A useful embarrassment about which tool won
One more finding deserves attention because of how the authors handle it. On merged-pull-request throughput, Copilot CLI outperformed Claude Code by more than two to one, even though public sentiment in early 2026 generally rated Claude Code as the stronger autonomous agent. The authors do not pretend to resolve this. They offer two hypotheses, that engineers reach for the two tools for different kinds of work, and that Microsoft owns GitHub while it merely buys Claude Code, so the Copilot harness was probably tuned to fit how Microsoft engineers actually work. They also flag their own position openly, since they are Microsoft employees studying a Microsoft-owned tool.
I appreciate this more than a cleaner result would have earned. It is a reminder that a throughput number measures a tool inside a specific organization with specific incentives and specific integration, and it is not a verdict on model quality in the abstract. Anyone who has deployed the same model in two different firms and watched it succeed in one and stall in the other already knows this. The harness, the integration, and the fit to real workflows often matter more than the raw model.
What I am taking into my own rollouts
Three things:
Make competent use visible, because adoption travels through people rather than memos;
Watch retention instead of sign-ups, because the tool that becomes someone’s first real habit is worth more than the one everyone tries once;
And refuse to let the throughput number stand in for value, because in law the gap between faster and better is not a rounding error, it is the entire professional question.
The engineers in this study merged more code. Whether the profession that adopts these tools produces better work, or merely more of it, is the measure we still have to build.
Which side of that line is your firm actually measuring?
Reference
Emerson Murphy-Hill, Jenna Butler, and Alexandra Savelieva, “Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft’s Early 2026 Rollout of Claude Code and GitHub Copilot CLI,” Microsoft, https://arxiv.org/abs/2607.01418 [cs.SE], July 1, 2026.