A Model That Sorts Its Own Doubt
To make a language model write faster, NVIDIA’s researchers took half of it and froze it solid. That half can no longer learn anything, and that is exactly why it works. The frozen half is what lets the other half run more than twice as fast at almost no cost to quality. I want to walk through why, because the same move is the one I keep reaching for when I try to make legal AI faster without making it less safe to rely on.
So They Built It Twice
Start with the problem the paper solves. Most models write the way a court stenographer takes dictation, one token after another in strict order, which is accurate but slow, because every word has to wait for the one before it. Diffusion language models try to escape that by generating text in parallel and then refining it. The trouble is that in the usual design a single network has to do two jobs at once. It has to hold a reliable picture of the text so far, and it has to clean up the noisy tokens it is producing. Those two jobs pull the same weights in opposite directions, so the model does neither as well as it could.
TwoTower, the model in the paper, refuses to make one network do both. It keeps two copies of a pretrained model and gives each a single job. The first copy, the context tower, reads the clean text causally, the way a normal model does, and it stays frozen, which means its weights never change during training. The second copy, the denoiser tower, is trained to take a block of masked tokens and resolve them by looking across at the frozen tower’s understanding, layer by layer. One tower understands, and the other writes. The understanding stays fixed and trustworthy, while the writing gets faster.
The Number That Should Not Work
The numbers are what earned my attention, because they resist the trade-off I expect.
Built on a 30B open-weight hybrid model and trained on a fraction of the tokens used to pretrain the backbone, TwoTower keeps 98.7 percent of the original model's benchmark quality while producing text 2.42 times faster in wall-clock terms.
You can push it past three times faster by loosening one threshold, though quality then starts to slip, and the authors say so plainly.
The Part I Did Not Expect
This is where the speed becomes the least interesting thing in the paper. The model does not hand over a finished block all at once. At each step it predicts every masked position, commits only the ones it is confident about, and leaves the uncertain positions blank for another pass. The early steps commit many tokens, and the later steps grind through the hard remainder. The model shows you its own uncertainty before you ask.
Read that again with your due diligence file open. A system that separates what it is sure of from what it is still guessing is doing, at the level of individual words, the same triage a junior associate does with a highlighter. That is a signal a verification layer can act on. A governance system no longer has to treat the model as a black box that slides a finished paragraph across the table. It can step in at the block boundary, where the model has already marked its own weak spots, and then decide what to check, what to escalate, and what to let through. Speed and control stop being enemies once the fast part of the system is built to admit what it does not know.
Where the Speed Quietly Does Nothing
Now the honest question, the one worth sitting with.
Where does this speed actually pay off in legal work, and where does it quietly do nothing?
It pays off in generation, and only in generation. The paper measures throughput as the time to produce the final answer, not the time to find or read the input, and that line is easy to skate past and expensive to ignore. If you are working a knowledge graph built from millions of documents, your bottleneck is retrieval, which is finding the right nodes and the paths between them. TwoTower does nothing for that first mile. What it changes is the last mile. Once the relevant slice of the graph is in context, generating the synthesis over it becomes much cheaper, and the frozen context tower is built to hold that retrieved slice as stable, reusable memory while the denoiser writes. So the honest version is that this helps you write the memo faster after the graph has done its work, and not walk the graph itself.
Where Volume Changes the Math
The non-performing loan due diligence case is the one that changes the math, because volume does the heavy lifting. Reviewing several hundred loan files is not one long generation, it is a few hundred short ones, the same structured extraction repeated file after file. Reading and pulling the data out of those PDFs is a separate problem, and it is better solved by processing each document on its own than by asking a model to hold the whole portfolio in its head. But once you reach the generation step, producing a standard summary or a risk flag for each file, a speedup of that size compounds across the entire portfolio. And the confidence-graded output sorts the work for you, because the files the model finished in a flash are the clean ones, and the files it kept reworking are the ones a person should open first.
What Actually Travels
None of this makes the released model ready to sign an opinion. It is a base model, measured before any instruction tuning or alignment, and the idea that travels is the architecture rather than the specific checkpoint. The lesson is that you can freeze the part of a system that has to be trusted and speed up the part that only has to be quick, and you can keep the two apart on purpose. A verification layer that governs legal output wants precisely that separation: a stable source of “truth” on one side, a fast producer on the other, and a clean boundary in between where the checking happens. TwoTower builds that boundary into the model, and the rest of it is our job.
Source: Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro, "Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context," NVIDIA, 2026. https://arxiv.org/abs/2606.26493