The Colleague Who Only Speaks Before You Hit Send

A new Meta AI paper about agent memory is, at first glance, a paper about information storage. Read more closely, however, and it becomes a paper about supervision, and about why the most useful kind of supervision is often the kind that says very little.

Consider a simple example: a customer tells an agent that she is a Gold member and asks what compensation Gold members receive; the agent has already retrieved her record, which identifies her as a Regular customer, but it pays out the Gold benefit anyway.

Nothing was missing from the system: the database was queried, the answer came back, and the agent had access to it. The problem was that the conversation continued, and by the time the decision had to be made, the verified fact had quietly stopped influencing the agent’s behaviour.

The paper calls this behavioural state decay. Anyone who has spent more than a year practising law will recognise the pattern immediately.

a dark room with a bunch of framed pictures on the wall
Photo by Yusuf Evli on Unsplash

The file is open, but nobody is reading it

Wu and his colleagues at Meta AI define behavioural state decay as the gradual loss of influence that information suffers during a long task.

The information remains in the transcript and may even remain inside the model’s context window, but it no longer meaningfully constrains the next decision.

Once you translate that out of machine language, it stops sounding like an AI-specific problem. It is the restriction you flagged in week one of due diligence, only to see it breached in month four while everyone was dealing with an unrelated timetable; it is the client instruction recorded at intake and then contradicted by the brief sent out in March; or it is the argument you deliberately decided not to run, for reasons you documented carefully, only to see it revived by a colleague who read the folder but missed the relevant memo.

In each case, the file is complete and nobody has lost the information. The real problem is that it is no longer operational at the moment when it should affect someone’s conduct.

That distinction is the first thing the paper gets right, and it is also where much of the legal AI conversation goes wrong. We tend to treat memory as a storage problem because storage is something we know how to buy: document management systems, matter workspaces, archive-wide search and retrieval. Yet law firms are already very good at storing information. What they are less good at (and what agents appear to struggle with in much the same way) is reactivation: bringing the right piece of stored knowledge back into the conversation when it needs to change what happens next.

A second agent that mostly stays quiet

The architecture described in the paper is deliberately modest. The working agent is left unchanged, while a second model runs alongside it. This supervisor sees the task, a recent window of activity and its own memory bank, and it performs two related functions.

It first maintains the memory bank, which is divided into three parts: a private status field; stable knowledge, including requirements, environmental facts and constraints; and procedural records of what was tried, what failed and what worked. It then decides whether any of that information needs to be surfaced at the next step. If it does, the supervisor sends one short reminder. If it does not, it says nothing.

The second point is the more important one. The status field is never shown to the working agent, so the supervisor retains its own private view of progress, open issues and unresolved risks without passing all of that information downstream.

Anyone who has watched a capable senior associate manage a matter will understand why this matters. You do not hand a junior every concern you have about a case, nor do you forward every half-formed suspicion as soon as it occurs to you. You keep those concerns in reserve and use them in short, carefully timed interventions when they are most likely to affect the work.

The result nobody will quote

The headline numbers are respectable. On Terminal-Bench 2.0, the weaker action agent improves from 37.6% to 45.9% pass@1 across the 85 evaluated tasks. On tau2-Bench, the task-weighted average rises from 55.0% to 61.8%.

When the researchers use a stronger action agent, the gains become smaller, but they do not disappear: 2.4 points on one benchmark and 2.5 on the other. That matters because it suggests the system is not merely compensating for a weak underlying model.

The ablation results are more revealing, however, because they show what kind of memory actually helps. If the memory bank is retained but shown to the working agent in full at every step, performance falls well below the selective version. If the bank and reminders are retained but the supervisor is forced to inject something after every step, the results remain competitive on the task-weighted average while weakening on the domain-balanced score; the authors interpret the small difference as run variance.

Removing the bank altogether and keeping only a model that watches and advises (the familiar advisor pattern) produces unstable results. It helps in one domain but pushes the airline domain below the no-memory baseline. Replacing the system with a production memory layer such as Mem0, using vector and keyword retrieval, improves the average but does not improve the airline domain at all.

These are four plausible ways of trying to make an agent more helpful. The strongest result comes from the least intrusive one: maintain state, but give the supervisor permission to remain silent.

The authors make the point directly: doing nothing is an explicit action. Silence is not a failure of the system to intervene, but It is one of the ways the system performs its job.

Where the useful interruption happens

The qualitative analysis is brief, but it is probably the most practically useful part of the paper. Successful interventions tend to arrive immediately before a state-changing call, at the point where the agent is about to commit to something that cannot easily be undone.

The reminder itself is usually only one sentence: the customer’s status is verified as Regular; the fare class cannot be modified; the authentication step has not yet taken place. The value lies less in the amount of information than in its timing.

That finding sits awkwardly beside the way legal quality control is normally organised. Most of our controls are front-loaded: the kickoff memo, the intake checklist, the engagement letter, the onboarding pack and the matter-opening call in which everything relevant is explained once, usually during the first month.

The paper suggests that this is not enough. Giving people or systems all the relevant context at the beginning of a matter does not guarantee that the context will still influence their decisions several weeks later. In fact, the paper’s full-context version (the one that exposes the entire memory bank at every step) is precisely the version that performs worse.

The irreversible acts in legal practice are not hard to identify. They include the filing that fixes your pleadings, the letter before action that establishes a tone you cannot easily take back, the waiver of an objection, the signed release and the email to the other side that concedes a point almost invisibly inside a subordinate clause.

Translated into legal language, the paper is making a fairly simple claim: one sentence delivered thirty seconds before one of those acts may be worth more than a thorough memo delivered eight weeks earlier. That does not make the memo unnecessary. The memo records the reasoning and creates an audit trail; the reminder brings the relevant conclusion back into view when it can still guide conduct.

The cost of a supervisor who talks too much

The failure analysis is unusually candid, and this is where the legal analogy becomes uncomfortable. When memory hurt performance, it was rarely because the system had stored the wrong information. More often, the problem was calibration: the memory agent presented a speculation too confidently, repeated something the working agent already knew or raised a plausible but unnecessary concern that triggered another round of verification.

Every one of those behaviours has a human equivalent in a law firm. There is the partner who states a risk as though it were a certainty, the colleague who explains to you what you have just explained to them and the reviewer whose margin comments are so consistently urgent that, after a while, none of them feel urgent anymore.

In technical language, the paper is describing how supervision can destroy its own authority. A channel that speaks constantly carries very little usable signal, and the recipient eventually learns to discount it. The problem is not simply that the supervisor is annoying; it is that excessive intervention changes the listener’s estimate of how much attention any individual warning deserves.

The training results make the same point in a more concrete way. The researchers fine-tune a smaller open model to act as the memory agent. When the model is used without training, it makes the system worse than having no memory at all, reducing average reward from 0.709 to 0.693. Supervised fine-tuning recovers the loss and raises the score to 0.720, while reinforcement learning takes it to 0.734.

Only after that training does the memory agent improve the frozen action agent on the held-out benchmark, raising its score from 37.6% to 41.1%.

This should be read as a procurement warning. An uncalibrated supervisor is not neutral but it is actively harmful. A memory layer that has not learned when to stay quiet resembles a junior who has been told to interrupt senior lawyers whenever something seems relevant, but has never been taught to distinguish a genuine warning from a passing thought.

The legal bench test

Here is how I would use the paper on Monday morning, whether the supervisor being designed is software or a person.

  1. The first question is where the point of no return lies. Identify the specific irreversible act in the workflow. If there is no such act, you probably do not need an intervention layer; you need a summary, and you should stop paying for the difference. If there are three irreversible acts, those are the three moments worth protecting and everything else is secondary.

  2. The second question is whether the relevant constraint comes from earlier in the process. The Gold-versus-Regular example only matters because the record was retrieved earlier and then lost its influence. If everything needed to make the decision correctly is already visible in the document currently open, there is nothing to reactivate and therefore little to gain. Reactivation creates value only when the relevant constraint and the final decision are separated in time.

  3. The third question is whether you can describe what the supervisor would not flag. Write that list down. If the honest answer is “anything materially relevant,” you have effectively chosen the full-context approach, which the paper shows performs worse. A supervision policy that cannot explain its own silence is not really a policy.

  4. The fourth question is who holds the private status. Someone has to carry the doubts without publishing all of them. In the architecture, that is a field the working agent never sees. In a firm, it is a person. If that person does not exist, the reminders are likely to arrive as anxiety rather than instruction.

Before buying anything, I would test the idea manually. Choose one recurring irreversible act, such as a particular type of filing or an outgoing demand letter. For two weeks, ask one person to keep a short private note for each matter and allow them exactly one sentence per matter, delivered only in the minutes before that act.

Then count how often the sentence actually changed what happened.

If the answer is never, the constraint was already visible and the supervision was mostly theatre. If the answer is often, you have found the precise point at which a memory layer might justify its cost, and you now know what information it needs to receive.

Either way, you have learned something useful for the price of two weeks of attention rather than the cost of a full procurement cycle.

The name is wrong

“Memory” is a misleading name for what this paper is really describing. The word suggests a warehouse: a place where information is stored and retrieved when someone asks for it.

What the system actually provides is closer to timing. The firm already has the warehouse. What it usually lacks (and what many legal AI products are not designed to provide) is the discipline of the colleague who reads everything, says almost nothing and places a hand on the door at exactly the right moment, just before you hit send.


Source

Wu, Zhang, Zhou, Wang, Peng, Li, Fan, Zhao (Meta AI), Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents, July 2026: https://arxiv.org/abs/2607.08716

Torna alle news