They Locked Memory Inside the Model, Then Tried to Empty it
Every time you open a new session with an AI assistant about a matter you have been handling for months, the same thing happens: the system retrieves the relevant documents, pastes them back into the prompt, reads them from scratch and only then answers. Then the conversation closes, all that work gets thrown away and with the next request the cycle starts over exactly as before. This is how RAG works, and it is the reason why memory in the legal products on the market today looks more like a very fast filing clerk than a colleague who knows you.
On August 4 a group of researchers from MemTensor, together with Renmin University, the National University of Singapore, Shanghai Jiao Tong and Tongji, published a paper that tries to change where memory lives. The system is called Metis and it is presented as the first prototype of what the authors call a memory foundation model. It is worth reading, although it is worth reading mostly for the mechanism, because on results the paper is unusually honest and spends most of its pages describing where that mechanism breaks down.
What Metis Changes
The underlying idea is that memory should stop being a module external to the model and become part of its internal computation.
Inside every layer, Metis maintains a fixed-size state that the model updates after each interaction with a single forward pass, without any training at runtime.
The weights of the base model stay frozen, and during training only the parameters dedicated to memory get optimized. When the next question arrives, the model reads that state in parallel with ordinary attention, so past history takes part in the reasoning without being rewritten into the prompt.
The interesting part for anyone working in a law firm concerns the operations the system learns to perform. The authors train the model on four explicit behaviors: remembering, updating, forgetting and connecting pieces of information that sit far apart. Anyone who has handled a file for three years recognizes that cycle, because it is exactly what we do when a client changes its registered office, when a position is transferred to another servicer, when a retainer is revoked and when two separate documents become relevant only if read together.
The Numbers, With No Discount
Without context, the starting models collapse. On LoCoMo, the conversational benchmark used in the paper, the Qwen3.5 backbones stripped of their history score close to zero, which confirms that the answers cannot come from the model’s prior knowledge. Metis in its largest version reaches an average of 26.74 on that same test, so it really does recover something that used to be lost. Meanwhile the same model with the full context available stays at 65.03, and that distance is the real headline of the experiment.
The rest of the limitations are even more instructive. Forgetting remains the hardest operation of all, with the lowest score at every scale tested, and the authors explain that suppressing information inside a shared latent space generalizes far worse than adding it. Over long trajectories the capacity degrades visibly: as updates pile up, the facts introduced first grow progressively weaker and even the intermediate ones become unstable, because each new update interferes with the whole state instead of overwriting only the oldest part of it. Semantically similar pieces of information end up getting confused with each other. Finally, when memory fills up with irrelevant material, the model’s general capabilities degrade: on the benchmark that measures instruction following, the score drops from 76.71 for the base model to 54.53 for the version with memory active.
The authors put the resulting conclusion in writing, namely that native memory cannot yet be considered a replacement for external memory. On a topic where marketing runs faster than research, a sentence like that in a paper signed by a memory vendor is worth quite a lot.
Why It Still Counts as a First Step
If performance were the only dimension, the work would be easy to file away. The point is that it shifts two cost items that weigh heavily in a firm.
The first one is storage. The cache that keeps a session alive today grows along with the client’s history, while the Metis memory state stays constant. At thirty-two thousand tokens of history the former takes up more than a gigabyte, the latter just under seventeen megabytes, and with low-rank compression it drops to a little over two megabytes while preserving 99.9 percent of average performance. A profile like that means a client’s history can be kept for years without the cost of keeping it alive going up month after month.
The second one is separation between clients, which for us is a matter of professional duty before it is a technical question. In the serving architecture proposed in the paper, each user’s state occupies a distinct slice of the batch, so interference between different sessions is ruled out by construction. Two hundred cross-user probes, in which one user was asked for another user’s information, produced no leakage at all. The authors add the observation that makes this concrete for anyone who has to answer to a data protection authority: deleting a user amounts to deleting a single state file. Anyone who has tried to carry out an erasure request inside a distributed vector index knows how much that sentence is worth.
There is a third element, less immediate. The paper shows that the same training also works on models from different families, so the technique is not tied to one vendor. If the direction holds, a client’s memory becomes a portable artifact that survives a change of model, and that changes a firm’s negotiating position when it sits across from a vendor.
The Bench Test
The practical way to use this paper does not involve buying anything, because there is nothing to buy yet. It involves taking the four questions the paper finally makes sensible and carrying them into your next meeting with whoever is selling you a feature called memory.
Where does my client’s memory physically live, and what happens when I ask for it to be deleted? If the answer describes a shared index from which references get removed, you are buying an archive, with everything that follows in terms of proving the deletion.
Can the system forget on request, and how do I verify it? The research says this is the most fragile operation, so an answer that comes back too confident is itself a signal.
What happens to the information I entered first after a year of use? The decay of older facts is documented and it affects any architecture that compresses history, so the question applies to products built differently as well.
Does having memory switched on make the system worse at following my instructions? This is the most underrated effect of all, because it shows up as an assistant that slowly becomes less precise exactly while it seems better informed.
None of these questions requires any machine learning expertise to ask, although all of them require that whoever answers has read something like this paper.
What I Am Watching
The authors close with a five-level roadmap that runs from persistent state all the way to a model capable of reorganizing its own experience. The level that concerns our profession is the second one, where the system autonomously decides what to keep, what to update and what to let go. On that day the question will stop being technical and will become a matter of professional responsibility, because a system that decides on its own what to forget about a case file is making a judgment call that so far has always been ours.
Metis does not get there today, and its authors are the first to say so. Still, the first step in one direction counts for more than the tenth step in another, and this direction is worth watching now, while it is still a research problem.
-
The paper: [Metis: Memory Foundation Model, arXiv:2607.26760](https://arxiv.org/abs/2607.26760)