Give an Agent Your Best Experience and Watch It Get Worse
Researchers took an agent that was already performing well, handed it a bank of lessons distilled from its own past runs, and watched its success rate fall from 76.4 percent to 70.1 percent.
The lessons were not wrong. They were not irrelevant either. They had been retrieved precisely because they were semantically close to the task in front of the agent. That closeness was the problem.
The paper is MemHarness, posted at the end of July by a group spanning Zhejiang University, the Shanghai AI Laboratory, and several other institutions. It is the clearest demonstration I have seen of something that should concern anyone building a knowledge base for legal work: a retrieval system optimized for relevance has no way of knowing whether the thing it found still applies.
The gap between relevant and applicable
One example from the appendix makes the failure concrete. The agent had learned a search heuristic from an earlier run: if you cannot find a mug after opening several nearby cabinets and drawers, shift your search to open surfaces like countertops and the sink. That is sound advice, earned from a real failure.
The researchers then changed one line in the current observation. Drawer 5, previously empty, now contains a mug.
The retrieved memory is still relevant by every measure a retriever can compute. It is also now actively harmful, because the premise it was built on, that the search had already failed, no longer holds. An agent replaying it verbatim walks away from the object it was looking for.
The legal parallel is immediate. A decision comes back because the query matched its facts, and it is on point in exactly that sense. Whether the holding survives contact with your facts is a separate question, and no similarity score has ever answered it.
Two design choices carried the result
MemHarness inserts two steps between retrieval and action: the model critiques the retrieved experience against the present situation, then rewrites it into guidance specific to that situation, or discards it and falls back on its own reasoning.
Success rates climbed to 85.2 percent and 75.6 percent on the two benchmarks, against 76.4 and 66.1 for the same model trained without any of this.
Two implementation details did the work, and both translate directly.
The first is that every memory is stored together with the observation that produced it. Not the lesson alone, but the state it was learned from. When the researchers stripped that source context out, success fell from 85.2 to 80.0 percent while the rejection rate stayed essentially flat, which means the agent kept accepting mismatched guidance without noticing anything was off. When they instead paired each memory with a source state drawn at random from a different episode, rejection climbed from 8.7 to 13.3 percent. The comparison is real, and it is doing something.
For a legal tool the implication is blunt. A stored clause, holding, or drafting principle detached from the facts that produced it is an instruction that cannot be falsified. It can only be obeyed.
The second detail is that the critique lives inside the same trained policy, not in a wrapper around it. When the authors replaced the internal reconstruction step with a general instruction-tuned model asked to do the same rewriting, performance dropped from 85.2 to 77.7 percent. Adding a “check whether this applies” instruction to a prompt is not the same thing as a system that has been optimized, against real outcomes, to make that judgment.
One number deserves a moment on its own. The trained agent rejected 8.7 percent of retrieved memories in the household environment and 56 percent in the shopping environment. Same architecture, same training procedure. The right amount of skepticism turned out to be a property of the domain rather than a virtue of the model. Any vendor quoting one accuracy figure across contract review, litigation research, and regulatory analysis is quoting a number that has been averaged into meaninglessness.
The finding that will not fit in a demo
When the researchers switched the memory bank off entirely at test time, the model still beat the baseline, 83.0 percent against 76.4. Training it to interrogate past experience had made it better at reasoning without any.
That is the part I keep coming back to. Judging whether prior experience applies is not a feature layered on top of reasoning. It is a large part of what reasoning is, in law more than most places.
Three questions to ask your tool this week
Can it show me where the memory came from? When it surfaces a precedent, a clause, or a house drafting rule, does the interface put the originating facts next to the recommendation, or only the recommendation? If the source context is not on screen, nothing in the system is comparing it to yours.
Can it say no? Take a matter that superficially resembles one you have handled, and change one dispositive fact. Then ask. A tool that adapts its answer silently, without flagging that the prior situation no longer governs, is replaying, whatever the marketing calls it.
How often does it say no, and does that number move? If the refusal rate is identical across your practice areas, it is not measuring applicability. It is measuring a threshold someone picked once.
Loftus and Palmer showed in 1974 that human memory rebuilds the past rather than replaying it, and that this is what makes recall both fallible and useful. Fifty-two years later the engineering has arrived at the same place. The interesting question for our profession is no longer whether an AI system can find the right precedent. It is whether it can tell you when the right precedent is the wrong one.
Source: https://arxiv.org/abs/2607.28272