When One AI Agent Optimizes Another Agent's Memory
This report describes an experiment on one snapshot of Lerim. Claude Code, running Opus 4.6, proposed changes to the prompts, schemas, tools, and harness around memory extraction. The extraction system ran on MiniMax M2.5. An evaluator measured each candidate against a set of reference cases and kept changes that improved the measured score.
Across the first round, the composite evaluation score rose from 0.61 to 0.86, a 41% relative increase. That number belongs to this composite evaluation and its cases. It is not a claim that memory quality improves by 41% for every workload.
The system
Lerim’s current repository describes an open-source context compiler for repeated workflows. It ingests traces, curates context records, and produces briefs that can be used in later work. Native trace support and other integration paths depend on the client. Hosted offerings, where available, are separate from the repository’s core functionality. See the Lerim repository for the current product boundary.
The experiment examined an earlier memory pipeline with several places where an optimizer could make a change: system prompts, DSPy signatures, tool descriptions, schema field descriptions, harness logic, and post-extraction filters.

The evaluator combined model-based judgements of whether a memory was atomic and actionable with deterministic checks for titles, search relevance, duplicate decisions, and maintenance precision. This combination produced one composite score for the optimization loop.
Evaluation design
The first round used 15 reference cases. Each trial changed one part of the system, ran the component evaluator, and was kept or discarded according to the measured result. Fourteen trials were run: seven were kept and seven were discarded.
The component evaluation took approximately 15 minutes per trial. The end-to-end lifecycle check took approximately seven minutes. These figures measure different stages and should not be read as a direct speed comparison.
Round 1 results
The component metrics changed as follows. Relative changes are approximate, calculated from the displayed scores.
| Metric | Before | After | Relative change |
|---|---|---|---|
| Extraction quality | 0.69 | 0.88 | ~28% |
| Search relevance, NDCG@5 | 0.91 | 0.91 | no change |
| Deduplication accuracy | 0.28 | 0.72 | ~157% |
| Maintain precision | 1.00 | 1.00 | no change |
| Composite evaluation score | 0.61 | 0.86 | ~41% |
The deduplication score moved the most. The baseline classified many nearly identical candidates as new memories. The optimized version identified more of those duplicates and updated existing records instead.

The largest single change was at a DSPy call site. The extraction module changed from dspy.Predict(MemoryExtractSignature) to dspy.ChainOfThought(MemoryExtractSignature). In this experiment, extraction output became more consistent and the resulting candidates were easier to classify.
Other retained changes improved the descriptions of schema fields and supplied explicit similarity thresholds for deduplication decisions. Changes to summarization, tool descriptions, restrictive extraction rules, and body-format guidance were discarded after they reduced the measured score or affected other flows unpredictably.
End-to-end check
The component result was followed by a lifecycle evaluation with three sequential sessions and a maintenance cycle. The maintain score improved by 29% even though the maintain prompt was not changed directly. Higher-quality extracted memories then reached the maintenance stage, which is consistent with the pipeline dependency. The measured result remains specific to this lifecycle evaluation.
| Lifecycle metric | Before | After |
|---|---|---|
| Sync composite | 0.904 | 0.925 |
| Maintain composite | 0.667 | 0.860 |
| Overall end-to-end score | 0.845 | 0.909 |

The deduplication result varied across trials, ranging from 0.17 to 0.72 in the plotted progression. The variation was attributed to nondeterministic classification outputs. The retained configuration shifted the observed results upward, but the spread is a reason to repeat evaluations before treating a small difference as meaningful.

What the first round taught
The keep or discard loop exposed several useful patterns in this evaluation:
- A schema description can shape an output more directly than a long general prompt. The
titlefield description was changed from a short label to a self-contained title format with a maximum length. - Explicit thresholds can make a classification boundary easier for a model to follow than phrases such as “very high similarity.”
- A change in one stage can affect later stages. The extraction module influenced the quality of candidates that deduplication and maintenance received.
- Seven of 14 trials were discarded. A candidate that looks sensible in isolation still needs an evaluation that can catch regressions.
Round 2: measuring useful memories
The first evaluator rewarded finding the reference memories, but it did not penalize extracting implementation details that an agent would not use later. The second round added a quality-alignment dimension and changed the weighting so that precision and quality alignment together represented half of the composite score. Completeness represented 15%.
The second dataset contained 327 cases across 20 categories, including 70 negative cases where no memory should be extracted and 50 mixed cases with one or two decisions buried in implementation noise. A 30-case evaluation was used for quick trials and the full 327-case set was used at checkpoints.
Ten trials were run in this round. Three were kept and seven were discarded. The extraction score rose from 0.819 to 0.847, a 3.4% relative increase.
| Round 2 metric | Before | After |
|---|---|---|
| Extraction | 0.819 | 0.847 |
| Search | 0.905 | 0.905 |
| Maintain | 1.000 | 1.000 |
The retained changes added positive quality criteria to the extraction signature, described a WHY and HOW TO APPLY structure for memory bodies, and included one positive example. Across the reported trials, restrictive guidance reduced recall. That observation is limited to the five restrictive-rule trials in these two rounds; it is not a general law about all negative instructions.
The second round also showed why a conservative duplicate decision matters. Weakening the default used when the system was uncertain reduced the maintain score from 1.000 to 0.667 in one trial. That result supports keeping the evaluated default for this system; it does not establish a universal rule for every memory architecture.
The discarded restrictive-instruction trials included:
| Round | Instruction tested | Decision |
|---|---|---|
| 1 | Do not extract lists | Reverted |
| 2 | If in doubt, do not extract | Reverted |
| 2 | More examples of information to skip | Reverted |
A reusable pattern
The experiment suggests a reusable division of responsibility. An optimization loop can propose and evaluate small changes, while a fixed evaluator defines what counts as an improvement. The objective can change between rounds, but the evaluation must state its cases, metrics, and limitations each time.
The experiment does not show that one model can improve every other model, or that a single module change will transfer to another memory system. It shows that, for this pipeline and these cases, measured changes to prompts, schemas, and modules produced different outcomes and that several plausible changes were worth discarding.
Work with Nazmi
Build your AI system with Nazmi.
Tell us what you are building, what exists today, and where your team needs help.
Start a conversation or book a 20-minute call →