← BlogAI Agents in Production

When One AI Agent Optimizes Another Agent's Memory

Isaac Kargar7 min read

  • AI Agents
  • Optimization
  • LLM
  • Memory

This report describes an experiment on one snapshot of Lerim. Claude Code, running Opus 4.6, proposed changes to the prompts, schemas, tools, and harness around memory extraction. The extraction system ran on MiniMax M2.5. An evaluator measured each candidate against a set of reference cases and kept changes that improved the measured score.

Across the first round, the composite evaluation score rose from 0.61 to 0.86, a 41% relative increase. That number belongs to this composite evaluation and its cases. It is not a claim that memory quality improves by 41% for every workload.

The system

Lerim’s current repository describes an open-source context compiler for repeated workflows. It ingests traces, curates context records, and produces briefs that can be used in later work. Native trace support and other integration paths depend on the client. Hosted offerings, where available, are separate from the repository’s core functionality. See the Lerim repository for the current product boundary.

The experiment examined an earlier memory pipeline with several places where an optimizer could make a change: system prompts, DSPy signatures, tool descriptions, schema field descriptions, harness logic, and post-extraction filters.

Lerim memory pipeline with extraction, search, deduplication, and maintenance stages
The evaluated pipeline combined extraction, search, deduplication, and maintenance. The optimizer could change instructions and supporting code around those stages while the evaluator remained fixed.

The evaluator combined model-based judgements of whether a memory was atomic and actionable with deterministic checks for titles, search relevance, duplicate decisions, and maintenance precision. This combination produced one composite score for the optimization loop.

Evaluation design

The first round used 15 reference cases. Each trial changed one part of the system, ran the component evaluator, and was kept or discarded according to the measured result. Fourteen trials were run: seven were kept and seven were discarded.

The component evaluation took approximately 15 minutes per trial. The end-to-end lifecycle check took approximately seven minutes. These figures measure different stages and should not be read as a direct speed comparison.

Round 1 results

The component metrics changed as follows. Relative changes are approximate, calculated from the displayed scores.

MetricBeforeAfterRelative change
Extraction quality0.690.88~28%
Search relevance, NDCG@50.910.91no change
Deduplication accuracy0.280.72~157%
Maintain precision1.001.00no change
Composite evaluation score0.610.86~41%

The deduplication score moved the most. The baseline classified many nearly identical candidates as new memories. The optimized version identified more of those duplicates and updated existing records instead.

Composite evaluation score across the first round of optimization trials
The composite score rose from 0.608 to 0.855 across the first round, rounded to 0.61 and 0.86 in the table. The curve distinguishes retained and discarded trials.

The largest single change was at a DSPy call site. The extraction module changed from dspy.Predict(MemoryExtractSignature) to dspy.ChainOfThought(MemoryExtractSignature). In this experiment, extraction output became more consistent and the resulting candidates were easier to classify.

Other retained changes improved the descriptions of schema fields and supplied explicit similarity thresholds for deduplication decisions. Changes to summarization, tool descriptions, restrictive extraction rules, and body-format guidance were discarded after they reduced the measured score or affected other flows unpredictably.

End-to-end check

The component result was followed by a lifecycle evaluation with three sequential sessions and a maintenance cycle. The maintain score improved by 29% even though the maintain prompt was not changed directly. Higher-quality extracted memories then reached the maintenance stage, which is consistent with the pipeline dependency. The measured result remains specific to this lifecycle evaluation.

Lifecycle metricBeforeAfter
Sync composite0.9040.925
Maintain composite0.6670.860
Overall end-to-end score0.8450.909
Changes tested by the first-round optimizer
The trials covered module choices, schema descriptions, thresholds, tool descriptions, and extraction guidance. Several changes were reverted after their measured score fell.

The deduplication result varied across trials, ranging from 0.17 to 0.72 in the plotted progression. The variation was attributed to nondeterministic classification outputs. The retained configuration shifted the observed results upward, but the spread is a reason to repeat evaluations before treating a small difference as meaningful.

Per-dimension progression across the first round
The plotted progression shows both the upward shift in deduplication results and the variability between trials.

What the first round taught

The keep or discard loop exposed several useful patterns in this evaluation:

  1. A schema description can shape an output more directly than a long general prompt. The title field description was changed from a short label to a self-contained title format with a maximum length.
  2. Explicit thresholds can make a classification boundary easier for a model to follow than phrases such as “very high similarity.”
  3. A change in one stage can affect later stages. The extraction module influenced the quality of candidates that deduplication and maintenance received.
  4. Seven of 14 trials were discarded. A candidate that looks sensible in isolation still needs an evaluation that can catch regressions.

Round 2: measuring useful memories

The first evaluator rewarded finding the reference memories, but it did not penalize extracting implementation details that an agent would not use later. The second round added a quality-alignment dimension and changed the weighting so that precision and quality alignment together represented half of the composite score. Completeness represented 15%.

The second dataset contained 327 cases across 20 categories, including 70 negative cases where no memory should be extracted and 50 mixed cases with one or two decisions buried in implementation noise. A 30-case evaluation was used for quick trials and the full 327-case set was used at checkpoints.

Ten trials were run in this round. Three were kept and seven were discarded. The extraction score rose from 0.819 to 0.847, a 3.4% relative increase.

Round 2 metricBeforeAfter
Extraction0.8190.847
Search0.9050.905
Maintain1.0001.000

The retained changes added positive quality criteria to the extraction signature, described a WHY and HOW TO APPLY structure for memory bodies, and included one positive example. Across the reported trials, restrictive guidance reduced recall. That observation is limited to the five restrictive-rule trials in these two rounds; it is not a general law about all negative instructions.

The second round also showed why a conservative duplicate decision matters. Weakening the default used when the system was uncertain reduced the maintain score from 1.000 to 0.667 in one trial. That result supports keeping the evaluated default for this system; it does not establish a universal rule for every memory architecture.

The discarded restrictive-instruction trials included:

RoundInstruction testedDecision
1Do not extract listsReverted
2If in doubt, do not extractReverted
2More examples of information to skipReverted

A reusable pattern

The experiment suggests a reusable division of responsibility. An optimization loop can propose and evaluate small changes, while a fixed evaluator defines what counts as an improvement. The objective can change between rounds, but the evaluation must state its cases, metrics, and limitations each time.

The experiment does not show that one model can improve every other model, or that a single module change will transfer to another memory system. It shows that, for this pipeline and these cases, measured changes to prompts, schemas, and modules produced different outcomes and that several plausible changes were worth discarding.

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →