← BlogContext Compaction in LLM Agents

Context Compaction in LLM Agents: Part 3 - Post-Hoc Compilation

Isaac Kargar5 min read

  • LLM
  • Context Management
  • Compaction
  • AI Agent
  • Lerim

Part 1 and Part 2 covered methods that reduce context during a single run. Semantic compression rewrites text, pruning selects a subset, and learned policies decide when to compact. Those methods help the current run fit its context budget. They do not by themselves make useful decisions, constraints, or corrections available to the next run.

Online and post-hoc decisions

Online compaction runs while an agent is working. It has to choose what to keep before the outcome is known. A completed trajectory gives the compiler more evidence about what mattered. Online compaction remains necessary while a task is running, because the current agent still needs to fit within its context window.

Post-hoc compilation can inspect the full trace and its outcome before extracting reusable signal. That extra evidence does not make every extracted record correct. It gives the compiler a better basis for selecting records and lets each record point back to the source session for review.

Lerim: a context compiler

Lerim is an open-source context compiler for repeated agent workflows. Its current README describes a pipeline that imports completed traces, filters them into evidence-backed records, and exposes cited context for future work. The open-source core and any hosted offering are separate; the Lerim repository describes that boundary.

Lerim compiles completed agent traces into cited, reusable context records
The pipeline turns completed traces into selected context records with links back to their source sessions. Read the Lerim repository for the current integration boundary.

1. Capture

Lerim can ingest completed local sessions through the trace sources documented in its agent support matrix. The current repository lists native trace parsing for Claude Code, Codex CLI, Letta Code, and OpenClaw, with MCP configuration or explicit submission for other supported clients. Cursor and OpenCode are not native completed-session sources in the current 0.4.0 support table. A workflow without a supported source can provide already-clean trajectory-v1 JSONL through a custom trace folder.

Capture happens after the run. The agent does not need to call a memory tool on every turn, and routine traces can produce no durable record.

2. Compile

The compiler filters a trace for signal that a later run may need:

  • Decisions that should not be debated again
  • Constraints discovered during the work
  • Preferences about a repeated workflow
  • Facts that remain useful in the project or domain
  • Corrections that record a failed path and its fix
  • Handoffs that carry context to another run or agent

The record keeps a source boundary. A source reference makes it possible to inspect the session that produced the record, but it does not prove that the record’s interpretation is correct. A review or downstream evaluation is still needed when the consequence matters.

3. Reuse

The current repository documents these ways to use compiled context:

  • MCP tools such as lerim_context_brief, lerim_context_answer, lerim_context_search, and lerim_trace_submit
  • CLI queries such as lerim answer "What decisions exist about caching?"
  • A startup context brief for the next run

The Lerim CLI overview and MCP quickstart contain the current commands. The agent still needs an online compaction strategy when the active run approaches its context limit.

Across-run context and live memory

Lerim’s post-hoc workflow differs from systems that write memories during a conversation. Mem0 and A-Mem are examples of online memory approaches: they can make a memory available during the current session, but they decide what to save before the full trajectory is known.

Mem0 extracts and retrieves memories during a conversation
Mem0 illustrates online extraction and retrieval during a conversation. Read the Mem0 paper for its evaluated memory tasks.
Mem0 results comparing long-term memory quality and context use
These Mem0 results belong to the benchmarks reported by the paper and should not be read as a comparison with Lerim. Read the Mem0 paper.
A-Mem creates linked memory notes during an agent conversation
A-Mem creates linked notes online, before the complete trajectory is available. Read the A-Mem paper for its method and evaluation.

These approaches support different purposes. Online memory can support the active session. Post-hoc compilation can select signal after the result is known and make the source available for later review. Combining the two is an architectural choice that should be tested on a repeated workflow.

Current integration surface

The repository’s support boundary is explicit:

Integration pathCurrent role
Native trace parsingCompleted sessions from the sources listed as native in the support matrix
MCPQuery context and explicitly submit a completed trace where the client supports it
Custom trace folderImport clean trajectory-v1 JSONL for other workflows
Signal profilesFocus extraction for coding, support, operations, research, compliance, or a custom workflow

Lerim ships the coding, support, ops, research, compliance, and generic profiles. The custom-agent documentation explains how to use a profile or provide a clean trace. The profile determines what the compiler looks for; it does not make the resulting records true by itself.

What to measure

For a repeated workflow, measure whether compiled context changes the next run’s task outcome, retrieval quality, or review effort. The repository links LerimBench and reports separate surfaces for retrieval, context budget, and extraction diagnostics. Those measurements are more useful than counting records, because a large record set can still fail to help the next agent.

Post-hoc compilation and online compaction solve different timing problems. Online methods keep an active trajectory usable. Post-hoc compilation uses the completed trajectory to propose reusable, traceable context for later runs.

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →