← BlogContext Compaction in LLM Agents

Context Compaction in LLM Agents: Part 2 - Learning to Compact

Isaac Kargar4 min read

  • LLM
  • Context Management
  • Compaction
  • AI Agent

In Part 1, we separated semantic compression, token or sentence pruning, and dynamic summarization. This post looks at three papers that add a learned or model-guided decision to that process. Their training requirements and evaluation settings differ, so their numbers should be read in those settings.

Inference-time self-compaction: SelfCompact

SelfCompact gives a model a compaction tool and a task-specific rubric. At probe intervals, the model decides whether a subtask has finished or a trajectory is converging. It can suppress compaction during a derivation or when the agent is stuck. The method does not fine-tune the model.

The paper evaluates four Qwen configurations on IMO-Answerbench, HMMT November 2025, and HMMT February 2026. Under token budgets matched to fixed-interval summaries, SelfCompact has the best score in 11 of 12 benchmark and model cells. For Qwen3.5-9B with thinking disabled, its score on HMMT February 2026 is 52.3%, compared with 34.2% for no compaction, a difference of 18.1 percentage points.

For a separate agentic-search evaluation, the paper uses GLM-4.7-Flash, MiniMax-M2.5, and Mimo-V2-Flash on BrowseComp, BrowseComp Plus, and DeepSearch QA. It reports gains of 5 to 9 accuracy points and 30% to 70% lower per-question token cost than its no-compaction baseline. The unit is per-question cost in that experiment, not a universal token-saving rate.

SelfCompact uses a rubric to decide when to summarize an agent trajectory
SelfCompact combines a compaction tool with a rubric that changes the decision at inference time. Read the SelfCompact paper for models, tasks, and baselines.

Reinforcement-learning compaction: CompactionRL

CompactionRL puts compaction inside the rollout. A single trainable policy generates ordinary execution actions and, when the context budget is nearly full, a summary. The rollout resumes from that summary plus a short recent tail. The final task reward is assigned to both execution and summary segments. The paper does not add a separate human-readable summary reward; it propagates credit across segments with a cross-trajectory advantage calculation.

This is different from a fixed summarizer paired with a separate execution model. The summary is sampled from the trainable policy and is part of the reinforcement-learning objective. The paper’s example keeps GLM-4.7-Flash as the execution agent and changes the summary agent: its reported SWE-bench Verified pass@1 values are 55.5 with Qwen3.5-27B, 50.5 with GLM-4.7-Flash, and 49.0 with Qwen3-30B-A3B. These figures describe that fixed-execution comparison and do not establish a general ranking of summarizers.

CompactionRL trains summary and execution segments together inside a long-horizon rollout
CompactionRL reconstructs a bounded context from a generated summary and recent interaction, then trains all generated segments with the task reward. Read the CompactionRL paper.

Agent-specific learned compaction: ACON

ACON targets agent trajectories, where observations, tool calls, outcomes, and reasoning steps have different roles. It first uses failure analysis and natural-language prompt optimization to improve compression guidelines for a teacher model. It can then distill the resulting compressor into a smaller model. At inference time, the distilled compressor replaces the teacher, and greedy decoding is used in the paper’s setup.

The paper evaluates history and observation compression on AppWorld, OfficeBench, and 8-objective QA. It reports 26% to 54% lower peak token use and reports task-success changes for the tested baselines. The evaluation is controlled and the paper notes that compression itself adds computation. Those results do not show that every agent workload benefits in the same way.

ACON optimizes and distills a compressor for tool calls, observations, and agent history
ACON learns compression guidelines from agent failures and can distill them into a smaller compressor. Read the ACON paper for its benchmarks and inference settings.

Comparing the decisions

Here, inspectability means how directly a reader can examine the decision rule. A written rubric can be read and edited. A policy learned from trajectories is observable through its outputs, but its internal reason for each choice is harder to inspect.

MethodTraining or optimizationWhat is compressedDecision at inferenceInspectability
SelfCompactNone; task-specific rubricAccumulated historyModel chooses compress or continue at probe intervalsThe rubric and calls are visible
CompactionRLReinforcement learning over execution and summary segmentsHistory when the context budget is lowOne policy generates a summary and resumesLearned behavior is harder to inspect
ACONPrompt optimization, then optional compressor distillationAgent history and observationsA learned compressor applies optimized guidelinesGuidelines are visible; distilled choices are less direct

The practical choice follows the available evidence. A SelfCompact-style rubric is easy to trial when there is no training pipeline. CompactionRL requires rollout collection and reinforcement-learning infrastructure. ACON is useful to study when tool traces have structure that a generic text compressor does not represent. In each case, evaluate task success and the cost of the compression calls on the workload that matters.

These approaches focus on the current trajectory. Part 3 examines what can be recorded after a run so later work can start from evidence instead of replaying the whole transcript.

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →