Context Compaction in LLM Agents: Part 2 - Learning to Compact
In Part 1, we separated semantic compression, token or sentence pruning, and dynamic summarization. This post looks at three papers that add a learned or model-guided decision to that process. Their training requirements and evaluation settings differ, so their numbers should be read in those settings.
Inference-time self-compaction: SelfCompact
SelfCompact gives a model a compaction tool and a task-specific rubric. At probe intervals, the model decides whether a subtask has finished or a trajectory is converging. It can suppress compaction during a derivation or when the agent is stuck. The method does not fine-tune the model.
The paper evaluates four Qwen configurations on IMO-Answerbench, HMMT November 2025, and HMMT February 2026. Under token budgets matched to fixed-interval summaries, SelfCompact has the best score in 11 of 12 benchmark and model cells. For Qwen3.5-9B with thinking disabled, its score on HMMT February 2026 is 52.3%, compared with 34.2% for no compaction, a difference of 18.1 percentage points.
For a separate agentic-search evaluation, the paper uses GLM-4.7-Flash, MiniMax-M2.5, and Mimo-V2-Flash on BrowseComp, BrowseComp Plus, and DeepSearch QA. It reports gains of 5 to 9 accuracy points and 30% to 70% lower per-question token cost than its no-compaction baseline. The unit is per-question cost in that experiment, not a universal token-saving rate.

Reinforcement-learning compaction: CompactionRL
CompactionRL puts compaction inside the rollout. A single trainable policy generates ordinary execution actions and, when the context budget is nearly full, a summary. The rollout resumes from that summary plus a short recent tail. The final task reward is assigned to both execution and summary segments. The paper does not add a separate human-readable summary reward; it propagates credit across segments with a cross-trajectory advantage calculation.
This is different from a fixed summarizer paired with a separate execution model. The summary is sampled from the trainable policy and is part of the reinforcement-learning objective. The paper’s example keeps GLM-4.7-Flash as the execution agent and changes the summary agent: its reported SWE-bench Verified pass@1 values are 55.5 with Qwen3.5-27B, 50.5 with GLM-4.7-Flash, and 49.0 with Qwen3-30B-A3B. These figures describe that fixed-execution comparison and do not establish a general ranking of summarizers.

Agent-specific learned compaction: ACON
ACON targets agent trajectories, where observations, tool calls, outcomes, and reasoning steps have different roles. It first uses failure analysis and natural-language prompt optimization to improve compression guidelines for a teacher model. It can then distill the resulting compressor into a smaller model. At inference time, the distilled compressor replaces the teacher, and greedy decoding is used in the paper’s setup.
The paper evaluates history and observation compression on AppWorld, OfficeBench, and 8-objective QA. It reports 26% to 54% lower peak token use and reports task-success changes for the tested baselines. The evaluation is controlled and the paper notes that compression itself adds computation. Those results do not show that every agent workload benefits in the same way.

Comparing the decisions
Here, inspectability means how directly a reader can examine the decision rule. A written rubric can be read and edited. A policy learned from trajectories is observable through its outputs, but its internal reason for each choice is harder to inspect.
| Method | Training or optimization | What is compressed | Decision at inference | Inspectability |
|---|---|---|---|---|
| SelfCompact | None; task-specific rubric | Accumulated history | Model chooses compress or continue at probe intervals | The rubric and calls are visible |
| CompactionRL | Reinforcement learning over execution and summary segments | History when the context budget is low | One policy generates a summary and resumes | Learned behavior is harder to inspect |
| ACON | Prompt optimization, then optional compressor distillation | Agent history and observations | A learned compressor applies optimized guidelines | Guidelines are visible; distilled choices are less direct |
The practical choice follows the available evidence. A SelfCompact-style rubric is easy to trial when there is no training pipeline. CompactionRL requires rollout collection and reinforcement-learning infrastructure. ACON is useful to study when tool traces have structure that a generic text compressor does not represent. In each case, evaluate task success and the cost of the compression calls on the workload that matters.
These approaches focus on the current trajectory. Part 3 examines what can be recorded after a run so later work can start from evidence instead of replaying the whole transcript.
Work with Nazmi
Build your AI system with Nazmi.
Tell us what you are building, what exists today, and where your team needs help.
Start a conversation or book a 20-minute call →