← BlogContext Compaction in LLM Agents

Context Compaction in LLM Agents: Part 4 - Theory and Safety

Isaac Kargar6 min read

  • LLM
  • Context Management
  • Compaction
  • AI Safety

The earlier posts covered semantic compression, pruning, dynamic summaries, learned compaction, and post-hoc compilation. They operate at different layers and use different units. A common way to compare them is to ask what information is retained, at what fidelity, under which resource budget.

A rate-distortion lens

The survey What to Keep, What to Forget: A Rate-Distortion View of Memory Compaction in LLMs and Agents, by Colaco and Lahjouji, proposes a seven-axis taxonomy for describing these choices:

  1. Layer: attention, prompt, model architecture, or agent workflow
  2. Granularity: tokens, sentences, turns, or sessions
  3. Signal: attention, perplexity, recency, or task relevance
  4. Timing: before inference, during a trajectory, or after a run
  5. Reversibility: whether discarded information can be recovered
  6. Adaptivity: whether the policy is fixed or learned
  7. Budget: tokens, memory, latency, or monetary cost

This taxonomy is a comparison framework from the survey. It does not show that every method has the same objective or that one axis predicts quality on every task. It helps make a design choice explicit. For example, a token pruner is often hard to reverse, while a post-hoc record can keep a source reference; a fixed recency rule is inspectable, while a learned policy may require behavioral tests to understand.

Seven-axis rate-distortion taxonomy for context compaction methods
The survey organizes compaction methods by layer, granularity, signal, timing, reversibility, adaptivity, and budget. Read the rate-distortion survey for the taxonomy and its scope.

The survey also identifies open comparison problems. Results often come from different tasks and budgets, and repeated compaction across several sessions is less studied than a single compression step. Those observations motivate measuring the actual workflow rather than assuming that a result transfers between layers.

Governance can be lost during compaction

Governance Decay, by Shiyang Chen, studies a specific threat model. A policy is delivered through a user turn, memory entry, or tool output. It is not an immutable system or developer message that the harness promises to preserve. The session then grows until the harness compacts the history, and a later request would violate that policy if it is no longer present.

The paper’s ConstraintRot benchmark contains nine tasks, three repetitions per model and condition, and 1,323 episodes in the main grid. It evaluates seven model families: DeepSeek-V4-Flash, GLM-5.1, Qwen3.6-27B, Kimi-K2.5, Claude-Sonnet-4.6, GPT-5.4-mini, and Gemini-3.5-flash. The terminal tool call is graded deterministically, so the violation metric is whether the prohibited effect appears in the action.

With the policy in the full context, the control condition recorded 0% violations for every model. After one compaction step, the pooled violation rate was 30%, ranging from 0% for GLM-5.1 to 59% for DeepSeek-V4 and Kimi-K2.5. The paper reports that when an independent judge found the constraint still present in a compacted context, violations were 0% in 90 cases; when the judge found it missing, violations were 38% in 315 cases.

ConditionDeepSeek-V4GLM-5.1Qwen3.6Kimi-K2.5Claude-4.6GPT-5.4-miniGemini-3.5-flashAll models
No policy (floor)37%44%48%48%33%56%63%47%
Policy in full context (control)0%0%0%0%0%0%0%0%
Policy then compaction59%0%30%59%19%41%4%30%
Volume attack48%7%26%44%0%37%19%26%
Fixed injection attack59%22%33%41%0%37%0%28%
Constraint pinning0%0%0%0%0%0%0%0%
Constraint pinning with fixed attack0%0%0%0%0%0%0%0%

Violation rate by condition and model in the paper’s main grid. Each cell has nine tasks and three repetitions (n = 27); the grid contains 1,323 episodes. Values are percentages from the Governance Decay paper.

Governance Decay results for compaction, attacks, and constraint pinning
The chart shows the benchmark’s per-model violation rates under compaction and attacks, with the control and pinning conditions as the comparison points. Read the Governance Decay paper for the threat model and repetitions.

Compaction-eviction attacks

The paper tests two ways an adversary that controls only ingested context can bias this benchmark. A volume attack supplies enough content to force or crowd the policy out of the summary. A fixed injection places an instruction in that content aimed at the summarizer. These attacks target the memory-management step, rather than changing the model or system prompt. They show a failure in this tested setup; they are not evidence that every compaction implementation has the same rates.

The paper also searches several injection strategies. On a three-model, five-soft-task evaluation, its optimized strategy produced violation rates of 100% for DeepSeek, 85% for GLM, and 65% for Claude. Those numbers describe the paper’s selected attack search and should be interpreted with its model set and task scope.

Constraint pinning

Constraint Pinning extracts a governance rule into a separate buffer, re-injects it after compaction, and checks that the post-compaction context still entails the rule. In the main benchmark condition, no violations were observed with pinning, including the fixed attack. That is a result for the evaluated tasks, models, and implementation rather than a universal zero-error guarantee.

The paper’s stress test shows the boundary. An operator-impersonation message in the recent context raised violations from 0% to 17% with naive pinning and to 10% with an explicit provenance instruction. A trusted channel outside the token stream would be needed to distinguish a genuine operator update from an in-context impersonation.

What a production design can learn

The benchmark supports a concrete design question: which constraints are allowed to be compacted, and which must remain protected state? A system could combine online compaction for task history, post-hoc records with source references, and a pinned buffer for explicit governance rules. That is a recommended design to test against the application’s threat model, not a universally proven best stack.

The paper also has limitations. It uses simulated tool calls, inference-only evaluation, modest repetitions, and policies that are explicit enough to extract. Preserving a system message does not automatically protect a policy carried in a memory or tool channel. A real deployment should test its own channels, compaction strategy, allowed actions, and audit trail.

Across this series, compaction is a resource and information-management decision. The useful question is not whether a method keeps an abstract essence. It is whether the retained representation preserves the specific state, evidence, and constraints that the next action requires, under a measured budget.

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →