Context Compaction in LLM Agents: Part 4 - Theory and Safety
The earlier posts covered semantic compression, pruning, dynamic summaries, learned compaction, and post-hoc compilation. They operate at different layers and use different units. A common way to compare them is to ask what information is retained, at what fidelity, under which resource budget.
A rate-distortion lens
The survey What to Keep, What to Forget: A Rate-Distortion View of Memory Compaction in LLMs and Agents, by Colaco and Lahjouji, proposes a seven-axis taxonomy for describing these choices:
- Layer: attention, prompt, model architecture, or agent workflow
- Granularity: tokens, sentences, turns, or sessions
- Signal: attention, perplexity, recency, or task relevance
- Timing: before inference, during a trajectory, or after a run
- Reversibility: whether discarded information can be recovered
- Adaptivity: whether the policy is fixed or learned
- Budget: tokens, memory, latency, or monetary cost
This taxonomy is a comparison framework from the survey. It does not show that every method has the same objective or that one axis predicts quality on every task. It helps make a design choice explicit. For example, a token pruner is often hard to reverse, while a post-hoc record can keep a source reference; a fixed recency rule is inspectable, while a learned policy may require behavioral tests to understand.

The survey also identifies open comparison problems. Results often come from different tasks and budgets, and repeated compaction across several sessions is less studied than a single compression step. Those observations motivate measuring the actual workflow rather than assuming that a result transfers between layers.
Governance can be lost during compaction
Governance Decay, by Shiyang Chen, studies a specific threat model. A policy is delivered through a user turn, memory entry, or tool output. It is not an immutable system or developer message that the harness promises to preserve. The session then grows until the harness compacts the history, and a later request would violate that policy if it is no longer present.
The paper’s ConstraintRot benchmark contains nine tasks, three repetitions per model and condition, and 1,323 episodes in the main grid. It evaluates seven model families: DeepSeek-V4-Flash, GLM-5.1, Qwen3.6-27B, Kimi-K2.5, Claude-Sonnet-4.6, GPT-5.4-mini, and Gemini-3.5-flash. The terminal tool call is graded deterministically, so the violation metric is whether the prohibited effect appears in the action.
With the policy in the full context, the control condition recorded 0% violations for every model. After one compaction step, the pooled violation rate was 30%, ranging from 0% for GLM-5.1 to 59% for DeepSeek-V4 and Kimi-K2.5. The paper reports that when an independent judge found the constraint still present in a compacted context, violations were 0% in 90 cases; when the judge found it missing, violations were 38% in 315 cases.
| Condition | DeepSeek-V4 | GLM-5.1 | Qwen3.6 | Kimi-K2.5 | Claude-4.6 | GPT-5.4-mini | Gemini-3.5-flash | All models |
|---|---|---|---|---|---|---|---|---|
| No policy (floor) | 37% | 44% | 48% | 48% | 33% | 56% | 63% | 47% |
| Policy in full context (control) | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| Policy then compaction | 59% | 0% | 30% | 59% | 19% | 41% | 4% | 30% |
| Volume attack | 48% | 7% | 26% | 44% | 0% | 37% | 19% | 26% |
| Fixed injection attack | 59% | 22% | 33% | 41% | 0% | 37% | 0% | 28% |
| Constraint pinning | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| Constraint pinning with fixed attack | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
Violation rate by condition and model in the paper’s main grid. Each cell has nine tasks and three repetitions (n = 27); the grid contains 1,323 episodes. Values are percentages from the Governance Decay paper.

Compaction-eviction attacks
The paper tests two ways an adversary that controls only ingested context can bias this benchmark. A volume attack supplies enough content to force or crowd the policy out of the summary. A fixed injection places an instruction in that content aimed at the summarizer. These attacks target the memory-management step, rather than changing the model or system prompt. They show a failure in this tested setup; they are not evidence that every compaction implementation has the same rates.
The paper also searches several injection strategies. On a three-model, five-soft-task evaluation, its optimized strategy produced violation rates of 100% for DeepSeek, 85% for GLM, and 65% for Claude. Those numbers describe the paper’s selected attack search and should be interpreted with its model set and task scope.
Constraint pinning
Constraint Pinning extracts a governance rule into a separate buffer, re-injects it after compaction, and checks that the post-compaction context still entails the rule. In the main benchmark condition, no violations were observed with pinning, including the fixed attack. That is a result for the evaluated tasks, models, and implementation rather than a universal zero-error guarantee.
The paper’s stress test shows the boundary. An operator-impersonation message in the recent context raised violations from 0% to 17% with naive pinning and to 10% with an explicit provenance instruction. A trusted channel outside the token stream would be needed to distinguish a genuine operator update from an in-context impersonation.
What a production design can learn
The benchmark supports a concrete design question: which constraints are allowed to be compacted, and which must remain protected state? A system could combine online compaction for task history, post-hoc records with source references, and a pinned buffer for explicit governance rules. That is a recommended design to test against the application’s threat model, not a universally proven best stack.
The paper also has limitations. It uses simulated tool calls, inference-only evaluation, modest repetitions, and policies that are explicit enough to extract. Preserving a system message does not automatically protect a policy carried in a memory or tool channel. A real deployment should test its own channels, compaction strategy, allowed actions, and audit trail.
Across this series, compaction is a resource and information-management decision. The useful question is not whether a method keeps an abstract essence. It is whether the retained representation preserves the specific state, evidence, and constraints that the next action requires, under a measured budget.
Work with Nazmi
Build your AI system with Nazmi.
Tell us what you are building, what exists today, and where your team needs help.
Start a conversation or book a 20-minute call →