Context Compaction in LLM Agents: Part 1 - The Fundamentals
Agents accumulate instructions, conversation turns, tool results, and retrieved data as they work. A model’s context window sets the maximum amount of text that can be supplied in one request. That capacity is different from effective context use: a model may technically accept a long prompt while its accuracy or ability to find a relevant detail changes as the prompt grows.
The change is task-dependent. A long transcript can dilute a small but important instruction, while a carefully structured long document may remain usable. Context management is the process of deciding what to include, where to put it, and how to represent it for the next model call. Compaction is one part of that process. It reduces the amount of context while trying to retain information needed for the current task.
Four families of compaction
These methods make different tradeoffs. A summary rewrites text, a pruner selects part of the original text, and a learned compressor can be trained for a particular retrieval or question format. None of them guarantees that every detail survives or that quality improves without an evaluation on the intended task.
Semantic compression
Semantic compression asks a language model to encode a long passage or code description in a shorter natural-language representation. The representation aims to retain task-relevant meaning rather than reproduce every original character. The LLM semantic compression paper and the code-oriented semantic compression study evaluate whether a model can answer questions or regenerate code from that shorter representation.

Loss-aware token pruning
Token pruning keeps a subset of the original prompt. One family scores tokens or spans by a language-model signal such as perplexity, then removes material under a token budget. The LongLLMLingua paper uses prompt compression and budget control, while LLMLingua’s documentation describes the implementation. The Perplexity-based prompt compression paper evaluates whether pruning can reduce context while keeping downstream performance close to the uncompressed prompt. These methods preserve selected source tokens, rather than producing a new summary of the whole input.


Question-aware sentence pruning
Some systems prune at the sentence or passage level after seeing the question. Provence trains a small pruner to remove sentences that are less useful for the query. This keeps selected evidence from a retrieved passage and does not ask a general-purpose summarizer to restate the entire context.

Dynamic summarization
Dynamic summarization maintains a rolling summary of older turns and keeps a recent window in full. When the conversation reaches a chosen budget, the system updates the summary with the next chunk of messages. This trades exact transcript access for a shorter representation that can be refreshed as the task changes.
LangChain’s short-term memory documentation shows this pattern, and LlamaIndex’s memory documentation describes summary-based memory as one option. A threshold such as 70% of the available token budget is an example configuration, not a general rule. A production system should measure when its summaries lose details that matter to its tasks.

The recursive dialogue summarization paper studies repeated updates of a summary as new dialogue arrives. It is an example of a summary that evolves over time, rather than a one-time truncation.

The DTCRS approach is another query-aware variant. It first classifies the question, then builds a hierarchy of summaries only when the question needs that extra compression. Its reported results apply to the datasets and retrieval tasks in the DTCRS paper.

Work with Nazmi
Build your AI system with Nazmi.
Tell us what you are building, what exists today, and where your team needs help.
Start a conversation or book a 20-minute call →