← BlogContext Compaction in LLM Agents

Context Compaction in LLM Agents: Part 1 - The Fundamentals

Isaac Kargar4 min read

  • LLM
  • Context Management
  • Production
  • Compaction

Agents accumulate instructions, conversation turns, tool results, and retrieved data as they work. A model’s context window sets the maximum amount of text that can be supplied in one request. That capacity is different from effective context use: a model may technically accept a long prompt while its accuracy or ability to find a relevant detail changes as the prompt grows.

The change is task-dependent. A long transcript can dilute a small but important instruction, while a carefully structured long document may remain usable. Context management is the process of deciding what to include, where to put it, and how to represent it for the next model call. Compaction is one part of that process. It reduces the amount of context while trying to retain information needed for the current task.

Four families of compaction

These methods make different tradeoffs. A summary rewrites text, a pruner selects part of the original text, and a learned compressor can be trained for a particular retrieval or question format. None of them guarantees that every detail survives or that quality improves without an evaluation on the intended task.

Semantic compression

Semantic compression asks a language model to encode a long passage or code description in a shorter natural-language representation. The representation aims to retain task-relevant meaning rather than reproduce every original character. The LLM semantic compression paper and the code-oriented semantic compression study evaluate whether a model can answer questions or regenerate code from that shorter representation.

Overview of semantic compression from a long input to a shorter task-oriented representation
Semantic compression rewrites an input into a shorter representation for a downstream task. See the code compression study for its measured tasks.

Loss-aware token pruning

Token pruning keeps a subset of the original prompt. One family scores tokens or spans by a language-model signal such as perplexity, then removes material under a token budget. The LongLLMLingua paper uses prompt compression and budget control, while LLMLingua’s documentation describes the implementation. The Perplexity-based prompt compression paper evaluates whether pruning can reduce context while keeping downstream performance close to the uncompressed prompt. These methods preserve selected source tokens, rather than producing a new summary of the whole input.

Loss-aware pruning selects the portions of a prompt that best fit a token budget
Loss-aware pruning selects source material using a model-based importance signal. Read the perplexity-based pruning paper for its quality and compression measurements.
LLMLingua prompt compression pipeline
LLMLingua applies token-level compression with a budget controller, so the retained prompt is still evaluated on the downstream task. Visit the LLMLingua documentation.

Question-aware sentence pruning

Some systems prune at the sentence or passage level after seeing the question. Provence trains a small pruner to remove sentences that are less useful for the query. This keeps selected evidence from a retrieved passage and does not ask a general-purpose summarizer to restate the entire context.

Provence uses a question-aware pruner to select useful sentences from retrieved passages
A query-aware pruner can remove irrelevant sentences while retaining passages that score as useful for the question. Read the Provence paper for the training and retrieval setup.

Dynamic summarization

Dynamic summarization maintains a rolling summary of older turns and keeps a recent window in full. When the conversation reaches a chosen budget, the system updates the summary with the next chunk of messages. This trades exact transcript access for a shorter representation that can be refreshed as the task changes.

LangChain’s short-term memory documentation shows this pattern, and LlamaIndex’s memory documentation describes summary-based memory as one option. A threshold such as 70% of the available token budget is an example configuration, not a general rule. A production system should measure when its summaries lose details that matter to its tasks.

Dynamic summarization keeps recent turns and updates a summary of older conversation
A rolling summary combines recent messages with a compact representation of older turns. Consult the LangChain short-term memory documentation for one implementation pattern.

The recursive dialogue summarization paper studies repeated updates of a summary as new dialogue arrives. It is an example of a summary that evolves over time, rather than a one-time truncation.

Recursive summarization updates a long-term dialogue summary as new turns arrive
Recursive summarization carries an earlier summary forward and incorporates a new dialogue chunk. Read the long-term dialogue memory paper.

The DTCRS approach is another query-aware variant. It first classifies the question, then builds a hierarchy of summaries only when the question needs that extra compression. Its reported results apply to the datasets and retrieval tasks in the DTCRS paper.

DTCRS builds a query-aware tree of summaries for multi-step retrieval questions
A query-aware summary tree can preserve different levels of detail for a multi-step question. See the DTCRS paper for the evaluated setting.

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →