Skip to main content

In plain terms

Long chats get too big. Before each model call, this shrinks the pile. This is contextCompactor.maxContextTokens — not a top-level Agent field.
History length for sessions is memory.maxMessages / memory.maxTokens. Compaction is a second pass inside the loop.

Configuration

Strategies

Trim

Drops oldest non-system messages first, keeping the system prompt and most recent exchanges intact.

Summarize

Uses a cheap model to summarize older messages into a single compact summary, preserving key context while reducing token count.

Hybrid

Trims first, then summarizes if still over budget. Best balance of speed and context preservation.

How It Works

The compactor hooks into the beforeLLMCall loop hook and runs before every LLM API call:
  1. Estimates token count of all messages
  2. If under budget, passes through unchanged
  3. If over budget, applies the configured strategy
  4. Returns the compacted messages to the LLM
System messages are always preserved. The most recent user/assistant exchanges are prioritized.