Skip to main content
Two systems handle conversations that exceed model context limits:
  1. Context compaction — automatically summarizes conversation history when the context window fills up, preserving recent turns and discarding older ones.
  2. Semantic memory — indexes discarded messages so the agent can retrieve past context on demand via the memory_search tool.
Together they allow agents to maintain coherent multi-turn sessions that exceed any single model’s context limit.

What this guide is for

Use this guide when you want to understand or configure:
  • long-horizon conversation behavior
  • semantic recall via memory_search
  • compaction thresholds
  • the relationship between compaction and semantic memory

Feature flags

Compaction and semantic memory have separate build-time and runtime switches.
The default CLI build includes memory backend/session wiring and effective session compaction. Its default mob feature pulls in meerkat-mob, which enables session-compaction on the shared facade and session crates through Cargo feature unification. The CLI’s own session-compaction alias is not listed directly in its defaults. For no-default/reduced builds that omit this dependency path, enable that alias explicitly when you need DefaultCompactor, and enable the memory features when you need memory_search.
At runtime, AgentFactory::memory(true) enables the semantic memory path (HnswMemoryStore + memory_search). The effective facade session-compaction feature installs DefaultCompactor independently of that memory toggle. Low-level core embeddings can instead supply a Compactor; facade builds use the factory/config pipeline rather than direct AgentBuilder::compactor injection.
That means these cases are all valid:
  • compaction enabled, semantic memory disabled
  • semantic memory enabled, compaction unavailable
  • both enabled together

CompactionConfig

Controls when and how compaction runs.

How compaction triggers

Compaction is checked at every turn boundary, just before the next LLM call. The decision flow:
1

Skip turn 0

The first turn always skips compaction.
2

Cadence check

Ordinary compaction attempts respect min_turns_between_compactions. Current-history or request-byte capacity recovery bypasses this cost guard.
3

Threshold evaluation

Compaction triggers if ANY of:
  • last_input_tokens >= auto_compact_threshold (input tokens from the last LLM response), OR
  • estimated_history_tokens >= auto_compact_threshold (a conservative current-history estimate: text and serialized structure use a bytes/4 heuristic, while images and video use media-aware estimates), OR
  • the measured request budget for this boundary crosses the threshold on its input side. This is the same measurement the pre-dispatch preflight performs - fully hydrated request messages, the exact visible tool definitions, and the effective output reserve - so tool schemas and hydrated media, which the provider counts but a transcript-only estimate cannot see, are visible to the trigger. It is available whenever the active model declares a context window, OR
  • request bytes cross four-fifths of the effective byte cap. An exact provider-lowered witness is preferred; otherwise an explicitly configured max_request_bytes is compared with the conservative transcript request estimate.

What happens during compaction

When compaction triggers:
1

Emit CompactionStarted event

Emitted with input/estimated token counts and message count.
2

Produce the summary

Without a curator, a provider-safe projection of the current history plus the compaction prompt is sent to the LLM with no tools and max_summary_tokens as the response limit. Media and reasoning are projected for summarization, and an oversized source projection is bounded to a head/tail handoff excerpt. With a curator, this LLM call is skipped.
3

Handle result

On ordinary failure: a CompactionFailed event is emitted and the session is not mutated (safe failure).If the summarization request itself exceeds provider capacity: the built-in path can use a deterministic mechanical handoff summary with zero provider usage for recognized capacity errors, but only when current history has no protected, unsummarized assistant transcript observations with source SpokenUnmeasured. If those observations remain, the compaction attempt emits CompactionFailed and preserves the original history rather than mechanically discarding them. This fallback is not used for curator failures.On success: DefaultCompactor::rebuild_history produces new messages:
  • Every unkeyed System message and the latest version of each keyed System prompt are preserved verbatim and in source order. Superseded keyed versions are discarded.
  • A User-channel summary message is injected with the typed CompactionSummary transcript role and the rendered prefix [Context compacted].
  • Up to recent_turn_budget complete turns are retained. The compactor shrinks that set when request-body pressure requires more room.
  • All other messages become discarded.
4

Project discarded memory

If semantic memory is enabled, the runtime pairs the validated transcript rewrite with a scoped memory batch. The in-memory store publishes the pair synchronously; the durable HNSW store stages the batch under the exact rewrite identity for session-runtime commit.
5

Commit the rewrite

The session messages become authoritative only after the memory projection accepts the same rewrite. A store rejection, transcript validation error, or runtime handoff refusal preserves the original history.
6

Record usage and emit completion event

Compaction usage is recorded against the session and budget. A CompactionCompleted event is emitted with summary token count and before/after message counts.

Host-supplied summary curation

AgentBuildConfig.compaction_curator_override accepts an Arc<dyn CompactionCurator>. When present, the curator receives the typed CompactionWindow and produces the summary instead of the summarization LLM call. Summary usage is recorded as zero. The runtime does not fall back to an LLM if the curator fails or returns an empty summary; it emits a typed CompactionFailed event and preserves the original history. The curator does not control trigger timing, retained/discarded provenance, memory indexing, or transcript commit. Those remain runtime-owned and are validated before mutation.
The compactor sends this prompt to the LLM:
You are performing a CONTEXT COMPACTION. Your job is to create a handoff summary so work can continue seamlessly. Include:
  • Current progress and key decisions made
  • Important context, constraints, or user preferences discovered
  • What remains to be done (clear next steps)
  • Any critical data, file paths, examples, or references needed to continue
  • Tool call patterns that worked or failed
Be concise and structured. Prioritize information the next context needs to act, not narrate.

Memory indexing after compaction

When both a Compactor and a MemoryStore are wired into the agent, discarded messages are indexed into semantic memory before compacted history is committed. If the memory store rejects indexing, Meerkat preserves the original history, emits a CompactionFailed event, and skips that compaction attempt instead of dropping the only authoritative copy of the discarded text. For each discarded message:
  • The producer carries message.indexable_content() into a typed MemoryIndexRequest. The store owns the include/exclude decision; the producer does not use an empty-string convention. Compaction summaries and injected context are excluded from ordinary semantic recall.
  • The request carries MemoryMetadata containing the session ID, the typed source handle (the offset range of the source message), and a timestamp.
This means previously discarded conversation content becomes searchable via the memory_search tool.

The memory_search tool

When memory is enabled, the agent gains a memory_search tool.

Tool definition

Parameters

string
required
Natural language search query describing what you want to recall.
integer
default:"5"
Maximum number of results to return. Capped at 20.

Response format

Returns a JSON array of result objects:
string
The text content of the memory entry.
number
Similarity score from 0.0 (no match) to 1.0 (exact match). Typical useful matches are above 0.7.
object
The half-open [start, end) offset range of the source message(s) the entry was indexed from. Memory is scoped to a single session; results do not carry a session_id, and there is no cross-session recall.

Memory store implementations

Uses:
  • hnsw_rs (v0.3) for approximate nearest-neighbor search with cosine distance.
  • SQLite for persistent metadata and text storage.
Storage layout: {store_path}/memory/memory.sqlite3Key characteristics:
  • Embedding: Bag-of-words TF with hash-based dimensionality reduction (4096-dimensional vectors, L2-normalized). Each word is hashed to a bucket and its presence increments that dimension.
  • Persistence: Data survives process restart. open() runs a one-time in-place schema migration (indexed session_id projection column + durable point-ID allocator table, inside one BEGIN IMMEDIATE transaction) and an idempotent heal of legacy rows; it scans and embeds nothing. A scope’s HNSW graph is built lazily on first use (search / index / enumerate / drop), so opening the store no longer pays for every session in the realm.
  • Scoping: One HNSW index per session owner, built on demand from the scope’s durable rows.
  • Lifecycle: drop_scope(owner) deletes a scope’s durable rows all-or-nothing and drops its live index (dropped point IDs are never reused); enumerate_scoped(scope, request) pages raw scope rows in durable-id order with optional source_range-overlap and indexed_after filters (a zero limit is rejected with a typed error).
  • Score conversion: HNSW cosine distance (0 = identical, 2 = opposite) is converted to a 0..1 similarity score: score = 1.0 - (distance / 2.0).
  • Thread safety: Point IDs are allocated transactionally from the durable allocator table (never reused, collision-free across concurrent store instances). Insertions, scope drops, and lazy scope loads are serialized via a Mutex; the scoped-index map sits behind a RwLock for concurrent searches.
  • Parameters (HnswParams defaults): max_nb_connection = 16, max_layer = 16, ef_construction = 200, ef_search = 200.

How memory gets wired

When the memory-store-session feature is compiled in and memory is enabled:
  1. An HnswMemoryStore is opened at {store_path}/memory/.
  2. The memory_search tool is added to the agent’s tool set.
  3. A DefaultCompactor is attached only if session-compaction is also enabled.
  4. The embedded memory-retrieval companion skill is available in the skill inventory when its capability gate is satisfied. It is not automatically preloaded; callers or the agent can activate it through the normal typed skill paths.

Examples

Custom CompactionConfig

See also