- Context compaction — automatically summarizes conversation history when the context window fills up, preserving recent turns and discarding older ones.
- Semantic memory — indexes discarded messages so the agent can retrieve past context on demand via the
memory_searchtool.
What this guide is for
Use this guide when you want to understand or configure:- long-horizon conversation behavior
- semantic recall via
memory_search - compaction thresholds
- the relationship between compaction and semantic memory
Feature flags
Compaction and semantic memory have separate build-time and runtime switches.The default CLI build includes memory backend/session wiring and effective
session compaction. Its default
mob feature pulls in meerkat-mob, which
enables session-compaction on the shared facade and session crates through
Cargo feature unification. The CLI’s own session-compaction alias is not
listed directly in its defaults. For no-default/reduced builds that omit this
dependency path, enable that alias explicitly when you need DefaultCompactor,
and enable the memory features when you need memory_search.AgentFactory::memory(true) enables the semantic memory path
(HnswMemoryStore + memory_search). The effective facade
session-compaction feature installs DefaultCompactor independently of that
memory toggle. Low-level core embeddings can instead supply a Compactor;
facade builds use the factory/config pipeline rather than direct
AgentBuilder::compactor injection.
- compaction enabled, semantic memory disabled
- semantic memory enabled, compaction unavailable
- both enabled together
CompactionConfig
Controls when and how compaction runs.How compaction triggers
Compaction is checked at every turn boundary, just before the next LLM call. The decision flow:1
Skip turn 0
The first turn always skips compaction.
2
Cadence check
Ordinary compaction attempts respect
min_turns_between_compactions.
Current-history or request-byte capacity recovery bypasses this cost guard.3
Threshold evaluation
Compaction triggers if ANY of:
last_input_tokens >= auto_compact_threshold(input tokens from the last LLM response), ORestimated_history_tokens >= auto_compact_threshold(a conservative current-history estimate: text and serialized structure use a bytes/4 heuristic, while images and video use media-aware estimates), OR- the measured request budget for this boundary crosses the threshold on its input side. This is the same measurement the pre-dispatch preflight performs - fully hydrated request messages, the exact visible tool definitions, and the effective output reserve - so tool schemas and hydrated media, which the provider counts but a transcript-only estimate cannot see, are visible to the trigger. It is available whenever the active model declares a context window, OR
- request bytes cross four-fifths of the effective byte cap. An exact
provider-lowered witness is preferred; otherwise an explicitly
configured
max_request_bytesis compared with the conservative transcript request estimate.
What happens during compaction
When compaction triggers:1
Emit CompactionStarted event
Emitted with input/estimated token counts and message count.
2
Produce the summary
Without a curator, a provider-safe projection of the current history plus
the compaction prompt is sent to the LLM with no tools and
max_summary_tokens as the response limit. Media and reasoning are
projected for summarization, and an oversized source projection is bounded
to a head/tail handoff excerpt. With a curator, this LLM call is skipped.3
Handle result
On ordinary failure: a CompactionFailed event is emitted and the session is not mutated (safe failure).If the summarization request itself exceeds provider capacity: the
built-in path can use a deterministic mechanical handoff summary with zero
provider usage for recognized capacity errors, but only when current
history has no protected, unsummarized assistant transcript observations
with source
SpokenUnmeasured. If those observations remain, the compaction
attempt emits CompactionFailed and preserves the original history rather
than mechanically discarding them. This fallback is not used for curator
failures.On success: DefaultCompactor::rebuild_history produces new messages:- Every unkeyed System message and the latest version of each keyed System prompt are preserved verbatim and in source order. Superseded keyed versions are discarded.
- A User-channel summary message is injected with the typed
CompactionSummarytranscript role and the rendered prefix[Context compacted]. - Up to
recent_turn_budgetcomplete turns are retained. The compactor shrinks that set when request-body pressure requires more room. - All other messages become
discarded.
4
Project discarded memory
If semantic memory is enabled, the runtime pairs the validated transcript
rewrite with a scoped memory batch. The in-memory store publishes the pair
synchronously; the durable HNSW store stages the batch under the exact
rewrite identity for session-runtime commit.
5
Commit the rewrite
The session messages become authoritative only after the memory projection
accepts the same rewrite. A store rejection, transcript validation error,
or runtime handoff refusal preserves the original history.
6
Record usage and emit completion event
Compaction usage is recorded against the session and budget. A CompactionCompleted event is emitted with summary token count and before/after message counts.
Host-supplied summary curation
AgentBuildConfig.compaction_curator_override accepts an
Arc<dyn CompactionCurator>. When present, the curator receives the typed
CompactionWindow and produces the summary instead of the summarization LLM
call. Summary usage is recorded as zero. The runtime does not fall back to an
LLM if the curator fails or returns an empty summary; it emits a typed
CompactionFailed event and preserves the original history.
The curator does not control trigger timing, retained/discarded provenance,
memory indexing, or transcript commit. Those remain runtime-owned and are
validated before mutation.
The compaction prompt
The compaction prompt
The compactor sends this prompt to the LLM:
You are performing a CONTEXT COMPACTION. Your job is to create a handoff summary so work can continue seamlessly. Include:Be concise and structured. Prioritize information the next context needs to act, not narrate.
- Current progress and key decisions made
- Important context, constraints, or user preferences discovered
- What remains to be done (clear next steps)
- Any critical data, file paths, examples, or references needed to continue
- Tool call patterns that worked or failed
Memory indexing after compaction
When both aCompactor and a MemoryStore are wired into the agent, discarded messages are indexed into semantic memory before compacted history is committed. If the memory store rejects indexing, Meerkat preserves the original history, emits a CompactionFailed event, and skips that compaction attempt instead of dropping the only authoritative copy of the discarded text.
For each discarded message:
- The producer carries
message.indexable_content()into a typedMemoryIndexRequest. The store owns the include/exclude decision; the producer does not use an empty-string convention. Compaction summaries and injected context are excluded from ordinary semantic recall. - The request carries
MemoryMetadatacontaining the session ID, the typed source handle (the offset range of the source message), and a timestamp.
memory_search tool.
The memory_search tool
When memory is enabled, the agent gains a memory_search tool.
Tool definition
Parameters
string
required
Natural language search query describing what you want to recall.
integer
default:"5"
Maximum number of results to return. Capped at 20.
Response format
Returns a JSON array of result objects:string
The text content of the memory entry.
number
Similarity score from 0.0 (no match) to 1.0 (exact match). Typical useful matches are above 0.7.
object
The half-open
[start, end) offset range of the source message(s) the entry
was indexed from. Memory is scoped to a single session; results do not carry a
session_id, and there is no cross-session recall.Memory store implementations
- HnswMemoryStore (production)
- SimpleMemoryStore (test-only)
Uses:
- hnsw_rs (v0.3) for approximate nearest-neighbor search with cosine distance.
- SQLite for persistent metadata and text storage.
{store_path}/memory/memory.sqlite3Key characteristics:- Embedding: Bag-of-words TF with hash-based dimensionality reduction (4096-dimensional vectors, L2-normalized). Each word is hashed to a bucket and its presence increments that dimension.
- Persistence: Data survives process restart.
open()runs a one-time in-place schema migration (indexedsession_idprojection column + durable point-ID allocator table, inside oneBEGIN IMMEDIATEtransaction) and an idempotent heal of legacy rows; it scans and embeds nothing. A scope’s HNSW graph is built lazily on first use (search / index / enumerate / drop), so opening the store no longer pays for every session in the realm. - Scoping: One HNSW index per session owner, built on demand from the scope’s durable rows.
- Lifecycle:
drop_scope(owner)deletes a scope’s durable rows all-or-nothing and drops its live index (dropped point IDs are never reused);enumerate_scoped(scope, request)pages raw scope rows in durable-id order with optionalsource_range-overlap andindexed_afterfilters (a zerolimitis rejected with a typed error). - Score conversion: HNSW cosine distance (0 = identical, 2 = opposite) is converted to a 0..1 similarity score:
score = 1.0 - (distance / 2.0). - Thread safety: Point IDs are allocated transactionally from the durable allocator table (never reused, collision-free across concurrent store instances). Insertions, scope drops, and lazy scope loads are serialized via a
Mutex; the scoped-index map sits behind aRwLockfor concurrent searches. - Parameters (
HnswParamsdefaults):max_nb_connection = 16,max_layer = 16,ef_construction = 200,ef_search = 200.
How memory gets wired
When thememory-store-session feature is compiled in and memory is enabled:
- An
HnswMemoryStoreis opened at{store_path}/memory/. - The
memory_searchtool is added to the agent’s tool set. - A
DefaultCompactoris attached only ifsession-compactionis also enabled. - The embedded
memory-retrievalcompanion skill is available in the skill inventory when its capability gate is satisfied. It is not automatically preloaded; callers or the agent can activate it through the normal typed skill paths.
Examples
- CLI
- SDK
Custom CompactionConfig
See also
- Configuration: memory and compaction - config file settings
- Architecture - how compaction fits into the agent loop
