Skip to main content
This page is the canonical contract for reading token usage off a Meerkat event stream: which number is per-call, which number is already a total, which calls the event stream does not report at all, how to attribute a call to the model that produced it, and what you must not sum.

Two accounts, one field name

Meerkat reports token usage in two places with the same field names and two different denominators. A session may span providers and models, so the cumulative account deliberately carries no single provider/model. Its detail counters are still comparable across providers, because every adapter’s cache counts are subsets of that call’s presented input:
  • Anthropic reports cache-write and cache-read input as disjoint components of the presented total.
  • OpenAI, Gemini and OpenAI-compatible backends report cached input as a detail inside their inclusive input total.
So on every cumulative value, on every provider:
A detail counter is absent until some recorded call reports it. Sessions saved by 0.8.22 through 0.8.42 carry no cumulative detail counters, so theirs start from zero and cover only later calls. Sessions saved before 0.8.22 carry raw summed cache counters that can exceed the input total. Every reporting surface clamps them to the invariant above: the run result, run and compaction rollback deltas, session read billing.usage, and mob run accounting. Reads are clamped to input first and writes to what reads leave, so such a session’s history reads as fully cached. The stored value is untouched until the next recorded call normalizes it. Reasoning is the provider’s thinking-token count: OpenAI Responses output_tokens_details.reasoning_tokens, Chat Completions completion_tokens_details.reasoning_tokens, and Gemini thoughtsTokenCount. Gemini bills thinking separately from candidates, so its output_tokens is candidates plus thoughts. Chat Completions backends differ: OpenAI counts reasoning inside completion_tokens, while others such as xAI count it beside them. A row whose total_tokens equals prompt plus completion plus reasoning exactly, with non-zero reasoning, is read as reasoning outside completion, and its output_tokens becomes completion plus reasoning. Anthropic reports no separate count and folds thinking into output_tokens, so its reasoning_tokens stays absent.
run_completed.usage is session-cumulative, not run-scoped. It snapshots Session::total_usage() at emission time. That account is persisted with the session and restored on resume, so on the second and later runs it already contains earlier runs’ recorded calls. Later calls can advance the account beyond this snapshot, notably structured-output extraction. There is no per-run token account on the event stream; the final run result carries one (see below).

Run results: per-run and per-request usage

The final run result (rkat run --output json, the RPC and REST run results, the Python and TypeScript SDK RunResult and the @rkat/web TurnResult) carries three usage views: run_usage and the sum of request_usage[*].accounting.presented_tokens agree on input for a run that records every call it makes. A run that suspends for callback tool results keeps one account once its staged results are applied: the next run of the agent continues it, whether that is run_pending or a content turn carrying a new prompt, so calls made before the suspension appear in that run’s result. A run with no applied results discards the suspended account. The carry-over lives in the agent’s memory: a resume on an agent rebuilt from storage starts a fresh account, and the pre-suspension calls appear only in the session total. Compaction summaries written by a host curator or the mechanical capacity fallback make no provider request and add no row. Legacy totals are clamped as described above. On a fresh session run_usage equals usage. A harness that runs several turns on one session should read run_usage, never sum usage.

Which calls emit a usage row

An agent-loop call publishes its own per-call row when its turn completes, and structured-output extraction publishes one row per extraction request. A turn that fails after the provider answered publishes none, even if its assistant row was already committed to history before the failure. Session usage advances when a measured, chargeable outcome reaches the accounting path; measurement alone does not guarantee recording. run_completed.usage includes only calls recorded before its emission. turn_completed pairs with the turn_started of the same call. When a compaction boundary sends the loop back to rebuild a request, the loop announces the same turn_number again before the call, so pair a turn_completed with the latest turn_started. Extraction publishes no turn_completed: its rows ride the extraction outcome event, and an attempt that fails validation still publishes its row there. Event logs written by earlier releases carry turn_completed only for the call that closed each run. Measured compaction summary usage is recorded once run_compaction returns an outcome, even if the subsequent rewrite handoff or projection fails. A measured summary rejected before an outcome is returned, such as an empty summary or a failed rebuild validation, need not appear in session usage. The cumulative account is therefore a record of charged outcomes, not a complete provider invoice. For structured output, run_completed with extraction_required: true precedes the extraction requests. Their measured usage can increase the final RunResult.usage and authoritative session account, but no replacement run_completed usage snapshot is emitted. Extraction outcome events do not carry that updated total. Use the final result or an authoritative session observation to include the extraction phase. At a comparable observation point, the normalized sum of the measured rows cannot exceed the session-cumulative account. For a run whose calls all committed and that did no compaction, the run’s rows fold to exactly its run_usage. Treat any residual as unattributed cost (compaction summaries, failed turns, unmeasured calls), not a reconciliation error. Do not compare later rows against an earlier completion snapshot.

Per-model attribution lives on the per-call event

turn_completed.usage.accounting is the single owner of attribution for that call:
provider and model name the resolved model the request was actually lowered to, not configuration intent, so a consumer reading a successfully projected durable event audit log (.rkat/sessions/<id>/events.jsonl) can attribute the calls it does see without joining against session metadata. The log is optional derived state and may halt independently of session commits. presented_tokens is the provider-normalized input presented to the model for that one call, and convention records how the provider’s counters were normalized to produce it. There is intentionally no second copy of the model or provider beside accounting. Attribution has one owner.
accounting first appears on turn_completed in 0.8.22. Rows written by 0.8.21 and earlier carry a bare usage object with no attribution; recover the model for those from session metadata.

When a turn has no accounting at all

A provider stream can end without ever sending a usage event. From 0.8.24 that is a degradation, not a failure: the turn completes, its assistant message is committed, and turn_completed is published with usage absent rather than with fabricated counters. A companion turn_usage_accounting_unmeasured event names the provider and model that went unaccounted, carrying the operator marker unmeasured:turn_usage_accounting from the same vocabulary runtime/health uses. An unaccounted turn moves nothing: it is not added to run_completed.usage, not charged against the token budget, and does not reset the last-input-token figure compaction reads. Consumers must skip an absent row. Treating it as zero understates the session and, worse, publishes a measurement nobody made. A related but distinct event, turn_usage_accounting_identity_disputed, means the counters DID arrive but named a different provider/model than the request they answered. Those counters are internally consistent, so they are recorded exactly as reported and the axis advances normally; only the attribution is contested, and it is never rewritten to the active identity.

OpenAI-compatible usage trailers

Chat Completions servers may emit finish_reason first and send the stream_options.include_usage counters in a later usage-only SSE event. The OpenAI-compatible adapter therefore latches the stop reason instead of ending the stream immediately. It continues reading until [DONE], stream end, or a single post-finish trailer window expires. The standard window is 30 seconds, and keepalive comments do not extend it. Usage that arrives after finish_reason is retained and normalized normally. If the connection closes, becomes undecodable, or reaches the trailer deadline after the semantic finish, Meerkat treats that as stream end rather than retrying an answer the caller already received. If no usage arrived, the ordinary unmeasured-accounting contract above applies; no counters are minted.

Worked example

One session, two runs, with no extraction or other session activity. The first run makes three Anthropic calls (the first two request tools, the third closes the run); the second run makes one call. What the stream publishes:
For this ordinary tool-loop example, read that as follows.
  • The session total is the latest run_completed.usage: 18920 tokens. The first run’s 14160 is the same account observed earlier, not a separate run’s cost.
  • What the turn rows attribute is 18420 presented input plus 500 output: all 18920 tokens. Every call committed and nothing else ran, so the four rows fold to the session total, and run 1’s three rows fold to its 14160.
These numbers are pinned by turn_rows_cover_every_call_while_the_run_total_is_session_cumulative in crates/meerkat-core/src/agent/usage_accounting_tests.rs, which drives the real agent loop, and the CumulativeUsage arithmetic is pinned by cumulative_usage_matches_documented_aggregation_example in crates/meerkat-core/src/types/tests.rs. If this page and those tests disagree, make docs-check fails.

What not to sum

Three specific mistakes, all of which produce a wrong number silently.
1. Do not sum run_completed.usage. Each one is already the recorded session total at emission. For the session above, adding the two observed run totals reports 14160 + 18920 = 33080 instead of 18920. This is the double-count the field report hit. Use one snapshot at the desired observation point; never add cumulative snapshots. A pre-extraction completion event is not the final result’s usage. 2. Do not sum per-call input_tokens. Across the four turn rows above that gives 1000 + 300 + 120 + 200 = 1620, against 18420 presented tokens, because the raw Anthropic counter excludes cache-write and cache-read input. Sum accounting.presented_tokens instead. 3. Do not compare total_tokens across the two accounts. On a per-call usage, input_tokens + output_tokens uses the uncached denominator (210 for call 3); the comparable per-call figure is accounting.presented_tokens + output_tokens (4510). On the wire, WireTurnUsage.total_tokens is already the normalized form, while WireUsage.total_tokens built from a cumulative usage is the session total. The two safe aggregations, using comparable observation points:
  • Recorded session cost: the latest final RunResult.usage or authoritative session usage observation. Nothing to add. A run_completed.usage snapshot suffices only if no later measured work, including extraction, has advanced the account; it is not unconditionally equal to the final result.
  • Per-model breakdown of the observable calls: sum accounting.presented_tokens and output_tokens over the turn_completed.usage rows and the extraction events’ request_usage rows, grouped by accounting.model. Report any residual against the session total explicitly (compaction summaries and turns that failed after the provider answered publish no row); do not present this breakdown as the full cost.

SDK surfaces

Python and TypeScript expose one Usage type for both the per-call and the cumulative account, so accounting is typed optional there and is always absent on the cumulative value. Rust keeps the two apart as TurnUsage and CumulativeUsage. The session.usage / session.text style accessors on the Python and TypeScript SDK session objects read RunResult, so session.usage is the session-cumulative account, not the last call’s usage.

Known limits

  • The event stream carries no per-run token account; the final run result’s run_usage is one. Without a run result, difference comparable final cumulative results or authoritative session observations before and after the run, including its extraction phase. Isolate or account for other activity on that same session, and report missing measurements separately rather than treating them as zero. Differencing pre-extraction run_completed.usage snapshots instead can charge one run’s extraction to the next interval.
  • compaction_completed.summary_tokens is the output size of the compaction summary call and carries no provider/model attribution and no presented-input count, so compaction cost cannot be attributed from the event stream alone.
  • A call whose turn fails after the provider answered (a denied hook, an exhausted budget) is charged but publishes no per-call row, and neither does a compaction summary call; RunResult.request_usage lists them.
  • run_completed.usage carries no per-model breakdown. For a session that switched models, group the per-call rows yourself; the cumulative total is model-agnostic on purpose.