Two accounts, one field name
Meerkat reports token usage in two places with the same field names and two different denominators.
A session may span providers and models, so the cumulative account deliberately
carries no single provider/model. Its detail counters are still comparable
across providers, because every adapter’s cache counts are subsets of that
call’s presented input:
- Anthropic reports cache-write and cache-read input as disjoint components of the presented total.
- OpenAI, Gemini and OpenAI-compatible backends report cached input as a detail inside their inclusive input total.
billing.usage, and mob run accounting.
Reads are clamped to input first and writes to what reads leave, so such a
session’s history reads as fully cached. The stored value is untouched until
the next recorded call normalizes it.
Reasoning is the provider’s thinking-token count: OpenAI Responses
output_tokens_details.reasoning_tokens, Chat Completions
completion_tokens_details.reasoning_tokens, and Gemini thoughtsTokenCount.
Gemini bills thinking separately from candidates, so its output_tokens is
candidates plus thoughts. Chat Completions backends differ: OpenAI counts
reasoning inside completion_tokens, while others such as xAI count it beside
them. A row whose total_tokens equals prompt plus completion plus reasoning
exactly, with non-zero reasoning, is read as reasoning outside completion, and
its output_tokens becomes completion plus reasoning. Anthropic reports no separate count and folds
thinking into output_tokens, so its reasoning_tokens stays absent.
Run results: per-run and per-request usage
The final run result (rkat run --output json, the RPC and REST run results,
the Python and TypeScript SDK RunResult and the @rkat/web TurnResult)
carries three usage views:
run_usage and the sum of request_usage[*].accounting.presented_tokens agree
on input for a run that records every call it makes. A run that suspends for
callback tool results keeps one account once its staged results are applied:
the next run of the agent continues it, whether that is run_pending or a
content turn carrying a new prompt, so calls made before the suspension appear
in that run’s result. A run with no applied results discards the suspended
account. The carry-over lives in the agent’s memory: a resume on an agent
rebuilt from storage starts a fresh account, and the pre-suspension calls
appear only in the session total. Compaction summaries written
by a host curator or the mechanical capacity fallback make no provider request
and add no row. Legacy totals are clamped as described above. On a fresh session
run_usage equals usage. A harness that runs several turns on one session
should read run_usage, never sum usage.
Which calls emit a usage row
An agent-loop call publishes its own per-call row when its turn completes, and structured-output extraction publishes one row per extraction request. A turn that fails after the provider answered publishes none, even if its assistant row was already committed to history before the failure. Session usage advances when a measured, chargeable outcome reaches the accounting path; measurement alone does not guarantee recording.run_completed.usage includes
only calls recorded before its emission.
turn_completed pairs with the turn_started of the same call. When a
compaction boundary sends the loop back to rebuild a request, the loop announces
the same turn_number again before the call, so pair a turn_completed with
the latest turn_started. Extraction publishes no turn_completed: its rows
ride the extraction outcome event, and an attempt that fails validation still
publishes its row there. Event logs written by earlier releases carry
turn_completed only for the call that closed each run.
Measured compaction summary usage is recorded once run_compaction returns an outcome,
even if the subsequent rewrite handoff or projection fails. A measured summary
rejected before an outcome is returned, such as an empty summary or a failed
rebuild validation, need not appear in session usage. The cumulative account is
therefore a record of charged outcomes, not a complete provider invoice.
For structured output, run_completed with extraction_required: true precedes
the extraction requests. Their measured usage can increase the final
RunResult.usage and authoritative session account, but no replacement
run_completed usage snapshot is emitted. Extraction outcome events do not
carry that updated total. Use the final result or an authoritative session
observation to include the extraction phase.
At a comparable observation point, the normalized sum of the measured rows
cannot exceed the session-cumulative account. For a run whose calls all
committed and that did no compaction, the run’s rows fold to exactly its
run_usage. Treat any residual as unattributed cost (compaction summaries,
failed turns, unmeasured calls), not a reconciliation error. Do not compare
later rows against an earlier completion snapshot.
Per-model attribution lives on the per-call event
turn_completed.usage.accounting is the single owner of attribution for that
call:
provider and model name the resolved model the request was actually lowered
to, not configuration intent, so a consumer reading a successfully projected
durable event audit log (.rkat/sessions/<id>/events.jsonl) can attribute the
calls it does see without joining against session metadata. The log is optional
derived state and may halt independently of session commits. presented_tokens is the provider-normalized
input presented to the model for that one call, and convention records how the
provider’s counters were normalized to produce it.
There is intentionally no second copy of the model or provider beside
accounting. Attribution has one owner.
accounting first appears on turn_completed in 0.8.22. Rows written by 0.8.21
and earlier carry a bare usage object with no attribution; recover the model
for those from session metadata.When a turn has no accounting at all
A provider stream can end without ever sending a usage event. From 0.8.24 that is a degradation, not a failure: the turn completes, its assistant message is committed, andturn_completed is published with usage absent rather than
with fabricated counters. A companion turn_usage_accounting_unmeasured event
names the provider and model that went unaccounted, carrying the operator marker
unmeasured:turn_usage_accounting from the same vocabulary runtime/health
uses.
An unaccounted turn moves nothing: it is not added to run_completed.usage, not
charged against the token budget, and does not reset the last-input-token figure
compaction reads. Consumers must skip an absent row. Treating it as zero
understates the session and, worse, publishes a measurement nobody made.
A related but distinct event, turn_usage_accounting_identity_disputed, means
the counters DID arrive but named a different provider/model than the request
they answered. Those counters are internally consistent, so they are recorded
exactly as reported and the axis advances normally; only the attribution is
contested, and it is never rewritten to the active identity.
OpenAI-compatible usage trailers
Chat Completions servers may emitfinish_reason first and send the
stream_options.include_usage counters in a later usage-only SSE event. The
OpenAI-compatible adapter therefore latches the stop reason instead of ending
the stream immediately. It continues reading until [DONE], stream end, or a
single post-finish trailer window expires. The standard window is 30 seconds,
and keepalive comments do not extend it.
Usage that arrives after finish_reason is retained and normalized normally.
If the connection closes, becomes undecodable, or reaches the trailer deadline
after the semantic finish, Meerkat treats that as stream end rather than
retrying an answer the caller already received. If no usage arrived, the
ordinary unmeasured-accounting contract above applies; no counters are minted.
Worked example
One session, two runs, with no extraction or other session activity. The first run makes three Anthropic calls (the first two request tools, the third closes the run); the second run makes one call.
What the stream publishes:
- The session total is the latest
run_completed.usage: 18920 tokens. The first run’s 14160 is the same account observed earlier, not a separate run’s cost. - What the turn rows attribute is 18420 presented input plus 500 output: all 18920 tokens. Every call committed and nothing else ran, so the four rows fold to the session total, and run 1’s three rows fold to its 14160.
turn_rows_cover_every_call_while_the_run_total_is_session_cumulative in
crates/meerkat-core/src/agent/usage_accounting_tests.rs, which drives the real agent
loop, and the CumulativeUsage arithmetic is pinned by
cumulative_usage_matches_documented_aggregation_example in
crates/meerkat-core/src/types/tests.rs. If this page and those tests disagree,
make docs-check fails.
What not to sum
1. Do not sumrun_completed.usage. Each one is already the recorded session
total at emission. For the session above, adding the two observed run totals
reports 14160 + 18920 = 33080 instead of 18920. This is the double-count the
field report hit. Use one snapshot at the desired observation point; never add
cumulative snapshots. A pre-extraction completion event is not the final
result’s usage.
2. Do not sum per-call input_tokens. Across the four turn rows above that
gives 1000 + 300 + 120 + 200 = 1620, against 18420 presented tokens, because
the raw Anthropic counter excludes cache-write and cache-read input. Sum
accounting.presented_tokens instead.
3. Do not compare total_tokens across the two accounts. On a per-call
usage, input_tokens + output_tokens uses the uncached denominator (210 for
call 3); the comparable per-call figure is
accounting.presented_tokens + output_tokens (4510). On the wire,
WireTurnUsage.total_tokens is already the normalized form, while
WireUsage.total_tokens built from a cumulative usage is the session total.
The two safe aggregations, using comparable observation points:
- Recorded session cost: the latest final
RunResult.usageor authoritative session usage observation. Nothing to add. Arun_completed.usagesnapshot suffices only if no later measured work, including extraction, has advanced the account; it is not unconditionally equal to the final result. - Per-model breakdown of the observable calls: sum
accounting.presented_tokensandoutput_tokensover theturn_completed.usagerows and the extraction events’request_usagerows, grouped byaccounting.model. Report any residual against the session total explicitly (compaction summaries and turns that failed after the provider answered publish no row); do not present this breakdown as the full cost.
SDK surfaces
Python and TypeScript expose one
Usage type for both the per-call and the
cumulative account, so accounting is typed optional there and is always absent
on the cumulative value. Rust keeps the two apart as TurnUsage and
CumulativeUsage.
The session.usage / session.text style accessors on the Python and
TypeScript SDK session objects read RunResult, so session.usage is the
session-cumulative account, not the last call’s usage.
Known limits
- The event stream carries no per-run token account; the final run result’s
run_usageis one. Without a run result, difference comparable final cumulative results or authoritative session observations before and after the run, including its extraction phase. Isolate or account for other activity on that same session, and report missing measurements separately rather than treating them as zero. Differencing pre-extractionrun_completed.usagesnapshots instead can charge one run’s extraction to the next interval. compaction_completed.summary_tokensis the output size of the compaction summary call and carries no provider/model attribution and no presented-input count, so compaction cost cannot be attributed from the event stream alone.- A call whose turn fails after the provider answered (a denied hook, an
exhausted budget) is charged but publishes no per-call row, and neither does
a compaction summary call;
RunResult.request_usagelists them. run_completed.usagecarries no per-model breakdown. For a session that switched models, group the per-call rows yourself; the cumulative total is model-agnostic on purpose.
