> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Cost Tracking

> Track token costs and price estimates per session across all modes, with real API data, budgets, persistence and /cost subcommands

ChatCLI's **Cost Tracking** monitors token consumption and estimates costs in real time during your sessions, using real API usage data from every provider that reports it. You can inspect the live session with `/cost`, close and restart the accounting period with `/cost reset`, review past sessions with `/cost last` and `/cost sessions`, export machine-readable snapshots with `/cost export`, and configure spending limits — including an optional hard stop.

***

## The /cost Command

`/cost` renders the live session summary; four subcommands manage the accounting period:

| Command | What it does |
| :- | :- |
| `/cost` | Live session summary: tokens, cache, per-model costs, budget |
| `/cost reset` | Close and persist the current period, start a fresh one at zero — every aggregate (cache storage, memory worker, embeddings, compaction, provider context edits) is cleared, so a hard stop armed by the old period never survives |
| `/cost last` | Show the previous persisted snapshot (also accepted: `--last`, `-l`) |
| `/cost sessions` | List recent persisted snapshots, most recent first |
| `/cost export [path]` | Write the current snapshot as JSON (defaults into the cost store); a `.csv` path writes one row per provider/model plus a total |

The summary is rendered as a boxed panel (values below are self-consistent for `claude-sonnet-5`):

```text theme={"system"}
╭─ $ Session cost ─────────────────────────────────────
│  Provider:  CLAUDEAI
│  Model:     claude-sonnet-5
│  Duration:  23m15s
│  Requests:  14
│  Source:    API usage data
│
│  Tokens:
│    Input:    45.2K   ████████████████████
│    Output:   12.8K   █████
│    Total:    58.0K
│    Processed: 98.2K  (input + cache reads + cache writes + output)
│
│  Cache Tokens:
│    Created:  2.1K
│    Read:     38.1K
│  Cache economics:
│    Read discount:  $0.1029   (38.1K reads billed below the input price)
│    Write premium:  -$0.0016  (2.1K writes billed above the input price)
│    Net savings:    $0.1013
│    Without cache:  $0.4482   actual $0.3469 · 23% saved
│
│  Prompt cache: 14 requests · 91% of input from cache · 0 unstable misses · 0 expired · 1 expected rebuilds warm (5m TTL, last activity 40s ago)
│  write/read ratio 0.06 (cache writes over cache reads; a prefix that only grows stays well under 1)
│
│  Cost:
│    claudeai/claude-sonnet-5: (API)
│      Input:   $0.1356
│      Output:  $0.1920
│      Cache:   $0.0193
│
│    Total:   $0.3469
╰──────────────────────────────────────────────────────
```

<Info>
  The **Source** row reports whether the session has real API usage; each model line in the cost breakdown additionally carries its own tag — `(API)` for provider-reported counts, `(estimate)` for character-based estimation — because one session can mix both. Reasoning tokens (o-series / GPT-5 / Gemini thinking) get their own informational line when present; they are already billed inside Output.
</Info>

Models that matched **no pricing table entry** are not silently dropped: `/cost` lists them explicitly (`No pricing table for: provider/model — this spend is NOT in the total.`), so an under-reported total is always visible as such.

***

## Prompt Caching by Provider

ChatCLI arranges every request so the part that does not change between turns comes first (system prompt, attached contexts, tool catalog), and then asks the provider to cache it wherever the API offers a handle for that. The **Prompt cache** line in `/cost` and the `cache NN%` figure in the chat/agent footer are provider-neutral: they are computed from the cache fields every provider reports, on both accounting schemas (Anthropic/Bedrock count cache reads *alongside* input; OpenAI/Gemini/Grok/Kimi count them as a *subset* of the prompt).

* **requests** — how many calls reported cache activity.
* **% of input from cache** — share of the session's input served from cache.
* **unstable misses** — requests that lost the prefix while it should have been warm. A request is judged against what the previous request of the same provider left readable: on Anthropic and Bedrock everything the previous request read plus what it wrote, on OpenAI, Gemini, xAI and Kimi the previous request's whole prompt. Reading back noticeably less than that (below 90%, or 80% on the block-granular subset caches) is the signature of a stable prefix that changed (model or effort switch, MCP tool set churn, refreshed attachment). A large write or a large prompt on its own is never a miss — that is what a big tool result or a pasted file looks like on a perfectly stable prefix. Three unstable misses in a row print a one-shot notice.
* **attribution** — every provider's prompt cache is prefix-keyed over the same regions in the same order: tool definitions, system blocks, messages. Before each request ChatCLI fingerprints those regions and diffs the list against the previous request's, so a lost prefix is booked against the first region that changed: `tool set changed`, `system block N changed`, `message N (tool result / skills block / turn context / user message …) changed`, `history shrank`. A loss on a request that only grew at its end is reported as `no change in the request (server-side or lifetime)` rather than blamed on the prompt. `/cost` prints the table (`lost prefixes by cause`) and the last five events (`#7 expected rebuild · cache lifetime promoted to 1h`). When `auto` promotes the lifetime to the hour after an observed expiry, the health line says when (`promoted to 1h at +13m`), and the one rewrite the new markers force on the next request is declared and attributed to the promotion instead of counted as instability.
* **lanes** — the health line, the miss counters and the write/read ratio describe the **main conversation only**. The memory worker's extraction prompt and a squad worker's own loop share the provider but not the prefix, so their requests are booked in their own lanes: they count toward every token and dollar figure, but never as a lost prefix of your conversation. Their raw cache buckets are still logged per request.
* **expired** — the prefix was lost after a pause the cache did not survive (more than 5 minutes idle, less than an hour). Counted apart because it says nothing about the prefix's stability; it is also the evidence the [hour-long TTL promotion](/context/prefix-curation) acts on.
* **expected rebuilds** — misses ChatCLI itself caused by changing the prefix on purpose: rewriting the conversation (auto-compact, `/compact`, microcompact, skill aging, overflow recovery, provider context edits) or switching what the prefix holds (`/switch` provider or model, `/agent attach|detach`, `/context attach|detach`, `/session load|attach`, MCP server start/stop/restart/reload/login/logout, skill `pin|unpin`). Not a problem.
* **per provider** — "first request" and the hit ratio are tracked per provider (each has its own cache and schema), so switching providers never spends the other's first, and the ratio shown is the last-used provider's own. On the subset-schema providers (OpenAI, Gemini, xAI, Kimi) a prompt at or above the cacheable minimum served with `cached_tokens: 0` after that provider's first request counts as a miss — so the 3-miss alert works there too. Bedrock reports the TTL the marker actually carried (1h on Claude 4.5+ when configured).
* **warm / cold** — whether the last activity is still within the cache TTL.

| Provider | Request-side cache handle | Telemetry |
| - | - | - |
| Anthropic (API key, token, OAuth) | `cache_control` on system blocks, the last tool definition and a **rolling breakpoint on the last user/tool-result message** (the conversation itself is now a cacheable prefix); `CHATCLI_PROMPT_CACHE_TTL=1h` for the extended lifetime | Read/write split incl. the 1-hour write share priced at 2× |
| AWS Bedrock — Claude (InvokeModel, Mantle) | Same markers, one lifetime on the system block and the history breakpoint (Anthropic requires longer-lived breakpoints to precede shorter ones); `cache_control.ttl=1h` on Claude 4.5+ when `CHATCLI_PROMPT_CACHE_TTL` resolves to `1h` (Claude 3.x and 4.0/4.1 stay on the 5-minute wire default) | ✅ incl. the 1-hour write share |
| AWS Bedrock — Converse (Amazon Nova, Claude via Converse) | `cachePoint` after the system blocks and on the last user message, with `ttl: 1h` under the same Claude 4.5+ rule; other Converse vendors reject the block and stay automatic | ✅ (`TokenUsage` cache read/write, `cacheDetails` 1-hour split) |
| OpenAI (Chat Completions, streaming, tools, Responses API) | `prompt_cache_key` derived from the system prompt so every turn of a session lands on the same cache shard; caching itself is automatic (≥ 1,024 tokens); on gpt-5.5 / gpt-5.5-pro `prompt_cache_retention: 24h` when `CHATCLI_PROMPT_CACHE_TTL` resolves to `1h` (gpt-5.6+ moved to `prompt_cache_options`, whose only ttl is 30m — nothing is sent) | ✅ (`cached_tokens`) |
| OpenRouter | `cache_control` on system and last user message for `anthropic/` (with the configured `ttl`) and `google/` models (5 minutes, fixed upstream); `prompt_cache_key` for `openai/` models; at most four `cache_control` blocks per request (the upstream cap); other vendors cache automatically | ✅ |
| Google Gemini | Implicit caching (automatic on 2.5+, favours the stable-first prefix). Opt-in **explicit `cachedContents`** with `CHATCLI_PROMPT_CACHE_EXPLICIT=true`: the system instruction becomes a cache resource referenced by every turn (see below) | ✅ (`cachedContentTokenCount`, cached reads priced at 10% of input) |
| xAI Grok, Z.AI GLM, Moonshot Kimi, MiniMax, GitHub Copilot, DeepSeek (via OpenRouter) | Automatic provider-side caching on repeated prefixes; xAI, Z.AI, Moonshot, MiniMax (OpenAI-compatible path) and Copilot also receive `prompt_cache_key` (ignored where the upstream does not route on it; the Bedrock OpenAI family is a schema-validated upstream and gets none; MiniMax's Anthropic-compatible API documents no `cache_control`); no request-side field in the API | ✅ where the API reports `cached_tokens` |
| Ollama | Local KV-cache reuse on identical prefixes | — |
| StackSpot, OpenAI Assistants, Devin CLI | Server-managed threads / no cache API | Devin reports its cache split; others — |

### Byte-stable prefix and the turn context message

Every provider caches by prefix, so a system message that changes between turns makes every breakpoint after it miss. ChatCLI keeps the system message **byte-stable for the session**: everything that changes per turn — the date (day resolution), proactive memory and session recall, auto-activated and manual skills, MCP channel pushes, the watcher snapshot, retrieved passages — rides as one user-role message flagged `turn_context`, placed right before your turn and persisted with the conversation, so the next request replays identical bytes up to the previous breakpoint. Chat, the RPC surfaces and agent/coder share the mechanism; the injected message is labeled in exports and ignored by memory extraction and recall hints. The OpenAI `prompt_cache_key` hashes only the cache-marked stable parts, and the prefix budget freezes its chars-per-token ratio per session so cached sections do not fold and unfold between turns.

### Explicit cache resources (Gemini)

Implicit caching is free and automatic. Gemini also exposes the cache as a **resource** (`cachedContents`) with a guaranteed discount on every read, billed as **storage per token-hour** (Flash $1.00/M/h, Pro $4.50/M/h, 3.x Flash \$0.50/M/h). Because that charge never appears in a response, it is opt-in: `CHATCLI_PROMPT_CACHE_EXPLICIT=true`. The policy is conservative by design:

* The resource holds the **entire system instruction**; the request references it and omits `system_instruction`.
* Only prompts of roughly **4K+ tokens** qualify, and the same prompt must be seen on **two consecutive requests** before a resource is created — a one-shot never pays for storage.
* Lifetime is `CHATCLI_PROMPT_CACHE_TTL` (`5m`, `1h` or `auto`), extended when less than half remains; the resource is deleted when the prompt changes and at shutdown.
* A resource the API rejects (expired, deleted, floor too low) is dropped and the turn is **retried inline** — a cache can never fail a turn. Refusals back the prompt off for 10 minutes, and a "too small" answer teaches the floor.
* `/cost` prints the storage bought (`Cache storage: N resource(s) · $x`) and counts it in the session total; the amount is persisted with the cost snapshot.

### Embedding cost

Every embedding call of the configured provider (knowledge retrieval and warm-ups, memory vectors, HyDE) is metered: tokens are estimated from characters at the provider's list rate per million tokens, and `/cost` prints `Embeddings: N call(s) · ~T tokens (estimated from characters) · $x`. Ollama meters at \$0. The amount joins the session total and persists with the snapshot.

### Background spend and attribution

Every request is booked to the call that made it. The Level 2 summarizer's usage is read from its own call (an external context engine, an early return or a failure books nothing; a retried summary bills both calls). The memory worker (extraction, rollups, memory compaction) runs on its own client — `CHATCLI_COMPACT_MODEL` when set, otherwise a dedicated instance of the session model — so it never clobbers the interactive turn's usage; it honors `CHATCLI_BUDGET_HARD_STOP` like every other caller and `/cost` prints `Memory worker: N call(s) · $x`. A streamed reply that ends without a usage block is booked as a character estimate (flagged as such) rather than as free, and providers reset their usage before each request so a response without usage can never re-book the previous call. Gemini thinking tokens (`thoughtsTokenCount`) are billed on top of the candidates count; OpenAI reasoning tokens are already inside `completion_tokens` and stay informational.

### Compaction cost

`/cost` also prints `Compactions: N (level 3: M) · summarizer $x` once the history was compacted: how many times, how many landed at Level 3 (emergency truncation) and what the Level 2 summarizer consumed. The summarizer's request is a real request on the session route (or on `CHATCLI_COMPACT_MODEL`), so it joins the session totals like any turn; the counters persist with the session. See [context recovery](/context/context-recovery#what-every-level-guarantees) for the back-off that stops paying the summarizer when it cannot converge.

Moonshot/Kimi context caching is **fully automatic** on the current platform (no cache ids, no management API; the previous request must exceed 256 prompt tokens), so nothing is created there — the `cached_tokens` split is what the tracker prices.

***

## Token counting by provider

The footer `ctx %`, the compaction budget and `/context status` share **one estimate** with four categories — system prompt, history, native tool definitions and the answer reserve (`max_tokens` as sent, capped at a quarter of the window) — the same breakdown Claude Code's `/context` shows. The footer renders the reserve apart — `ctx 2% (+14% reserve)` — so a fresh conversation never reads as a capped window: the first number is what already occupies the window, the second is the space the next request must leave free for the answer (providers require `input + max_tokens ≤ window`; lower `/max-tokens` shrinks it). `/context status` prints the two last lines, and the compactor reserves the prompt and the tool definitions next to the history it measures. The retrieved-passages budget scales with the window (at most 24K chars, 15% of the window in chars, 4K floor). The chars-per-token ratio behind that estimate is learned from every provider's **real usage** (the same 14-provider coverage as the table below), so it is exact after the first turn everywhere. Providers that also expose a **counting API**, and the GPT family through a **local tokenizer**, anchor the ratio before a request goes out — one free count every 8 chat turns and on `/context status`, which then prints the provider-counted size of the live history:

| Provider | Counting mechanism | Used for |
| - | - | - |
| **GPT everywhere** — OpenAI (Chat Completions, Responses), Copilot, OpenRouter `openai/*`, Bedrock OpenAI family | local tokenizer with the model's own BPE encoding (`o200k_base` for 4o/4.1/4.5/5.x and the o-series, `cl100k_base` for GPT-4/3.5); the vocabulary is fetched once from OpenAI's public blob and cached under `~/.chatcli/tokenizers` (no key, nothing embedded in the binary), warmed at startup for GPT sessions; a cold cache loads in the background and never holds a turn | exact calibration, `/context status` |
| Anthropic (API key, token, OAuth) | `POST /v1/messages/count_tokens` — same messages, system blocks and cache markers as the send | exact calibration, `/context status` |
| Google Gemini | `models/{model}:countTokens` with the full `generateContentRequest` | exact calibration, `/context status` |
| AWS Bedrock | `CountTokens` (Converse input for the Converse family, InvokeModel body for Claude; GPT via the local tokenizer) | exact calibration, `/context status` |
| Moonshot Kimi | `POST /v1/tokenizers/estimate-token-count` | exact calibration, `/context status` |
| Z.AI GLM | `POST /api/paas/v4/tokenizer` (coding-plan endpoint honored) | exact calibration, `/context status` |
| xAI, MiniMax, Ollama, StackSpot, Devin, non-GPT OpenRouter models | no counting API published | usage-calibrated ratio (exact per turn from `prompt_tokens`) |

A failed or slow count (10 s bound) never affects a turn: the learned ratio simply stands.

The learned ratios are **persisted** in `~/.chatcli/calibration.json` (atomic write, flushed at exit) and reloaded on the next start, so a new process does not begin at 4 chars per token again. Under the [gateway](/gateway/chat-gateway) the file lives under each tenant's root, so tenants never share ratios. Every estimate in the CLI reads the same learned ratio: the file-processing budget, `/metrics`, `/context list` and attach feedback, the compression-savings footer, the compaction budget and the context manager's chunker, validator and digests.

## Real API Data — Provider Coverage

The tracker prefers real usage from the provider's response and falls back to a character estimate (`chars/4`, marked `IsReal=false`) only when the provider reports nothing:

| Provider | Real usage | Notes |
| :- | :- | :- |
| OpenAI (Chat Completions, streaming, Responses API, Assistants) | ✅ | `stream_options: {include_usage: true}` on streams; `response.completed` on Responses; cached + reasoning details |
| Anthropic (API key and OAuth, buffered, stream and tool paths) | ✅ | Cache creation/read tokens; per-client state, no cross-attribution between parallel clients |
| AWS Bedrock (InvokeModel, Converse, OpenAI-compatible family, Mantle endpoint) | ✅ | Converse reports typed `TokenUsage` including cache read/write |
| Google Gemini | ✅ | Also captures `cachedContentTokenCount` (cache) and `thoughtsTokenCount` (reasoning) |
| xAI, Z.AI, MiniMax, Moonshot, Copilot | ✅ | Buffered and native tool paths |
| OpenRouter | ✅ | Also records `usage.cost` — the **actually billed amount**, authoritative over local tables |
| Ollama, StackSpot | ✅ (tokens) | Real token counts, cost \$0 by design (unmetered backends) |
| Devin CLI | ✅ | Each turn asks the CLI for its ATIF trajectory (`--export`) and reads the real `input_tokens` / `output_tokens` / cache split back — from the Devin-native metrics block or from the ATIF standard block enterprise builds emit, where an Anthropic cache write arrives under `extra.cache_creation_input_tokens` and the prompt count already includes the cached share (the tracker prices that share once, at the cache rate); an older build without `--export` is detected once and falls back to the estimate (`DEVIN_CLI_USAGE_EXPORT=false` opts out) |
| Fallback chain | ✅ | Usage is forwarded from — and attributed to — the entry that actually served the request |

### Modes covered

Every LLM call a session makes is accounted for and attributed to the model that actually served it (skill hints and `@model` route overrides included):

* **Chat** — streaming, buffered, and the `/ask` tool-exception path
* **Agent / Coder** — every ReAct turn, plus **dispatched workers/subagents** (each worker's client records every LLM call live, under the worker's own provider+model — and the budget hard stop gates each call)
* **MoA panels** — native-tools, XML and plain paths, attributed per participant
* **One-shot (`-p`)**, **RPC/Gateway/ACP** turns and **scheduled runs** (the scheduler runs on a dedicated client and receives token/cost figures back — a deliberate self-contained estimate, immune to races with interactive turns)

***

## Pricing Tables

Prices in USD per 1M tokens, matching `cli/cost_tracker.go` (pinned by `cost_tracker_pricing_test.go`). Matching is by model-id substring, most specific first.

### Anthropic (cache write = 1.25× input, cache read = 10% of input — 5% on Opus 5.5 and Sonnet 5.5, 2.5% on Fable 5.1)

| Model | Input | Output |
| :- | :- | :- |
| claude-fable-5-1 (cache read \$0.25) | \$10.00 | \$50.00 |
| claude-fable-5 (cache read \$1.00) | \$10.00 | \$50.00 |
| claude-opus-5-5 (cache read $0.20; fast mode $8/\$40 not modeled) | \$4.00 | \$20.00 |
| claude-opus-5 | \$5.00 | \$25.00 |
| claude-opus-4-5 / 4-6 / 4-7 / 4-8 | \$5.00 | \$25.00 |
| claude-opus (4.1 and older, legacy) | \$15.00 | \$75.00 |
| claude-sonnet-5-5 (cache read \$0.10) | \$2.00 | \$10.00 |
| claude-sonnet-5 (permanent list price) | \$2.00 | \$10.00 |
| claude-sonnet (4.5 / 4.6 and older) | \$3.00 | \$15.00 |
| claude-haiku-5-5 (cache read $0.01; prompt over 100K: every line 5× — $0.50 / $2.50, cache read $0.05) | \$0.10 | \$0.50 |
| claude-haiku-4-5 | \$1.00 | \$5.00 |
| claude-haiku (legacy) | \$0.25 | \$1.25 |

### OpenAI (cache read = 50% of input and no write surcharge unless noted)

| Model | Input | Output |
| :- | :- | :- |
| gpt-6.1-sol (cache write 1.25×, read 5%) | \$2.00 | \$10.00 |
| gpt-6-astra (cache write 1.25×, read 10%; >272K input: 2× input / 1.5× output, modeled per call) | \$10.00 | \$50.00 |
| gpt-6-sol (cache write 1.25×, read 10%) | \$2.00 | \$10.00 |
| gpt-6-luna (cache write 1.25×, read 10%) | \$0.10 | \$0.50 |
| gpt-5.6 (Sol / family alias — since Aug 21, 2026, promo through at least Nov 21, 2026) | \$4.00 | \$20.00 |
| gpt-5.6-terra | \$2.00 | \$12.00 |
| gpt-5.6-luna | \$0.20 | \$1.20 |
| gpt-5.5-pro | \$30.00 | \$180.00 |
| gpt-5.5 (cache read 10%) | \$5.00 | \$30.00 |
| gpt-5.4-pro | \$30.00 | \$180.00 |
| gpt-5.4-mini | \$0.75 | \$4.50 |
| gpt-5.4-nano | \$0.20 | \$1.25 |
| gpt-5.4 (cache read 10%) | \$2.50 | \$15.00 |
| gpt-5.3-codex | \$1.75 | \$14.00 |
| gpt-5.2-pro | \$21.00 | \$168.00 |
| gpt-5.2 | \$1.75 | \$14.00 |
| gpt-5-pro | \$15.00 | \$120.00 |
| gpt-5-mini | \$0.25 | \$2.00 |
| gpt-5-nano | \$0.05 | \$0.40 |
| gpt-5 / gpt-5.1 | \$1.25 | \$10.00 |
| gpt-4.1 | \$2.00 | \$8.00 |
| gpt-4o | \$2.50 | \$10.00 |
| gpt-4o-mini | \$0.15 | \$0.60 |
| gpt-4-turbo | \$10.00 | \$30.00 |
| gpt-4 | \$30.00 | \$60.00 |
| gpt-3.5 | \$0.50 | \$1.50 |
| o3-mini / o4-mini | \$1.10 | \$4.40 |
| o3 | \$10.00 | \$40.00 |
| o1-mini | \$3.00 | \$12.00 |
| o1 | \$15.00 | \$60.00 |

### Google (cache read = 25% of input)

| Model | Input | Output |
| :- | :- | :- |
| gemini-3.8-flash / 3.7-flash / 3.6-flash (intro pricing through Dec 2026) | \$0.75 | \$3.75 |
| gemini-3.5-flash | \$1.50 | \$9.00 |
| gemini-3.5-flash-lite | \$0.30 | \$2.50 |
| gemini-3.1-pro | \$2.00 | \$12.00 |
| gemini-3.1-flash-lite | \$0.25 | \$1.50 |
| gemini-3-flash | \$0.50 | \$3.00 |
| gemini-3 (other 3.x ids, incl. the retired gemini-3-pro) | \$2.00 | \$12.00 |
| gemini-2.5-pro | \$1.25 | \$10.00 |
| gemini-2.5-flash | \$0.30 | \$2.50 |
| gemini-2.5-flash-lite | \$0.10 | \$0.40 |
| gemini-2.0 (shut down Jun 1, 2026 — kept for old session logs) | \$0.075 | \$0.30 |
| gemini-1.5-pro | \$1.25 | \$5.00 |
| gemini-1.5-flash | \$0.075 | \$0.30 |

### xAI (Grok) — cached input = 25% on grok-4.7/4.6, 15% on grok-4.5, 16% elsewhere; no write surcharge

| Model | Input | Output |
| :- | :- | :- |
| grok-4.7 / grok-4.6 / grok-4.5 | \$2.00 | \$6.00 |
| grok-4.3 / grok-4.20 | \$1.25 | \$2.50 |
| grok-build / grok-code-fast | \$1.00 | \$2.00 |
| grok-3 / grok-4-\* (retired May 15, 2026 — xAI redirects them to grok-4.3 and bills grok-4.3 rates) | \$1.25 | \$2.50 |
| grok-2 | \$2.00 | \$10.00 |
| grok (generic) | \$5.00 | \$15.00 |

### Z.AI (GLM)

| Model | Input | Output |
| :- | :- | :- |
| glm-5.3 / glm-5.2 / glm-5.1 (cache read \$0.26) | \$1.40 | \$4.40 |
| glm-5.3-flashx (cache read \$0.075) | \$0.37 | \$1.25 |
| glm-5.3-flash (cache read \$0.03) | \$0.15 | \$0.50 |
| glm-5-turbo / glm-5v-turbo | \$1.20 | \$4.00 |
| glm-5 (cache read \$0.20) | \$1.00 | \$3.20 |
| glm-4.7 / glm-4.6 / glm-4.5 (cache read \$0.11 on glm-4.7 only; 4.6/4.5 bill cached tokens at the input price) | \$0.60 | \$2.20 |
| glm-4.7-flashx | \$0.07 | \$0.40 |
| glm-4.7-flash / glm-4.5-flash / glm-4.6v-flash | free | free |
| glm-4.6v-flashx | \$0.04 | \$0.40 |
| glm-4.6v | \$0.30 | \$0.90 |
| glm-4.5-air | \$0.20 | \$1.10 |
| glm-4.5-airx | \$1.10 | \$4.50 |
| glm-4.5-x | \$2.20 | \$8.90 |
| glm-4.5v (cache read \$0.11) | \$0.60 | \$1.80 |
| other Z.AI ids | \$0.50 | \$0.50 |

### DeepSeek (peak price; cache read = 25% of input)

| Model | Input | Output |
| :- | :- | :- |
| deepseek-v4-pro | \$1.32 | \$3.96 |
| deepseek-v4 | \$0.44 | \$1.32 |
| deepseek-r1 / deepseek-reasoner | \$0.55 | \$2.19 |
| deepseek (generic) | \$0.27 | \$1.10 |

### Moonshot (Kimi) — cache-miss price; cache read = 10% of input on K3, 20% on K2.7 Code, \~17% on K2.6

| Model | Input | Output |
| :- | :- | :- |
| kimi-k3 | \$3.00 | \$15.00 |
| kimi-k2.7-code-highspeed | \$1.90 | \$8.00 |
| kimi / moonshot (generic, incl. k2.7-code, k2.6 — the retired kimi-k2.5 / moonshot-v1-\* keep this tier so old session logs still price) | \$0.95 | \$4.00 |

### Others

| Provider | Pricing |
| :- | :- |
| MiniMax | MiniMax-M3 and MiniMax-M2.7 $0.30 input / $1.20 output, MiniMax-M2.7-highspeed $0.60 / $2.40 (cache read $0.06 on all three; MiniMax-M3 bills 2× on both above 512K input); other models $0.20 / \$1.10 (flat) |
| GitHub Copilot | $2.50 / $10.00 (flat, model-independent) |
| OpenRouter | Re-dispatched to the underlying family (claude/gpt/gemini/deepseek), plus flat rates for llama ($0.20/$0.20), mistral ($0.20/$0.60), qwen ($0.15/$0.15) — and **overridden by the billed `usage.cost`** whenever the response carries it |
| Ollama, StackSpot | \$0 — unmetered from ChatCLI's viewpoint |
| Devin CLI | Three sources, in order: `CHATCLI_MODEL_PRICING` when you pinned a rate; the **per-account rate the CLI lists** (`devin models list --format json`, `cost_summary`) — registered when models are listed and persisted in `~/.chatcli/devin_models.json` so one-shot runs price too; otherwise a **static table of Cognition's rates** mirrored from an account listing (Sep 2026), so an enterprise build whose listing carries no `cost_summary` still estimates instead of reporting $0. A variant or alias without its own rate prices at its family's; `-fast`/`-priority` variants carry their own surcharge. A family the listing never priced (`adaptive`, unlisted newcomers) stays a known zero — and `/cost` says so instead of showing a silent $0 |

<Tip>
  Prices are updated in ChatCLI releases and pinned by tests. An unlisted model is reported as **unpriced** in `/cost` (never silently zero), and OpenRouter's billed cost covers its long tail regardless of the local tables.
</Tip>

***

## Cache Accounting

Cache tokens are priced per family — and, crucially, with the correct **semantics** per provider:

* **Anthropic / Bedrock Anthropic** report cache tokens *alongside* `input_tokens` (additive): write is billed at 1.25× input, read at 10% (5% on Opus 5.5 and Sonnet 5.5, 2.5% on Fable 5.1).
* **OpenAI, xAI, Gemini, DeepSeek, Moonshot** report cache reads as a *subset* of the prompt count: the cached slice is carved out of the input and billed once at the discounted rate — never full price plus discount.
* A cache write on a model with no published write rate is billed at the **input price** (it used to be booked as free).

The `Cache economics` block is the session's cache balance in dollars, computed from each model's actual rates — not a fixed percentage. **Read discount** is what the cache reads saved against the input price. **Write premium** is what the cache writes cost above the input price: 1.25x on Anthropic and Bedrock, 2x with the hour-long TTL, nothing on OpenAI, Gemini, xAI and Kimi (the line is omitted when it is zero). **Net savings** is the difference, and **Without cache** is what the same tokens would have cost with no cache at all, next to the actual total. Reads alone overstate the saving: a session that rewrote its prefix a few times can have given back a third of what the reads saved, and this is where that shows. The `Processed` token row is every token the provider tokenized, cached or not, since `Input` on the additive schemas is only the uncached share.

***

## Session Persistence

Cost data is written through to disk as the session runs (throttled, atomic writes):

```text theme={"system"}
~/.chatcli/costs/<start-timestamp>-<pid>.json
```

**Every surface finalizes.** One-shot (`-p`), the gateway daemon and the MCP/ACP servers persist the cost snapshot and the daily spend and release paid provider caches on exit, exactly like the REPL's cleanup; embedding spend counts toward the daily budget. When the budget hard stop trips inside the agent loop the run is **parked** (resume at the next local midnight, or `/parked resume` once the budget is raised) instead of dying with the work lost.

**Long-context tiers, subscriptions and the daily budget.** A call past the provider's long-context threshold is priced at its tier per call — Claude Sonnet 4/4.5 (1M beta) past 200K (2× input, 1.5× output; Claude 4.6 and later run the full 1M at the standard price, except Haiku 5.5, which bills every line 5× past 100K), Gemini 2.5 Pro and 3.1 Pro past 200K (2× / 1.5×; Flash lines have a single tier), xAI Grok 4.x and grok-build once the prompt reaches 200K (2× / 2×), OpenAI GPT-6 family, GPT-5.6 tiers, gpt-5.5 and gpt-5.4 above 272K input (2× / 1.5×), MiniMax-M3 above 512K (2× / 2×) — and booked as a billed amount so the aggregate never averages it away; the footer estimate agrees. GitHub Copilot is a subscription (premium requests, not tokens) and is booked as known-zero. Every known-zero model that carried tokens is named under the total (`No per-token rate known for: …`) so a \$0 never reads as free when it means unknown, and `CHATCLI_MODEL_PRICING` attributes a rate to it. Amazon Nova Micro/Lite/Pro/Premier carry their on-demand list rows; Nova 2 generations are reported as unpriced until their list price is verified rather than guessed. A daily-only budget (`CHATCLI_DAILY_BUDGET_USD` without a session limit) announces its own warning and exceeded notices, and when both limits are set the daily one speaks when it tripped; `/cost export path.csv` writes one CSV row per provider/model plus a total.

The snapshot (`SessionCostData`) carries per-model usage records, totals, timestamps and the bound session name when one is set. Snapshots are also saved on shutdown and on `/cost reset`, and pruned after **90 days**.

* `/cost last` — the previous session's (or period's) spend
* `/cost sessions` — the 10 most recent snapshots with date, requests, tokens and total, plus each session's cache record: hit share, write/read ratio, unstable misses, expiries, expected rebuilds and the TTL in effect. The prompt-cache telemetry is persisted with the cost snapshot and restored with the session, so prefix stability can be compared across sessions instead of dying with the process; the same two figures are exported to OTLP as `chatcli.cache.hit_pct` and `chatcli.cache.write_read_ratio`.
* `/cost export [path]` — current snapshot as JSON, for pipelines and audits

***

## Session Budget

| Environment Variable | Description | Default |
| :- | :- | :- |
| `CHATCLI_SESSION_BUDGET_USD` | Maximum spending limit per session in USD | 0 (no limit) |
| `CHATCLI_DAILY_BUDGET_USD` | Maximum spend per calendar day across every session under the store directory (per tenant on the gateway); shares the warning share and hard stop below | 0 (no limit) |
| `CHATCLI_MODEL_PRICING` | Pin a per-token rate by hand for any provider: `PROVIDER:model=input/output` in USD per million tokens, entries separated by `;`, model `*` prices the whole provider. Outranks the account listing and the static tables — the knob for a metered wrapper that never tells ChatCLI its price (an enterprise Devin CLI whose listing has no `cost_summary`), a paid Ollama host or a subscription you want attributed. Example: `DEVIN:claude-sonnet-4.6=3/15;DEVIN:*=1/5`. `/config` shows how many entries took effect and which ones did not parse. | — |
| `CHATCLI_BUDGET_WARNING_PCT` | Fraction of the limit at which the warning fires | 0.80 (80%) |
| `CHATCLI_BUDGET_HARD_STOP` | Refuse new LLM turns once the limit is exhausted | false |

All three are re-read on `/reload` — no restart needed to change a budget mid-session.

### Budget Levels

| Level | Condition | Behavior |
| :- | :- | :- |
| `BudgetOK` | Below the warning threshold | Normal |
| `BudgetWarning` | Between warning threshold and the limit | One-shot notice printed the turn the threshold is crossed |
| `BudgetExceeded` | At or above the limit | One-shot notice; with hard stop armed, new turns are refused |

The notice is **proactive**: chat, agent and coder print it on the turn the session crosses a level — you don't have to remember to run `/cost`. Without `CHATCLI_BUDGET_HARD_STOP`, exceeding the budget never interrupts the session; with it, new turns (chat, agent loop, dispatched squad workers mid-wave, MoA participants, scheduled runs, RPC/gateway) are blocked until the limit is raised or `/cost reset` starts a new period.

```bash theme={"system"}
# Limit the session to $5.00, warn at 70%, block once exhausted
export CHATCLI_SESSION_BUDGET_USD=5.00
export CHATCLI_BUDGET_WARNING_PCT=0.70
export CHATCLI_BUDGET_HARD_STOP=true
```

***

## Next Steps

<CardGroup cols={2}>
  <Card title="Conversation Control" icon="clock-rotate-left" href="/usage/conversation-control">
    Use /compact to reduce tokens and costs.
  </Card>

  <Card title="One-Shot Mode" icon="terminal" href="/usage/non-interactive-mode">
    Monitor costs in automated pipelines.
  </Card>

  <Card title="Token Efficiency" icon="gauge-high" href="/context/token-efficiency">
    The optimizations /cost lets you verify.
  </Card>

  <Card title="Model Routing" icon="route" href="/agents/model-routing">
    Per-model attribution of every routed turn.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.