> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Supported AI Models

> Complete reference of AI models natively supported by ChatCLI for each provider, with capabilities, output limits, and API details.

**ChatCLI** supports a wide range of models from major AI providers. Switch models at any time with `/switch --model <name>`.

**Capabilities legend:**

* 👁 **Vision** — accepts images as input
* 🔧 **Tools** — native tool use (function calling)
* 📋 **JSON Mode** — guaranteed structured JSON output
* 💻 **Code Exec** — native code execution on the provider

All providers support **streaming** via SSE (Server-Sent Events). ChatCLI enables streaming automatically.

***

<Tabs>
  <Tab title="OpenAI">
    Models ideal for code generation and complex reasoning. Support both **Chat Completions API** and **Responses API**.

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `gpt-6.1-sol` | `gpt-6-1-sol`, `gpt-6.1` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools (Responses API), 📋 JSON Mode |
    | `gpt-6-astra` | `gpt-6` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-6-sol` | — | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-6-luna` | — | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.6-sol` | `gpt-5.6` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.6-terra` | -- | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.6-luna` | -- | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.5` | -- | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.5-pro` | -- | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.4` | `gpt-5.4-pro` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.4-mini` | `gpt-5.4-nano` | 400K tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.3-codex` | -- | 400K tokens | 128K tokens | 🔧 Tools, 📋 JSON Mode |
    | `gpt-5.2` | `gpt-5.2-pro` | 400K tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-5` | `gpt-5.1`, `gpt-5-mini`, `gpt-5-nano`, `gpt-5-pro` | 400K tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `o3` / `o3-mini` / `o4-mini` | -- | 200K tokens | 100K tokens | 🔧 Tools, 📋 JSON Mode, 🧠 Reasoning |
    | `gpt-4.1` | `gpt-4.1-mini`, `gpt-4.1-nano` | **1.05M tokens** | 32K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
    | `gpt-4o` | `gpt-4o-mini` | 128K tokens | 16K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |

    <Info>
      **GPT-6.1 Sol** (`gpt-6.1-sol`, GA Sep 29, 2026): 1,050,000-token context, 128K output, knowledge cutoff Apr 30, 2026. $2/$10 per MTok, cached input $0.10 (5% of input), cache write $2.50. `reasoning_effort` accepts `low`, `medium` (default), `high`, `xhigh` and `max` — there is no `none`/`minimal`. Tool calling requires the **Responses API** (Chat Completions serves it without tools). There is no 6.1 Astra/Terra/Luna, and the `-pro` variant exists only on OpenRouter (not cataloged). Also on **Copilot** (`gpt-6.1-sol`), **OpenRouter** (`openai/gpt-6.1-sol`), **Bedrock** (`global.openai.gpt-6.1-sol`) and **Devin** (`gpt-6.1-sol`). The bare `gpt-6` alias still resolves to `gpt-6-astra`. Effort hints (`low`/`medium`/`high`) now reach the whole GPT-6 family.
    </Info>

    <Info>
      **GPT-5.6** (GA Jul 9, 2026) ships in three named tiers: **Sol** (flagship), **Terra** (balanced everyday) and **Luna** (fast and affordable) — $4/$20, $2/$12 and $0.20/$1.20 per MTok respectively (Sol dropped from $5/$30 to $4/$20 on Aug 21, 2026 — promotional rate valid at least through Nov 21, 2026). All three work with an **API key** and with **ChatGPT OAuth** (`/auth login openai-codex`); on the Codex backend ChatCLI sends the required client-identification headers automatically (without them the backend returns 404 for Luna).
    </Info>

    <Note>
      **Specs re-verified Sep 2026:** `gpt-5.4` runs the same 1.05M / 128K profile as 5.5/5.6 (its `gpt-5.4-pro` alias is Responses-only, $30/$180); `gpt-5.4-mini`/`-nano` are a separate 400K / 128K entry; `gpt-5.3-codex` (Responses-only) and `gpt-5.2` (+ `gpt-5.2-pro`, Responses-only, $21/$168) are 400K / 128K. The aliases `gpt-5.3`, `gpt-5.3-mini`, `gpt-5.3-nano`, `gpt-5.2-mini` and `gpt-5.2-nano` were removed — no such models exist on the API — and so was `gpt-5.3-codex-spark` (Codex-app-only, never a platform id). **Upcoming OpenAI shutdowns:** `o3-mini`, `o4-mini` and `gpt-4.1-nano` on Oct 23, 2026; the `gpt-5`, `gpt-5-mini`, `gpt-5-nano`, `gpt-5-pro` and `o3` snapshots on Dec 11, 2026; `gpt-5.3-codex`, `gpt-5.4-nano` and `gpt-5.1` on Apr 1, 2027.
    </Note>

    <Info>
      Routing between **Chat Completions** and the **Responses API** is automatic per model via the catalog (`gpt-5.x`, `gpt-4.1` and o-series prefer Responses; `gpt-4o` stays on Chat Completions). Force Responses for every model with `OPENAI_USE_RESPONSES=true`. OAuth sessions always use the Responses API. Streaming is enabled for all models.
    </Info>
  </Tab>

  <Tab title="Anthropic (Claude)">
    Large context windows and excellent ability to follow complex instructions. All models support streaming via SSE.

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `claude-fable-5-1` | `fable-5-1`, `claude-fable-5.1`, `fable-5.1`, `fable` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (always on), ✉️ Mid-conv system |
    | `claude-fable-5` (legacy) | `fable-5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, ✉️ Mid-conv system |
    | `claude-opus-5-5` | `opus-5-5`, `claude-opus-5.5`, `opus-5.5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (always on), ⚡ Fast mode, ✉️ Mid-conv system |
    | `claude-sonnet-5-5` | `sonnet-5-5`, `claude-sonnet-5.5`, `sonnet-5.5`, `claude-5.5-sonnet` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (always on), ✉️ Mid-conv system |
    | `claude-haiku-5-5` | `haiku-5-5`, `claude-haiku-5.5`, `haiku-5.5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, ✉️ Mid-conv system |
    | `claude-opus-5` | `opus-5`, `claude-5-opus` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
    | `claude-sonnet-5` | `sonnet-5`, `claude-5-sonnet` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
    | `claude-opus-4-8` | `opus-4-8` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, ⚡ Fast mode, ✉️ Mid-conv system, 💾 1K-token cache floor |
    | `claude-opus-4-7` | `opus-4-7` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
    | `claude-opus-4-6` | `opus-4-6` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools |
    | `claude-sonnet-4-6` | `sonnet-4-6` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools |
    | `claude-haiku-4-5-20251001` | `claude-haiku-4-5`, `haiku-4-5` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `claude-opus-4-5` | `opus-4-5` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `claude-sonnet-4-5` (deprecated, retires Nov 30, 2026) | `claude-4-5-sonnet`, `sonnet-4-5` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |

    <Info>
      **Claude Fable 5.1** (`claude-fable-5-1`, released Sep 1, 2026) is Anthropic's most capable model — the successor to Fable 5 in the **tier above Opus**: 1M context, 128K output, $10/$50 per MTok, with \*\*cache reads at $0.25/MTok** (2.5% of input instead of the usual 10% — Fable 5 stays at $1). Thinking is **always on** (adaptive): an explicit `thinking:{type:"disabled"}` or `budget_tokens` returns 400 — ChatCLI omits the field unless effort routing fires. Three breaking changes vs Fable 5: a forced `tool_choice` (`any`/`tool`) returns 400 (ChatCLI only ever sends `auto`), thinking blocks are bound to the model that produced them (other models drop them), and editing earlier turns invalidates thinking blocks. No fast mode, no Priority Tier; requires 30-day data retention. Shortcut: `/model fable` — the bare `fable` alias now tracks 5.1, while `fable-5` stays pinned to Fable 5 (legacy, still served). Also on **Bedrock** as `anthropic.claude-fable-5-1` and on **OpenRouter** as `anthropic/claude-fable-5.1`.

      **Claude Fable 5** (`claude-fable-5`) is now legacy but still served. Same API surface as Opus 4.7/4.8 (adaptive thinking only, no `temperature`/`top_p`/`top_k`) with one extra constraint: an explicit `thinking:{type:"disabled"}` returns 400 — the field must be **omitted** to run without thinking (ChatCLI's client already does). Also available on **Bedrock** as `anthropic.claude-fable-5` (dateless ID — the new generation has no ARN-versioned IDs).

      **Claude Sonnet 5** (`claude-sonnet-5`) is the Sonnet-tier successor (Anthropic skipped 4.7/4.8 for Sonnet): 1M context, 128K output, adaptive thinking, **$2/$10 per MTok** — the launch rate became the permanent list price (Anthropic cancelled the Sep 1, 2026 increase to $3/$15). Sonnet 5 does **not** support task budgets — ChatCLI no longer advertises the capability on it. On **Bedrock** it is served exclusively by the Messages endpoint (`anthropic.claude-sonnet-5`) — ChatCLI routes it automatically; see the AWS Bedrock tab.

      **Claude Opus 5** (`claude-opus-5`, Jul 2026) succeeds Opus 4.8 for complex agentic coding and enterprise work: 1M context, 128K output, adaptive thinking (`effort` defaults to `high` server-side), $5/$25 per MTok — the same price as Opus 4.5-4.8. Shortcut: `/model opus-5`. On **Bedrock** it is served through the Messages endpoint (`anthropic.claude-opus-5`), routed automatically like Sonnet 5; on **OpenRouter** the slug is `anthropic/claude-opus-5`.

      **Claude Opus 5.5** (`claude-opus-5-5`, Sep 22 2026) succeeds Opus 5 in the Opus line at a **lower** price: **$4/$20 per MTok**, cache reads $0.20 (5% of input), fast mode $8/\$40. 1M context, 128K output, thinking always on (adaptive — `disabled`/`budget_tokens` return 400), `effort` defaults to `medium` server-side (one level below Opus 5), forced `tool_choice` returns 400. It runs cyber **and** bio classifiers plus `reasoning_extraction`: a refusal arrives as HTTP 200 with `stop_reason: refusal`, and ChatCLI shows the category and the recommended model (see [routing](/agents/model-routing#when-the-provider-refuses-a-turn)). Shortcut: `/model opus-5.5`. Through the Claude Code OAuth surface it requires release **2.1.280+** — ChatCLI presents 2.1.293 and learns from the `claude_code_version_too_old` error if the API asks for newer. On **Bedrock** it is `anthropic.claude-opus-5-5` (Messages endpoint, `global.`/`us.`/`eu.`/`au.` profiles); on **OpenRouter** the slug is `anthropic/claude-opus-5.5`; on **Copilot** `claude-opus-5.5`.

      **Claude Sonnet 5.5** (`claude-sonnet-5-5`, Sep 28 2026) succeeds Sonnet 5 at the same **$2/$10 per MTok**, with cache reads at $0.10 (5% of input, like Opus 5.5) and cache writes at $2.50 (5m) / \$4 (1h). 1M context, 128K output (300K on the Batch API with the `output-300k` beta). Adaptive thinking is on by default and **cannot be disabled** (`thinking` disabled / `budget_tokens` return 400); `effort` ranges `low`..`max`, default `high`. A forced `tool_choice` (`any`/`tool`) returns 400 (ChatCLI only ever sends `auto`), and text the model writes between tool calls comes back as thinking blocks (hidden by default). Supports task budgets and mid-conversation system messages; no fast mode. Shortcut: `/model sonnet-5.5`. Through the Claude Code OAuth surface it requires release **2.1.284+**. On **Bedrock** it is `global.anthropic.claude-sonnet-5-5` (served by `bedrock-runtime`, not Mantle); on **OpenRouter** `anthropic/claude-sonnet-5.5`; on **Copilot** `claude-sonnet-5.5`.

      **Claude Haiku 5.5** (`claude-haiku-5-5`, Oct 7 2026): 1M context, 128K output, priced by **prompt length** — $0.10/$0.50 per MTok (cache read $0.01, 10%) up to 100K prompt tokens; above 100K **every line bills 5×** ($0.50/$2.50, cache read $0.05), and ChatCLI's cost tracker applies that tier per call. Adaptive thinking is on by default and can be disabled only at effort `high` or below; `effort` defaults to `medium`. A forced `tool_choice` is accepted. It uses a newer tokenizer (\~30% more tokens than Haiku 4.5 for the same text). Supports task budgets and mid-conversation system messages. Shortcut: `/model haiku-5.5`. Through the Claude Code OAuth surface it requires release **2.1.293+**. On **Bedrock** it is `global.anthropic.claude-haiku-5-5`; on **OpenRouter** `anthropic/claude-haiku-5.5`; on **Copilot** `claude-haiku-5.5`.

      `claude-opus-4-8` and `claude-opus-4-7` ship with **1M native context** (no extra flag). `claude-opus-4-6` can also use **1M context** by setting `ANTHROPIC_1MTOKENS_SONNET=true`. `claude-sonnet-4-6` now allows **128K output** on the Claude API (was 64K before Mar 2026; Bedrock still caps it at 64K). Different models may use distinct `anthropic-version` headers, managed automatically by the catalog.
    </Info>

    <Note>
      **Retired on the Claude API (removed from the catalog, Sep 2026):** Opus 4.1 (Aug 5, 2026), Opus 4 and Sonnet 4 (Jun 15, 2026) and the whole Claude 3.x line (Sonnet 3.7, Sonnet 3.5, Haiku 3.5, Opus 3, Haiku 3). Requests naming them fail server-side, so their aliases (`opus-4-1`, `opus-4`, `sonnet-4`, `claude-3-7-sonnet`, …) no longer resolve. There was never a Sonnet 4.7 — Anthropic went from Sonnet 4.6 straight to Sonnet 5, and the `sonnet-4-7` placeholder was dropped. Bedrock keeps its own lifecycle (see the AWS Bedrock tab).

      **Deprecated:** Sonnet 4.5 (`claude-sonnet-4-5`) was deprecated on the Claude API on Sep 30, 2026 and retires Nov 30, 2026 — the replacement is `claude-sonnet-5-5`.

      **Catalog order**: entries are declared newest-first in the registry. The alias resolver matches by prefix in registry order, so an older entry whose alias is a prefix of a newer id (`fable-5` ⊂ `fable-5-1`, `opus-4-5` ⊂ `opus-4-5-…`) must come after the newer one — otherwise `fable-5-1` would silently resolve to Fable 5 (wrong cache price, wrong capability flags). If you add a new Claude generation, keep this newest-first order.
    </Note>

    ### Claude Opus 4.8 — what's new

    Released **May 28, 2026**. Same default 1M / 128K profile as Opus 4.7 but with four new launch capabilities the catalog tracks as feature flags:

    | Capability | What it means |
    | :- | :- |
    | `adaptive_thinking` | Only thinking mode accepted by 4.7+. ChatCLI emits `thinking:{type:"adaptive"}` when a skill provides an `effort:` hint — the model decides per turn whether to reason. Sending `budget_tokens` returns HTTP 400. |
    | `fast_mode` | Research-preview faster output (\~2.5× tokens/sec) at premium pricing. Opt in with `ANTHROPIC_SPEED=fast`. |
    | `mid_conversation_system` | Server accepts `role:"system"` after the first user turn, preserving prompt-cache hits across instruction updates. ChatCLI's message builder already passes structured system blocks through unchanged. |
    | `low_cache_minimum` | Minimum cacheable prompt drops from previous models' floor to **1,024 tokens**. Prompts that didn't qualify on 4.7 now create cache entries with no code change. |

    Skill `effort: medium|high|max` continues to work — on Opus 4.7+, Sonnet 5, Sonnet 5.5, Haiku 5.5, Opus 5, Opus 5.5 and Fable 5/5.1 it maps to adaptive thinking automatically; on older 4.x (4.5/4.6) it falls back to budgeted extended thinking (`thinking:{type:"enabled", budget_tokens:N}`).
  </Tab>

  <Tab title="AWS Bedrock (full catalog)">
    Full AWS Bedrock catalog — Anthropic, OpenAI, Llama, Nova, Mistral, Cohere, AI21, DeepSeek, Moonshot Kimi, MiniMax, Qwen, Z.AI/GLM, Gemma, Nemotron, TwelveLabs, and any provider AWS adds. Auth uses the AWS SDK's default credentials chain (IAM role, `~/.aws/credentials`, env vars) — no API key from the original providers is needed.

    <Warning>
      Modern models (Claude 4.x/4.5/4.6 and equivalents from other providers) **do not accept direct on-demand invocation** by base ID — they require an **inference profile ID** (prefixes `global.`, `us.`, `eu.`, `apac.`). ChatCLI **automatically filters** non-invokable base IDs from `/switch --model`, so only what works appears. See [AWS Bedrock](/providers/bedrock) for details.
    </Warning>

    <Note>
      **New generation = dateless IDs.** Fable 5.1, Fable 5, Opus 5.5, Sonnet 5.5, Haiku 5.5, Opus 5, Sonnet 5, Opus 4.8 and Opus 4.7 have **no ARN-versioned IDs** on Bedrock — the IDs are `anthropic.claude-fable-5-1`, `anthropic.claude-fable-5`, `anthropic.claude-opus-5-5`, `global.anthropic.claude-sonnet-5-5`, `global.anthropic.claude-haiku-5-5`, `anthropic.claude-opus-5`, `anthropic.claude-sonnet-5`, `anthropic.claude-opus-4-8` and `anthropic.claude-opus-4-7`. Opus 4.8/4.7 are invoked through the `global.` inference profile (the bare ID is not on-demand invokable on `InvokeModel`). The old dated IDs (e.g. `global.anthropic.claude-opus-4-8-20260528-v1:0`) keep resolving as aliases. **Opus 5.5, Opus 5, Sonnet 5, Fable 5 and Fable 5.1 are served by the Messages endpoint** (`bedrock-mantle.{region}.api.aws/anthropic/v1/messages`) — ChatCLI detects and routes it automatically (SigV4 `bedrock-mantle` or `AWS_BEARER_TOKEN_BEDROCK`), and if the Mantle call fails it retries the same request through the legacy InvokeModel runtime under the `global.` inference profile; see [AWS Bedrock](/providers/bedrock).

      **Sonnet 5.5 and Haiku 5.5 are the exception:** they are **not** Mantle-only — `bedrock-runtime` serves them (Messages, Converse and InvokeModel). There is no in-region option, so the canonical ID is the `global.` inference profile (`global.anthropic.claude-sonnet-5-5`, `global.anthropic.claude-haiku-5-5`); bare first-party IDs are upgraded automatically.
    </Note>

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `anthropic.claude-fable-5-1` | `bedrock-fable-5-1`, `global.anthropic.claude-fable-5-1`, `us.anthropic.claude-fable-5-1`, `claude-fable-5-1`, `fable-5-1` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint, 30-day retention) |
    | `anthropic.claude-fable-5` | `bedrock-fable-5`, `claude-fable-5`, `fable-5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint, 30-day retention) |
    | `anthropic.claude-opus-5-5` | `bedrock-opus-5-5`, `global./us./eu./au.anthropic.claude-opus-5-5`, `claude-opus-5-5`, `opus-5.5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint) |
    | `anthropic.claude-opus-5` | `bedrock-opus-5`, `claude-opus-5`, `opus-5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint) |
    | `global.anthropic.claude-sonnet-5-5` | `anthropic.claude-sonnet-5-5`, `us./eu.anthropic.claude-sonnet-5-5`, `claude-sonnet-5-5`, `sonnet-5.5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (bedrock-runtime) |
    | `global.anthropic.claude-haiku-5-5` | `us./eu./au./jp.anthropic.claude-haiku-5-5`, `claude-haiku-5-5`, `haiku-5.5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (bedrock-runtime) |
    | `anthropic.claude-sonnet-5` | `bedrock-sonnet-5`, `claude-sonnet-5`, `sonnet-5` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint) |
    | `global.anthropic.claude-opus-4-8` | `bedrock-opus-4-8`, `claude-opus-4-8` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, 💾 1K-token cache floor |
    | `global.anthropic.claude-opus-4-7` | `bedrock-opus-4-7`, `claude-opus-4-7` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
    | `global.anthropic.claude-sonnet-4-6` | `bedrock-sonnet-4-6`, `claude-sonnet-4-6` | **1M tokens** | 64K tokens (AWS card) | 👁 Vision, 🔧 Tools |
    | `global.anthropic.claude-opus-4-6-v1` | `bedrock-opus-4-6`, `claude-opus-4-6` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools |
    | `global.anthropic.claude-haiku-4-5-20251001-v1:0` | `bedrock-haiku-4-5`, `claude-haiku-4-5` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `global.anthropic.claude-sonnet-4-5-20250929-v1:0` | `bedrock-sonnet-4-5`, `claude-sonnet-4-5` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `us.anthropic.claude-sonnet-4-5-20250929-v1:0` | `bedrock-sonnet-4-5-us` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `global.anthropic.claude-opus-4-5-20251101-v1:0` | `bedrock-opus-4-5`, `claude-opus-4-5`, `global.anthropic.claude-opus-4-5-20251001-v1:0` (old misspelling, alias only) | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `global.anthropic.claude-sonnet-4-20250514-v1:0` (Legacy, EOL Oct 14, 2026) | `bedrock-sonnet-4`, `claude-sonnet-4` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `us.anthropic.claude-sonnet-4-20250514-v1:0` (Legacy, EOL Oct 14, 2026) | `bedrock-sonnet-4-us` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `eu.anthropic.claude-sonnet-4-20250514-v1:0` (Legacy, EOL Oct 14, 2026) | `bedrock-sonnet-4-eu` | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
    | `us.anthropic.claude-opus-4-1-20250805-v1:0` (Legacy, extended access since Oct 8, 2026, EOL Jan 8, 2027) | `bedrock-opus-4-1`, `claude-opus-4-1` | 200K tokens | 32K tokens | 👁 Vision, 🔧 Tools |

    <Note>
      **Retired on Bedrock and removed from the catalog (Sep 2026):** Opus 4 (`us.anthropic.claude-opus-4-20250514-v1:0`), Sonnet 3.7 (`us./eu.anthropic.claude-3-7-sonnet-20250219-v1:0`), Sonnet 3.5 v1/v2, Haiku 3.5, Opus 3 and Haiku 3 (`anthropic.claude-3-haiku-20240307-v1:0`, EOL Sep 10, 2026). The placeholder `global.anthropic.claude-sonnet-4-7-20260401-v1:0` never existed on AWS and was dropped. Sonnet 4 and Opus 4.1 are **Legacy** but still invokable until their EOL dates above; Opus 4.1 entered the higher-priced extended-access period on Oct 8, 2026. Opus 4.5's real snapshot date is `20251101` — the previous `20251001` spelling never existed on AWS and is kept only as an alias.
    </Note>

    **OpenAI GPT-5.6 (frontier, Jul 2026)** — Sol, Terra and Luna on Bedrock. They speak Converse, Responses and Chat Completions but **not `InvokeModel`**, so the catalog flags them `bedrock_converse_only` and ChatCLI routes them through the **Converse API** automatically (no env var needed — the `openai.` prefix alone would have sent them down the GPT-OSS `InvokeModel` path). Served only through inference profiles: Sol via `us.`/`global.`, Terra/Luna via `us.`/`in.`/`global.`. Global-endpoint prices: Sol $4/$20, Terra $2/$12, Luna $0.20/$1.20 per MTok. **GPT-6.1 Sol** (`global.openai.gpt-6.1-sol`, `us.` profile) takes the same Converse path: 1M context, 131,072 max output, $2/$10 per MTok (cache write $2.50, cache read $0.10).

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `global.openai.gpt-6.1-sol` | `us.openai.gpt-6.1-sol` | **1M tokens** | 131,072 tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
    | `global.openai.gpt-6-astra` | `bedrock-gpt-6-astra`, `openai.gpt-6-astra`, `us.openai.gpt-6-astra` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
    | `global.openai.gpt-6-sol` | `bedrock-gpt-6-sol`, `openai.gpt-6-sol`, `us.openai.gpt-6-sol` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
    | `global.openai.gpt-6-luna` | `bedrock-gpt-6-luna`, `openai.gpt-6-luna`, `us.openai.gpt-6-luna` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
    | `global.openai.gpt-5.6-sol` | `bedrock-gpt-5.6-sol`, `openai.gpt-5.6-sol`, `us.openai.gpt-5.6-sol` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
    | `global.openai.gpt-5.6-terra` | `bedrock-gpt-5.6-terra`, `openai.gpt-5.6-terra`, `us.openai.gpt-5.6-terra`, `in.openai.gpt-5.6-terra` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
    | `global.openai.gpt-5.6-luna` | `bedrock-gpt-5.6-luna`, `openai.gpt-5.6-luna`, `us.openai.gpt-5.6-luna`, `in.openai.gpt-5.6-luna` | **1.05M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |

    **OpenAI GPT-OSS (open-weights)** — OpenAI models hosted on Bedrock. Use the OpenAI Chat Completions schema (auto-detected by `openai.*` prefix, or forced via `BEDROCK_PROVIDER=openai`).

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `openai.gpt-oss-120b-1:0` | `bedrock-gpt-oss-120b`, `gpt-oss-120b` | 128K tokens | 16K tokens | 🔧 Tools, 📋 JSON |
    | `openai.gpt-oss-20b-1:0` | `bedrock-gpt-oss-20b`, `gpt-oss-20b` | 128K tokens | 16K tokens | 🔧 Tools, 📋 JSON |

    **xAI Grok and Amazon Nova 2 (static entries)** — Grok 4.6 landed on Bedrock on Aug 18, 2026 (Converse, `us.`/`global.` profiles only, $2/$6 per MTok on the global endpoint), followed by **Grok 4.7** (`global.xai.grok-4.7`, `us.` profile, 500K context, $2/$6 per MTok, cached input \$0.50). Nova 2 Lite succeeds Nova Premier (Legacy since Mar 13, 2026, EOL Sep 14, 2026 — removed from the catalog): 1M context, 64K output, served **only** through inference profiles (`us.`/`eu.`/`jp.`/`global.` — there is no bare in-region id). Nova 2 Pro/Omni are not on Bedrock.

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `global.xai.grok-4.7` | `us.xai.grok-4.7` | 500K tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `global.xai.grok-4.6` | `bedrock-grok-4.6`, `xai.grok-4.6`, `us.xai.grok-4.6` | 500K tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `global.amazon.nova-2-lite-v1:0` | `nova-2-lite`, `amazon.nova-2-lite-v1:0`, `us./eu./jp.amazon.nova-2-lite-v1:0` | **1M tokens** | 64K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |

    **Moonshot Kimi K3 and Z.AI GLM-5.3 (static entries)** — Kimi K3 is served through `us.`/`in.` profiles (1M context, image input, $3/$15 per MTok, cache read \$0.30); GLM-5.3 through the `us.` profile (1M context, 128K output, text only — the AWS card publishes no price, so ChatCLI's cost tracker uses the Z.AI list price).

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `global.moonshotai.kimi-k3` | `us./in.moonshotai.kimi-k3` | **1M tokens** | 131K tokens | 👁 Vision, 🔧 Tools |
    | `global.zai.glm-5.3` | `us.zai.glm-5.3` | **1M tokens** | 128K tokens | 🔧 Tools |

    **Other providers via Converse API** — beyond the static entries above, Llama, Nova, Mistral, Cohere, AI21, DeepSeek, Moonshot Kimi, MiniMax, Qwen, Z.AI/GLM, Gemma, Nemotron, TwelveLabs, etc. are **not hardcoded** in the catalog — they appear dynamically in `/switch --model` based on what your AWS account has access to. ChatCLI routes these models through the **AWS Converse API** (unified schema), so adding a new provider doesn't require a release.

    Examples of IDs seen via ListFoundationModels (your actual list depends on account + region):

    | Provider | Example Model ID |
    | :- | :- |
    | Moonshot AI | `moonshotai.kimi-k2.6`, `moonshotai.kimi-k2-thinking` |
    | MiniMax | `minimax.m-2-5`, `minimax.m-2` |
    | Z.AI | `zai.glm-4-7`, `zai.glm-4-7-flash` |
    | Qwen | `qwen.qwen3-32b`, `qwen.qwen3-coder-480b` |
    | Meta Llama | `meta.llama3-70b-instruct-v1:0`, `us.meta.llama3-1-70b-...` |
    | Amazon Nova | `amazon.nova-pro-v1:0`, `amazon.nova-lite-v1:0`, `global.amazon.nova-2-lite-v1:0` |
    | Mistral | `mistral.mistral-large-2407-v1:0` |
    | DeepSeek | `us.deepseek.r1-v1:0` |
    | Google | `google.gemma-3-27b-pt` |
    | NVIDIA | `nvidia.nemotron-nano-9b-v2` |
    | TwelveLabs | `twelvelabs.pegasus-v1.2` |

    <Info>
      The dynamic listing (`/switch --model`) merges `bedrock:ListFoundationModels` (filtered by `ByOutputModality: TEXT` + `InferenceTypesSupported: ON_DEMAND`) and `bedrock:ListInferenceProfiles` with the static catalog above. **No allowlist** — any Bedrock provider your account can access shows up automatically. Use the command to see what your AWS account can actually invoke in the configured region.
    </Info>

    **Embeddings via Bedrock** — `amazon.titan-embed-text-v2:0` (default, 1024-dim, configurable 256/512/1024), `amazon.titan-embed-text-v1` (1536-dim), Cohere `cohere.embed-english-v3` / `cohere.embed-multilingual-v3` (1024-dim), `cohere.embed-v4:0` (1536-dim default, 128K context, also via `us.`/`eu.`/`global.` profiles), and `amazon.nova-2-multimodal-embeddings-v1:0` (Nova MME, 3072-dim default, configurable 256/384/1024/3072). Enable with `CHATCLI_EMBED_PROVIDER=bedrock`. See [RAG + HyDE](/agents/harness/rag-hyde).
  </Tab>

  <Tab title="Google (Gemini)">
    Advanced multimodal capabilities and massive context windows. Support streaming via SSE.

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `gemini-3.8-flash` | `gemini-3.8-flash-latest` | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
    | `gemini-3.7-flash` | `gemini-3.7-flash-latest` | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
    | `gemini-3.6-flash` | `gemini-3.6-flash-latest` | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
    | `gemini-3.5-flash` | `gemini-3.5-flash-latest` | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
    | `gemini-3.5-flash-lite` | -- | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `gemini-3.1-pro-preview` | `gemini-3.1-pro` | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
    | `gemini-3.1-flash-lite` | -- | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `gemini-3-flash-preview` | -- | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `gemini-2.5-pro` | `gemini-2.5-pro-latest` | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
    | `gemini-2.5-flash` | -- | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `gemini-2.5-flash-lite` | -- | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |

    <Info>
      The whole Gemini 3.x generation shares the 1,048,576-token input window with a 65,536-token output cap (`gemini-3.8-flash`, GA Sep 2 2026, is the current workhorse; `gemini-3.7-flash` stays served). `gemini-3.1-pro` without `-preview` is **only an alias**, not a real model code. Gemini 2.5 Flash Lite also supports **Multimodal Live** for real-time interactions; models with JSON Mode can return structured output via `response_mime_type`.
    </Info>

    <Note>
      **Shut down by Google and removed from the catalog:** `gemini-2.0-flash` and `gemini-2.0-flash-lite` (Jun 1, 2026), and `gemini-3` / `gemini-3-pro` / `gemini-3-pro-preview` (Mar 9, 2026 — use `gemini-3.1-pro-preview`). Next scheduled: `gemini-3.1-flash-lite` retires May 7, 2027 — migrate to `gemini-3.5-flash-lite`. The introductory price of 3.7/3.6 Flash ($0.75/$3.75) doubles on Jan 1, 2027.
    </Note>
  </Tab>

  <Tab title="xAI (Grok)">
    Real-time information integration and large context windows. Support streaming.

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `grok-4.7` | `grok-4.7-latest` | 500K tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `grok-4.6` | `grok-4.6-latest` | 500K tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `grok-4.5` | `grok-4.5-latest`, `grok-build-latest` | 500K tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `grok-4.3` | `grok-4.3-latest` | 1M tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `grok-4.20-0309-reasoning` | `grok-4.20-reasoning`, `grok-4.20` | 1M tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `grok-4.20-0309-non-reasoning` | `grok-4.20-non-reasoning` | 1M tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `grok-4.20-multi-agent-0309` | `grok-4.20-multi-agent` | 1M tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `grok-build-0.1` | `grok-build`, `grok-code-fast-1` | 256K tokens | -- | 👁 Vision, 🔧 Tools, 📋 JSON |

    <Note>
      Grok models use the OpenAI-compatible API. xAI publishes no per-model output cap, so limits are managed by the provider. `grok-4.7` (Sep 21 2026) is the current flagship — 500K context, `low|medium|high|xhigh` effort, the same $2/$6 tier as 4.6 (cached input \$0.50; once the prompt reaches 200K xAI bills 2× on input and output, and ChatCLI's cost tracker applies it per call). `grok-4.6` stays served. The `grok-4.7-fast` variant is Cursor/Grok Build only, not on the public API, and therefore not in the catalog. The provider default (`XAI_MODEL`) is now `grok-4.3`.

      **Retired by xAI on May 15, 2026 and removed from the catalog:** `grok-4-fast` (+ `grok-4-fast-reasoning*`, `grok-4-0709`), `grok-4-1` (+ `grok-4-1-fast*`), `grok-3` and `grok-3-mini`. The API keeps those slugs alive but **redirects them to `grok-4.3`** and bills grok-4.3 rates — pinned configs keep working, and ChatCLI's cost tracker prices them accordingly. `grok-code-fast-1` became an alias of `grok-build-0.1`. Cached-input pricing is tracked too: $0.50/MTok on grok-4.7/4.6, $0.30 on grok-4.5, \$0.20 on the 4.3/4.20 tier.
    </Note>
  </Tab>

  <Tab title="GitHub Copilot">
    Use models from the Copilot platform with your subscription (Individual, Business, Enterprise). Authenticate via `/auth login github-copilot`.

    The table below shows models registered in the static catalog. With **dynamic listing**, ChatCLI queries the Copilot API and automatically discovers all models available for your account.

    | Model (ID) | Context |
    | :- | :- |
    | `gpt-6.1-sol` | 1M tokens |
    | `gpt-6-astra` | 1M tokens |
    | `gpt-6-sol` | 1M tokens |
    | `gpt-6-luna` | 1M tokens |
    | `claude-sonnet-5.5` | 1M tokens |
    | `claude-haiku-5.5` | 1M tokens |
    | `claude-opus-5.5` | 1M tokens |
    | `gpt-4o` | 128K tokens |
    | `gpt-4o-mini` | 128K tokens |
    | *+ dynamic models* | *via API* |

    <Info>
      Available models vary depending on your plan and region. Use `/switch --model` to see the full list fetched directly from the Copilot API.
    </Info>

    <Note>
      **Retired by GitHub and removed from the static catalog:** `claude-sonnet-4` (May 1, 2026) and `gemini-2.0-flash` (Oct 23, 2025). GitHub has also scheduled GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, Gemini 3.7 Flash and Grok 4.5 for retirement on Oct 19, 2026.
    </Note>
  </Tab>

  <Tab title="ZAI (Zhipu AI)">
    Chinese AI models from Zhipu AI (z.ai) with strong multilingual and coding capabilities. OpenAI-compatible API with native tool calling support.

    | Model (ID) | Context | Max Output | Capabilities |
    | :- | :- | :- | :- |
    | `glm-5.3` | **1M tokens** | 128K tokens | 🔧 Tools, 📋 JSON |
    | `glm-5.3-flashx` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `glm-5.3-flash` | **1M tokens** | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `glm-5.2` | **1M tokens** | 128K tokens | 🔧 Tools, 📋 JSON |
    | `glm-5.1` | 200K tokens | 128K tokens | 👁 Vision, 🔧 Tools |
    | `glm-5-turbo` | 200K tokens | 128K tokens | 👁 Vision, 🔧 Tools |
    | `glm-5` | 200K tokens | 128K tokens | 👁 Vision, 🔧 Tools |
    | `glm-4.7` | 200K tokens | 128K tokens | 🔧 Tools |
    | `glm-4.6` | 200K tokens | 128K tokens | 🔧 Tools |
    | `glm-4.5` | 128K tokens | 96K tokens | 🔧 Tools |
    | `glm-4.5-flash` | 128K tokens | 16K tokens | 🔧 Tools |
    | `glm-5v-turbo` | 128K tokens | 16K tokens | 👁 Vision, 🔧 Tools |
    | `glm-4.5v` | 128K tokens | 16K tokens | 👁 Vision |

    <Note>
      **GLM-5.3** (released Aug 18, 2026) is Zhipu's open-weight flagship: same base model as GLM-5.2 with scaled-up post-training focused on coding and long-horizon agentic work — **1M-token context**, 128K output, reasoning always on, function calling and structured output (text-only input). **GLM-5.3-Flash** (Aug 26, 2026) is the natively multimodal MIT-licensed sibling: 1M context, 128K output, image/video/file input. List prices: GLM-5.3 $1.40/$4.40, GLM-5.3-Flash $0.15/$0.50, GLM-5.3-FlashX (Sep 18 2026, the high-speed tier of Flash) $0.37/$1.25 per MTok (GLM-5: $1.00/$3.20) — ChatCLI's cost tracker uses these rates, and it now prices the GLM-4.x line per tier as well (4.7 $0.60/$2.20, 4.7-flashx $0.07/$0.40, 4.7-flash free, 4.5-air $0.20/$1.10, 4.5-airx $1.10/$4.50, 4.5-x $2.20/$8.90, 4.5v $0.60/$1.80, 4.5-flash free) instead of the old flat \$0.50 fallback. The provider default model remains `glm-5`; use `/switch --model glm-5.3` to switch. `codegeex-4` was removed — it is no longer served by the Z.AI international API.
    </Note>

    <Info>
      ZAI uses an OpenAI-compatible API at `https://api.z.ai/api/paas/v4/chat/completions`. Authentication is via `ZAI_API_KEY` Bearer token. Model IDs are case-sensitive. Subscribers of the **GLM Coding Plan** can set `ZAI_USE_CODING_PLAN=true` to use the subscription endpoint (`/api/coding/paas/v4`) with the same key — usage draws from the plan instead of pay-as-you-go credits and `/cost` reports it at \$0.
    </Info>

    <Note>
      **Automatic JWT authentication:** Keys in `id.secret` format automatically enable JWT token rotation (HMAC-SHA256), cached for 30 minutes. Keys without "." work as traditional Bearer tokens. No additional configuration needed.
    </Note>
  </Tab>

  <Tab title="MiniMax">
    High-performance models from MiniMax with large context windows and native tool calling. OpenAI-compatible API.

    | Model (ID) | Context | Max Output | Capabilities |
    | :- | :- | :- | :- |
    | `MiniMax-M3` | **1M tokens** | 131K tokens | 👁 Vision, 🔧 Tools |
    | `MiniMax-M2.7` | 204K tokens | 131K tokens | 👁 Vision, 🔧 Tools |
    | `MiniMax-M2.7-highspeed` | 204K tokens | -- | 👁 Vision, 🔧 Tools |
    | `MiniMax-M2.5` | 196K tokens | 65K tokens | 👁 Vision, 🔧 Tools |
    | `MiniMax-M2.5-highspeed` | 196K tokens | -- | 👁 Vision, 🔧 Tools |
    | `MiniMax-Text-01` | 128K tokens | 2K tokens | 📋 JSON Mode |

    <Warning>
      MiniMax model IDs are **case-sensitive** (e.g., `MiniMax-M2.7`, not `minimax-m2.7`). The API uses a `base_resp` field for error handling with `status_code` and `status_msg`.
    </Warning>

    <Note>
      **Anthropic-compatible endpoint:** Set `MINIMAX_API_COMPAT=anthropic` to use `https://api.minimax.io/anthropic/v1/messages` with Anthropic Messages format. Native tool calling is disabled in this mode (falls back to XML). Model listing always uses the native endpoint.
    </Note>
  </Tab>

  <Tab title="Moonshot (Kimi)">
    Moonshot AI's Kimi family — the K3 flagship (Jul 2026) is a 2.8T-parameter MoE with 104B activated, 1M-token context via Kimi Delta Attention and multimodal input; K2.6 (1T/32B, 256K) remains fully supported, with the native MoonViT vision encoder and explicit "thinking" mode across the line. OpenAI-compatible API at `https://api.moonshot.ai/v1/chat/completions`.

    | Model (ID) | Aliases | Context | Max Output | Capabilities |
    | :- | :- | :- | :- | :- |
    | `kimi-k3` | `kimi-k-3`, `k3`, `k-3` | 1M tokens | 131K tokens | 🔧 Tools, 👁 Vision, 🧠 Thinking, 📋 JSON Mode |
    | `kimi-k2.7-code` | `kimi-k2.7`, `k2.7` | 256K tokens | 32K tokens | 🔧 Tools, 🧠 Thinking, 📋 JSON Mode |
    | `kimi-k2.7-code-highspeed` | `kimi-k2-7-code-highspeed` | 256K tokens | 32K tokens | 🔧 Tools, 🧠 Thinking, 📋 JSON Mode |
    | `kimi-k2.6` | `kimi-k2-6`, `k2.6`, `k2-6` | 256K tokens | 131K tokens | 🔧 Tools, 👁 Vision, 🧠 Thinking, 📋 JSON Mode |

    <Note>
      **Thinking mode:** Set `MOONSHOT_THINKING=enabled|disabled|auto` to toggle between Thinking (explicit reasoning, default for K3/K2.6) and Instant (direct response, cheaper). The default `auto` lets the model choose; models without the `thinking` capability ignore the flag.
    </Note>

    <Tip>
      **Public pricing (Aug 2026):** kimi-k3 is $3.00/M input tokens (cache miss; cache hit $0.30/M) and $15.00/M output. kimi-k2.7-code and kimi-k2.6 are $0.95/M input (cache miss) and $4.00/M output; kimi-k2.7-code-highspeed is 2× that ($1.90/\$8.00). Cache-hit input is cheaper on all of them, and the ChatCLI cost tracker prices the cached slice at each model's rate (10% of input on K3, 20% on K2.7 Code, \~17% on K2.6). **Retired by Moonshot and removed from the catalog** (the API now returns 404 for them): `kimi-k2.5` and the `moonshot-v1-128k/32k/8k` series (Aug 31, 2026), `kimi-k2-turbo-preview` (May 25, 2026), `kimi-latest` (Jan 28, 2026) and `kimi-thinking-preview` (Nov 11, 2025). Migration target for all of them is `kimi-k3`; the provider default stays `kimi-k2.6`.
    </Tip>
  </Tab>

  <Tab title="StackSpot">
    Accepts all compatible models on the StackSpotAI platform, selected during Agent creation.
  </Tab>

  <Tab title="OpenRouter">
    **Multi-provider API gateway** — access 200+ models from OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, and more through a single API key. Uses an OpenAI-compatible API at `https://openrouter.ai/api/v1/chat/completions`.

    Models use the `provider/model-name` format:

    | Model (ID) | Provider | Capabilities |
    | :- | :- | :- |
    | `openai/gpt-6.1-sol` | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `openai/gpt-6-astra` | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `openai/gpt-6-sol` | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `openai/gpt-6-luna` | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `openai/gpt-4o` | OpenAI | 👁 Vision, 🔧 Tools |
    | `openai/gpt-4o-mini` | OpenAI | 👁 Vision, 🔧 Tools |
    | `anthropic/claude-opus-5.5` | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `anthropic/claude-opus-5` | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `anthropic/claude-sonnet-5.5` | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `anthropic/claude-haiku-5.5` | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `anthropic/claude-sonnet-5` | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `anthropic/claude-fable-5.1` | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `anthropic/claude-fable-5` | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `anthropic/claude-sonnet-4` | Anthropic | 👁 Vision, 🔧 Tools |
    | `anthropic/claude-opus-4` | Anthropic | 👁 Vision, 🔧 Tools |
    | `google/gemini-2.5-pro` | Google | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `google/gemini-2.5-flash` | Google | 👁 Vision, 🔧 Tools, 📋 JSON |
    | `meta-llama/llama-4-maverick` | Meta | 🔧 Tools |
    | `deepseek/deepseek-r1` | DeepSeek | 🔧 Tools |
    | `mistralai/mistral-large` | Mistral | 🔧 Tools |

    <Info>
      The table above shows popular defaults. OpenRouter provides **200+ models** — ChatCLI discovers them dynamically via the `/api/v1/models` endpoint. Use `/switch --model` to browse the full list.
    </Info>

    <Tip>
      OpenRouter supports **native fallback routing** via `OPENROUTER_FALLBACK_MODELS`. If your primary model is unavailable, OpenRouter automatically routes to the next model in the list — handled server-side before ChatCLI's own fallback chain kicks in.
    </Tip>
  </Tab>

  <Tab title="Devin CLI (Cognition)">
    Served through the local Devin CLI wrapper — ChatCLI keeps the whole conversation and harness; Devin is only the transport. See [Devin Provider](/providers/devin-provider). Family slugs use **dots** (`claude-sonnet-4.6`, not `4-6`); variant ids keep the CLI spelling (`claude-opus-5-high`, `swe-1-6-fast`). `/switch --model` lists what **your account** can invoke via `devin models list --format json` (tagged `[api]`); the table below is the static fallback.

    | Family | Models |
    | :- | :- |
    | Anthropic | `claude-fable-5.1` · `claude-fable-5` · `claude-opus-5.5` · `claude-opus-5` · `claude-sonnet-5` · `claude-opus-4.8` / `4.7` / `4.6` / `4.5` · `claude-sonnet-4.6` / `4.5` / `4` · `claude-haiku-4.5` |
    | OpenAI | `gpt-6.1-sol` · `gpt-6-astra` / `-sol` / `-luna` · `gpt-5.6-sol` / `-terra` / `-luna` · `gpt-5.5` · `gpt-5.4` / `-mini` · `gpt-5.3-codex` · `gpt-5.2` · `gpt-5.1` · `gpt-4.1` |
    | Google | `gemini-3.8-flash` · `gemini-3.7-flash` · `gemini-3.6-flash` · `gemini-3.5-flash` · `gemini-3.1-pro` · `gemini-3-flash` |
    | xAI | `grok-4.6` · `grok-4.5` |
    | Others | `glm-5.3` / `5.2` · `kimi-k3` / `k2.7` / `k2.6` · `deepseek-v4-pro` / `v4-flash` |
    | Cognition (SWE) | `swe-1.7-lightning` · `swe-1.7` · `swe-1.6-fast` · `swe-1.6` |

    <Info>
      Auth belongs to the binary (`devin auth login`, corporate SSO) — no key in ChatCLI. Default model: `claude-sonnet-4.6` (`DEVIN_MODEL`). The CLI reports no token usage, so cost tracking shows zero — cost lives in the Cognition subscription.
    </Info>
  </Tab>

  <Tab title="Ollama (Local)">
    Supports any local model via Ollama. Configure in `.env`:

    ```env theme={"system"}
    OLLAMA_ENABLED=true
    OLLAMA_MODEL="llama3"
    ```

    Or switch interactively: `/switch --model llama3`

    Use `ollama pull <model>` to download new models.
  </Tab>
</Tabs>

***

## How model selection works

ChatCLI determines which model to use with the following priority (highest to lowest):

1. **`--model` flag** on the command line: `chatcli --model gpt-5.4`
2. **`/switch` command** during a session: `/switch --model claude-sonnet-4-6`
3. **`MODEL` environment variable**: sets the default model
4. **`LLM_PROVIDER` environment variable**: determines the provider (openai, anthropic, google, xai, etc.)
5. **Provider's default model**: each provider has a default model defined in the catalog

```bash theme={"system"}
# Example: set provider and model via .env
LLM_PROVIDER=CLAUDEAI
ANTHROPIC_MODEL=claude-sonnet-5-5
```

## Model aliases

Each model has **aliases** for easier typing. ChatCLI automatically resolves aliases to the canonical model ID. For example:

| Alias typed | Resolved model |
| :- | :- |
| `claude-4-5-sonnet` | `claude-sonnet-4-5` |
| `sonnet-4-5` | `claude-sonnet-4-5` |
| `opus-4-6` | `claude-opus-4-6` |
| `opus-4-7` | `claude-opus-4-7` |
| `opus-4-8` | `claude-opus-4-8` |
| `sonnet-5-5` / `sonnet-5.5` | `claude-sonnet-5-5` |
| `haiku-5-5` / `haiku-5.5` | `claude-haiku-5-5` |
| `sonnet-5` | `claude-sonnet-5` |
| `opus-5-5` / `opus-5.5` | `claude-opus-5-5` |
| `opus-5` | `claude-opus-5` |
| `fable` | `claude-fable-5-1` (tracks the newest Fable) |
| `fable-5-1` / `fable-5.1` | `claude-fable-5-1` |
| `fable-5` | `claude-fable-5` (pinned) |
| `glm-5-3` | `glm-5.3` |
| `glm-5-3-flashx` | `glm-5.3-flashx` |
| `glm-5-3-flash` | `glm-5.3-flash` |
| `glm-5-2` | `glm-5.2` |
| `gpt-6.1` / `gpt-6-1-sol` | `gpt-6.1-sol` |
| `gpt-6` | `gpt-6-astra` (family flagship) |
| `gpt-5.6` | `gpt-5.6-sol` (family flagship) |
| `gpt-5.4-pro` | `gpt-5.4` (Responses-only tier) |
| `gpt-5-mini` | `gpt-5` (mini variant) |
| `gemini-3.1-pro` | `gemini-3.1-pro-preview` |
| `grok-code-fast-1` | `grok-build-0.1` |
| `grok-build-latest` | `grok-4.5` |

Aliases are defined in the model catalog and accepted in all contexts: `--model`, `/switch`, and the `MODEL` variable.

## Catalog system

Models are registered in the `llm/catalog` package with complete metadata. ChatCLI uses the catalog to automatically determine:

* **API version** — which endpoint and protocol version to use for each model
* **Max tokens** — context and output limits for managing prompts and responses
* **Capabilities** — which features are available (vision, tools, JSON mode, etc.)
* **Provider-specific headers** — for example, the `anthropic-version` header varies per model

This means that when switching models, ChatCLI automatically adjusts all request parameters without manual configuration.

## Dynamic model listing

ChatCLI fetches available models **directly from each provider's API**, using the configured token or API key. This ensures you see exactly which models your account has access to — including new models not yet in the static catalog.

### How it works

1. When ChatCLI starts or when you switch providers (via `/switch`, `/auth login`, etc.), a background request queries the active provider's models endpoint
2. Discovered models are cached for use in the `/switch --model` **autocomplete**
3. Each suggestion indicates its origin: **`[API]`** (dynamic) or **`[catalog]`** (static)

### Endpoints per provider

| Provider | Endpoint | Auth |
| :- | :- | :- |
| OpenAI | `GET /v1/models` | API Key or OAuth |
| Anthropic | `GET /v1/models` | API Key or OAuth |
| Google AI | `GET /v1beta/models` | API Key |
| xAI | `GET /v1/models` | API Key |
| GitHub Copilot | `GET /models` | OAuth (Device Flow) |
| Ollama | `GET /api/tags` | No auth (local) |
| ZAI (Zhipu AI) | `GET /models` | API Key |
| MiniMax | `GET /models` | API Key |
| Moonshot (Kimi) | `GET /v1/models` | API Key (Bearer) |
| OpenRouter | `GET /api/v1/models` | API Key |
| StackSpot | — | Not supported (model fixed per agent) |

### Smart autocomplete

When typing `/switch --model` and pressing **Tab**, ChatCLI suggests available models:

```
> /switch --model [Tab]
  gpt-4o             GPT-4o (Copilot) [API]
  claude-sonnet-5.5  Claude Sonnet 5.5 (Copilot) [API]
  o4-mini            o4-mini (Copilot) [API]
```

If the API is not reachable, it falls back to the static catalog:

```
> /switch --model [Tab]
  gpt-4o           GPT-4o (Copilot) [catalog]
  gpt-4o-mini      GPT-4o mini (Copilot) [catalog]
```

<Tip>
  Pressing **Enter** with `/switch --model` (no value) lists all available models with source indication (API or catalog).
</Tip>

### OAuth and dynamic listing

Dynamic listing works with both **API key** and **OAuth**:

* **Anthropic OAuth**: uses `?beta=true` and Chrome-like headers, with automatic gzip decompression
* **OpenAI OAuth**: queries the ChatGPT backend (`/backend-api/models`) instead of the standard endpoint
* **GitHub Copilot OAuth**: uses the Device Flow token to query `api.githubcopilot.com/models`

After an `/auth login`, the model cache is automatically refreshed to reflect the new provider.

## Anthropic API versioning

Claude models may use different `anthropic-version` header values in API requests. The catalog manages this automatically:

* Newer models (claude-fable-5-1, claude-fable-5, claude-opus-5-5, claude-sonnet-5-5, claude-haiku-5-5, claude-opus-5, claude-sonnet-5, claude-opus-4-8, claude-opus-4-7, claude-sonnet-4-6) use the latest API version
* Every entry currently in the catalog ships with the default version; the retired Claude 3.x models that used older headers are no longer registered
* ChatCLI sends the correct header for each model without any user intervention


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.