/switch --model <name>.
Capabilities legend:
- 👁 Vision — accepts images as input
- 🔧 Tools — native tool use (function calling)
- 📋 JSON Mode — guaranteed structured JSON output
- 💻 Code Exec — native code execution on the provider
- OpenAI
- Anthropic (Claude)
- AWS Bedrock (full catalog)
- Google (Gemini)
- xAI (Grok)
- GitHub Copilot
- ZAI (Zhipu AI)
- MiniMax
- Moonshot (Kimi)
- StackSpot
- OpenRouter
- Devin CLI (Cognition)
- Ollama (Local)
Models ideal for code generation and complex reasoning. Support both Chat Completions API and Responses API.
| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
gpt-6.1-sol | gpt-6-1-sol, gpt-6.1 | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools (Responses API), 📋 JSON Mode |
gpt-6-astra | gpt-6 | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-6-sol | — | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-6-luna | — | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.6-sol | gpt-5.6 | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.6-terra | — | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.6-luna | — | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.5 | — | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.5-pro | — | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.4 | gpt-5.4-pro | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.4-mini | gpt-5.4-nano | 400K tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5.3-codex | — | 400K tokens | 128K tokens | 🔧 Tools, 📋 JSON Mode |
gpt-5.2 | gpt-5.2-pro | 400K tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-5 | gpt-5.1, gpt-5-mini, gpt-5-nano, gpt-5-pro | 400K tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
o3 / o3-mini / o4-mini | — | 200K tokens | 100K tokens | 🔧 Tools, 📋 JSON Mode, 🧠 Reasoning |
gpt-4.1 | gpt-4.1-mini, gpt-4.1-nano | 1.05M tokens | 32K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
gpt-4o | gpt-4o-mini | 128K tokens | 16K tokens | 👁 Vision, 🔧 Tools, 📋 JSON Mode |
GPT-6.1 Sol (
gpt-6.1-sol, GA Sep 29, 2026): 1,050,000-token context, 128K output, knowledge cutoff Apr 30, 2026. 2/10 per MTok, cached input 0.10(52.50. reasoning_effort accepts low, medium (default), high, xhigh and max — there is no none/minimal. Tool calling requires the Responses API (Chat Completions serves it without tools). There is no 6.1 Astra/Terra/Luna, and the -pro variant exists only on OpenRouter (not cataloged). Also on Copilot (gpt-6.1-sol), OpenRouter (openai/gpt-6.1-sol), Bedrock (global.openai.gpt-6.1-sol) and Devin (gpt-6.1-sol). The bare gpt-6 alias still resolves to gpt-6-astra. Effort hints (low/medium/high) now reach the whole GPT-6 family.GPT-5.6 (GA Jul 9, 2026) ships in three named tiers: Sol (flagship), Terra (balanced everyday) and Luna (fast and affordable) — 4/20, 2/12 and 0.20/1.20 per MTok respectively (Sol dropped from 5/30 to 4/20 on Aug 21, 2026 — promotional rate valid at least through Nov 21, 2026). All three work with an API key and with ChatGPT OAuth (
/auth login openai-codex); on the Codex backend ChatCLI sends the required client-identification headers automatically (without them the backend returns 404 for Luna).Specs re-verified Sep 2026:
gpt-5.4 runs the same 1.05M / 128K profile as 5.5/5.6 (its gpt-5.4-pro alias is Responses-only, 30/180); gpt-5.4-mini/-nano are a separate 400K / 128K entry; gpt-5.3-codex (Responses-only) and gpt-5.2 (+ gpt-5.2-pro, Responses-only, 21/168) are 400K / 128K. The aliases gpt-5.3, gpt-5.3-mini, gpt-5.3-nano, gpt-5.2-mini and gpt-5.2-nano were removed — no such models exist on the API — and so was gpt-5.3-codex-spark (Codex-app-only, never a platform id). Upcoming OpenAI shutdowns: o3-mini, o4-mini and gpt-4.1-nano on Oct 23, 2026; the gpt-5, gpt-5-mini, gpt-5-nano, gpt-5-pro and o3 snapshots on Dec 11, 2026; gpt-5.3-codex, gpt-5.4-nano and gpt-5.1 on Apr 1, 2027.Routing between Chat Completions and the Responses API is automatic per model via the catalog (
gpt-5.x, gpt-4.1 and o-series prefer Responses; gpt-4o stays on Chat Completions). Force Responses for every model with OPENAI_USE_RESPONSES=true. OAuth sessions always use the Responses API. Streaming is enabled for all models.Large context windows and excellent ability to follow complex instructions. All models support streaming via SSE.
Skill
| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
claude-fable-5-1 | fable-5-1, claude-fable-5.1, fable-5.1, fable | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (always on), ✉️ Mid-conv system |
claude-fable-5 (legacy) | fable-5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, ✉️ Mid-conv system |
claude-opus-5-5 | opus-5-5, claude-opus-5.5, opus-5.5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (always on), ⚡ Fast mode, ✉️ Mid-conv system |
claude-sonnet-5-5 | sonnet-5-5, claude-sonnet-5.5, sonnet-5.5, claude-5.5-sonnet | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (always on), ✉️ Mid-conv system |
claude-haiku-5-5 | haiku-5-5, claude-haiku-5.5, haiku-5.5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, ✉️ Mid-conv system |
claude-opus-5 | opus-5, claude-5-opus | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
claude-sonnet-5 | sonnet-5, claude-5-sonnet | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
claude-opus-4-8 | opus-4-8 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, ⚡ Fast mode, ✉️ Mid-conv system, 💾 1K-token cache floor |
claude-opus-4-7 | opus-4-7 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
claude-opus-4-6 | opus-4-6 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools |
claude-sonnet-4-6 | sonnet-4-6 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools |
claude-haiku-4-5-20251001 | claude-haiku-4-5, haiku-4-5 | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
claude-opus-4-5 | opus-4-5 | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
claude-sonnet-4-5 (deprecated, retires Nov 30, 2026) | claude-4-5-sonnet, sonnet-4-5 | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
Claude Fable 5.1 (
claude-fable-5-1, released Sep 1, 2026) is Anthropic’s most capable model — the successor to Fable 5 in the tier above Opus: 1M context, 128K output, 10/50 per MTok, with **cache reads at 0.25/MTok∗∗(2.51). Thinking is always on (adaptive): an explicit thinking:{type:"disabled"} or budget_tokens returns 400 — ChatCLI omits the field unless effort routing fires. Three breaking changes vs Fable 5: a forced tool_choice (any/tool) returns 400 (ChatCLI only ever sends auto), thinking blocks are bound to the model that produced them (other models drop them), and editing earlier turns invalidates thinking blocks. No fast mode, no Priority Tier; requires 30-day data retention. Shortcut: /model fable — the bare fable alias now tracks 5.1, while fable-5 stays pinned to Fable 5 (legacy, still served). Also on Bedrock as anthropic.claude-fable-5-1 and on OpenRouter as anthropic/claude-fable-5.1.Claude Fable 5 (claude-fable-5) is now legacy but still served. Same API surface as Opus 4.7/4.8 (adaptive thinking only, no temperature/top_p/top_k) with one extra constraint: an explicit thinking:{type:"disabled"} returns 400 — the field must be omitted to run without thinking (ChatCLI’s client already does). Also available on Bedrock as anthropic.claude-fable-5 (dateless ID — the new generation has no ARN-versioned IDs).Claude Sonnet 5 (claude-sonnet-5) is the Sonnet-tier successor (Anthropic skipped 4.7/4.8 for Sonnet): 1M context, 128K output, adaptive thinking, 2/10 per MTok — the launch rate became the permanent list price (Anthropic cancelled the Sep 1, 2026 increase to 3/15). Sonnet 5 does not support task budgets — ChatCLI no longer advertises the capability on it. On Bedrock it is served exclusively by the Messages endpoint (anthropic.claude-sonnet-5) — ChatCLI routes it automatically; see the AWS Bedrock tab.Claude Opus 5 (claude-opus-5, Jul 2026) succeeds Opus 4.8 for complex agentic coding and enterprise work: 1M context, 128K output, adaptive thinking (effort defaults to high server-side), 5/25 per MTok — the same price as Opus 4.5-4.8. Shortcut: /model opus-5. On Bedrock it is served through the Messages endpoint (anthropic.claude-opus-5), routed automatically like Sonnet 5; on OpenRouter the slug is anthropic/claude-opus-5.Claude Opus 5.5 (claude-opus-5-5, Sep 22 2026) succeeds Opus 5 in the Opus line at a lower price: 4/20 per MTok, cache reads 0.20(58/$40. 1M context, 128K output, thinking always on (adaptive — disabled/budget_tokens return 400), effort defaults to medium server-side (one level below Opus 5), forced tool_choice returns 400. It runs cyber and bio classifiers plus reasoning_extraction: a refusal arrives as HTTP 200 with stop_reason: refusal, and ChatCLI shows the category and the recommended model (see routing). Shortcut: /model opus-5.5. Through the Claude Code OAuth surface it requires release 2.1.280+ — ChatCLI presents 2.1.293 and learns from the claude_code_version_too_old error if the API asks for newer. On Bedrock it is anthropic.claude-opus-5-5 (Messages endpoint, global./us./eu./au. profiles); on OpenRouter the slug is anthropic/claude-opus-5.5; on Copilot claude-opus-5.5.Claude Sonnet 5.5 (claude-sonnet-5-5, Sep 28 2026) succeeds Sonnet 5 at the same 2/10 per MTok, with cache reads at 0.10(52.50 (5m) / $4 (1h). 1M context, 128K output (300K on the Batch API with the output-300k beta). Adaptive thinking is on by default and cannot be disabled (thinking disabled / budget_tokens return 400); effort ranges low..max, default high. A forced tool_choice (any/tool) returns 400 (ChatCLI only ever sends auto), and text the model writes between tool calls comes back as thinking blocks (hidden by default). Supports task budgets and mid-conversation system messages; no fast mode. Shortcut: /model sonnet-5.5. Through the Claude Code OAuth surface it requires release 2.1.284+. On Bedrock it is global.anthropic.claude-sonnet-5-5 (served by bedrock-runtime, not Mantle); on OpenRouter anthropic/claude-sonnet-5.5; on Copilot claude-sonnet-5.5.Claude Haiku 5.5 (claude-haiku-5-5, Oct 7 2026): 1M context, 128K output, priced by prompt length — 0.10/0.50 per MTok (cache read 0.01,100.50/2.50,cacheread0.05), and ChatCLI’s cost tracker applies that tier per call. Adaptive thinking is on by default and can be disabled only at effort high or below; effort defaults to medium. A forced tool_choice is accepted. It uses a newer tokenizer (~30% more tokens than Haiku 4.5 for the same text). Supports task budgets and mid-conversation system messages. Shortcut: /model haiku-5.5. Through the Claude Code OAuth surface it requires release 2.1.293+. On Bedrock it is global.anthropic.claude-haiku-5-5; on OpenRouter anthropic/claude-haiku-5.5; on Copilot claude-haiku-5.5.claude-opus-4-8 and claude-opus-4-7 ship with 1M native context (no extra flag). claude-opus-4-6 can also use 1M context by setting ANTHROPIC_1MTOKENS_SONNET=true. claude-sonnet-4-6 now allows 128K output on the Claude API (was 64K before Mar 2026; Bedrock still caps it at 64K). Different models may use distinct anthropic-version headers, managed automatically by the catalog.Retired on the Claude API (removed from the catalog, Sep 2026): Opus 4.1 (Aug 5, 2026), Opus 4 and Sonnet 4 (Jun 15, 2026) and the whole Claude 3.x line (Sonnet 3.7, Sonnet 3.5, Haiku 3.5, Opus 3, Haiku 3). Requests naming them fail server-side, so their aliases (
opus-4-1, opus-4, sonnet-4, claude-3-7-sonnet, …) no longer resolve. There was never a Sonnet 4.7 — Anthropic went from Sonnet 4.6 straight to Sonnet 5, and the sonnet-4-7 placeholder was dropped. Bedrock keeps its own lifecycle (see the AWS Bedrock tab).Deprecated: Sonnet 4.5 (claude-sonnet-4-5) was deprecated on the Claude API on Sep 30, 2026 and retires Nov 30, 2026 — the replacement is claude-sonnet-5-5.Catalog order: entries are declared newest-first in the registry. The alias resolver matches by prefix in registry order, so an older entry whose alias is a prefix of a newer id (fable-5 ⊂ fable-5-1, opus-4-5 ⊂ opus-4-5-…) must come after the newer one — otherwise fable-5-1 would silently resolve to Fable 5 (wrong cache price, wrong capability flags). If you add a new Claude generation, keep this newest-first order.Claude Opus 4.8 — what’s new
Released May 28, 2026. Same default 1M / 128K profile as Opus 4.7 but with four new launch capabilities the catalog tracks as feature flags:| Capability | What it means |
|---|---|
adaptive_thinking | Only thinking mode accepted by 4.7+. ChatCLI emits thinking:{type:"adaptive"} when a skill provides an effort: hint — the model decides per turn whether to reason. Sending budget_tokens returns HTTP 400. |
fast_mode | Research-preview faster output (~2.5× tokens/sec) at premium pricing. Opt in with ANTHROPIC_SPEED=fast. |
mid_conversation_system | Server accepts role:"system" after the first user turn, preserving prompt-cache hits across instruction updates. ChatCLI’s message builder already passes structured system blocks through unchanged. |
low_cache_minimum | Minimum cacheable prompt drops from previous models’ floor to 1,024 tokens. Prompts that didn’t qualify on 4.7 now create cache entries with no code change. |
effort: medium|high|max continues to work — on Opus 4.7+, Sonnet 5, Sonnet 5.5, Haiku 5.5, Opus 5, Opus 5.5 and Fable 5/5.1 it maps to adaptive thinking automatically; on older 4.x (4.5/4.6) it falls back to budgeted extended thinking (thinking:{type:"enabled", budget_tokens:N}).Full AWS Bedrock catalog — Anthropic, OpenAI, Llama, Nova, Mistral, Cohere, AI21, DeepSeek, Moonshot Kimi, MiniMax, Qwen, Z.AI/GLM, Gemma, Nemotron, TwelveLabs, and any provider AWS adds. Auth uses the AWS SDK’s default credentials chain (IAM role,
OpenAI GPT-5.6 (frontier, Jul 2026) — Sol, Terra and Luna on Bedrock. They speak Converse, Responses and Chat Completions but not
OpenAI GPT-OSS (open-weights) — OpenAI models hosted on Bedrock. Use the OpenAI Chat Completions schema (auto-detected by
xAI Grok and Amazon Nova 2 (static entries) — Grok 4.6 landed on Bedrock on Aug 18, 2026 (Converse,
Moonshot Kimi K3 and Z.AI GLM-5.3 (static entries) — Kimi K3 is served through
Other providers via Converse API — beyond the static entries above, Llama, Nova, Mistral, Cohere, AI21, DeepSeek, Moonshot Kimi, MiniMax, Qwen, Z.AI/GLM, Gemma, Nemotron, TwelveLabs, etc. are not hardcoded in the catalog — they appear dynamically in
Embeddings via Bedrock —
~/.aws/credentials, env vars) — no API key from the original providers is needed.Modern models (Claude 4.x/4.5/4.6 and equivalents from other providers) do not accept direct on-demand invocation by base ID — they require an inference profile ID (prefixes
global., us., eu., apac.). ChatCLI automatically filters non-invokable base IDs from /switch --model, so only what works appears. See AWS Bedrock for details.New generation = dateless IDs. Fable 5.1, Fable 5, Opus 5.5, Sonnet 5.5, Haiku 5.5, Opus 5, Sonnet 5, Opus 4.8 and Opus 4.7 have no ARN-versioned IDs on Bedrock — the IDs are
anthropic.claude-fable-5-1, anthropic.claude-fable-5, anthropic.claude-opus-5-5, global.anthropic.claude-sonnet-5-5, global.anthropic.claude-haiku-5-5, anthropic.claude-opus-5, anthropic.claude-sonnet-5, anthropic.claude-opus-4-8 and anthropic.claude-opus-4-7. Opus 4.8/4.7 are invoked through the global. inference profile (the bare ID is not on-demand invokable on InvokeModel). The old dated IDs (e.g. global.anthropic.claude-opus-4-8-20260528-v1:0) keep resolving as aliases. Opus 5.5, Opus 5, Sonnet 5, Fable 5 and Fable 5.1 are served by the Messages endpoint (bedrock-mantle.{region}.api.aws/anthropic/v1/messages) — ChatCLI detects and routes it automatically (SigV4 bedrock-mantle or AWS_BEARER_TOKEN_BEDROCK), and if the Mantle call fails it retries the same request through the legacy InvokeModel runtime under the global. inference profile; see AWS Bedrock.Sonnet 5.5 and Haiku 5.5 are the exception: they are not Mantle-only — bedrock-runtime serves them (Messages, Converse and InvokeModel). There is no in-region option, so the canonical ID is the global. inference profile (global.anthropic.claude-sonnet-5-5, global.anthropic.claude-haiku-5-5); bare first-party IDs are upgraded automatically.| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
anthropic.claude-fable-5-1 | bedrock-fable-5-1, global.anthropic.claude-fable-5-1, us.anthropic.claude-fable-5-1, claude-fable-5-1, fable-5-1 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint, 30-day retention) |
anthropic.claude-fable-5 | bedrock-fable-5, claude-fable-5, fable-5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint, 30-day retention) |
anthropic.claude-opus-5-5 | bedrock-opus-5-5, global./us./eu./au.anthropic.claude-opus-5-5, claude-opus-5-5, opus-5.5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint) |
anthropic.claude-opus-5 | bedrock-opus-5, claude-opus-5, opus-5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint) |
global.anthropic.claude-sonnet-5-5 | anthropic.claude-sonnet-5-5, us./eu.anthropic.claude-sonnet-5-5, claude-sonnet-5-5, sonnet-5.5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (bedrock-runtime) |
global.anthropic.claude-haiku-5-5 | us./eu./au./jp.anthropic.claude-haiku-5-5, claude-haiku-5-5, haiku-5.5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (bedrock-runtime) |
anthropic.claude-sonnet-5 | bedrock-sonnet-5, claude-sonnet-5, sonnet-5 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking (Messages endpoint) |
global.anthropic.claude-opus-4-8 | bedrock-opus-4-8, claude-opus-4-8 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking, 💾 1K-token cache floor |
global.anthropic.claude-opus-4-7 | bedrock-opus-4-7, claude-opus-4-7 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 🧠 Adaptive thinking |
global.anthropic.claude-sonnet-4-6 | bedrock-sonnet-4-6, claude-sonnet-4-6 | 1M tokens | 64K tokens (AWS card) | 👁 Vision, 🔧 Tools |
global.anthropic.claude-opus-4-6-v1 | bedrock-opus-4-6, claude-opus-4-6 | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools |
global.anthropic.claude-haiku-4-5-20251001-v1:0 | bedrock-haiku-4-5, claude-haiku-4-5 | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
global.anthropic.claude-sonnet-4-5-20250929-v1:0 | bedrock-sonnet-4-5, claude-sonnet-4-5 | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
us.anthropic.claude-sonnet-4-5-20250929-v1:0 | bedrock-sonnet-4-5-us | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
global.anthropic.claude-opus-4-5-20251101-v1:0 | bedrock-opus-4-5, claude-opus-4-5, global.anthropic.claude-opus-4-5-20251001-v1:0 (old misspelling, alias only) | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
global.anthropic.claude-sonnet-4-20250514-v1:0 (Legacy, EOL Oct 14, 2026) | bedrock-sonnet-4, claude-sonnet-4 | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
us.anthropic.claude-sonnet-4-20250514-v1:0 (Legacy, EOL Oct 14, 2026) | bedrock-sonnet-4-us | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
eu.anthropic.claude-sonnet-4-20250514-v1:0 (Legacy, EOL Oct 14, 2026) | bedrock-sonnet-4-eu | 200K tokens | 64K tokens | 👁 Vision, 🔧 Tools |
us.anthropic.claude-opus-4-1-20250805-v1:0 (Legacy, extended access since Oct 8, 2026, EOL Jan 8, 2027) | bedrock-opus-4-1, claude-opus-4-1 | 200K tokens | 32K tokens | 👁 Vision, 🔧 Tools |
Retired on Bedrock and removed from the catalog (Sep 2026): Opus 4 (
us.anthropic.claude-opus-4-20250514-v1:0), Sonnet 3.7 (us./eu.anthropic.claude-3-7-sonnet-20250219-v1:0), Sonnet 3.5 v1/v2, Haiku 3.5, Opus 3 and Haiku 3 (anthropic.claude-3-haiku-20240307-v1:0, EOL Sep 10, 2026). The placeholder global.anthropic.claude-sonnet-4-7-20260401-v1:0 never existed on AWS and was dropped. Sonnet 4 and Opus 4.1 are Legacy but still invokable until their EOL dates above; Opus 4.1 entered the higher-priced extended-access period on Oct 8, 2026. Opus 4.5’s real snapshot date is 20251101 — the previous 20251001 spelling never existed on AWS and is kept only as an alias.InvokeModel, so the catalog flags them bedrock_converse_only and ChatCLI routes them through the Converse API automatically (no env var needed — the openai. prefix alone would have sent them down the GPT-OSS InvokeModel path). Served only through inference profiles: Sol via us./global., Terra/Luna via us./in./global.. Global-endpoint prices: Sol 4/20, Terra 2/12, Luna 0.20/1.20 per MTok. GPT-6.1 Sol (global.openai.gpt-6.1-sol, us. profile) takes the same Converse path: 1M context, 131,072 max output, 2/10 per MTok (cache write 2.50,cacheread0.10).| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
global.openai.gpt-6.1-sol | us.openai.gpt-6.1-sol | 1M tokens | 131,072 tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
global.openai.gpt-6-astra | bedrock-gpt-6-astra, openai.gpt-6-astra, us.openai.gpt-6-astra | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
global.openai.gpt-6-sol | bedrock-gpt-6-sol, openai.gpt-6-sol, us.openai.gpt-6-sol | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
global.openai.gpt-6-luna | bedrock-gpt-6-luna, openai.gpt-6-luna, us.openai.gpt-6-luna | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
global.openai.gpt-5.6-sol | bedrock-gpt-5.6-sol, openai.gpt-5.6-sol, us.openai.gpt-5.6-sol | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
global.openai.gpt-5.6-terra | bedrock-gpt-5.6-terra, openai.gpt-5.6-terra, us.openai.gpt-5.6-terra, in.openai.gpt-5.6-terra | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
global.openai.gpt-5.6-luna | bedrock-gpt-5.6-luna, openai.gpt-5.6-luna, us.openai.gpt-5.6-luna, in.openai.gpt-5.6-luna | 1.05M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON (Converse only) |
openai.* prefix, or forced via BEDROCK_PROVIDER=openai).| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
openai.gpt-oss-120b-1:0 | bedrock-gpt-oss-120b, gpt-oss-120b | 128K tokens | 16K tokens | 🔧 Tools, 📋 JSON |
openai.gpt-oss-20b-1:0 | bedrock-gpt-oss-20b, gpt-oss-20b | 128K tokens | 16K tokens | 🔧 Tools, 📋 JSON |
us./global. profiles only, 2/6 per MTok on the global endpoint), followed by Grok 4.7 (global.xai.grok-4.7, us. profile, 500K context, 2/6 per MTok, cached input $0.50). Nova 2 Lite succeeds Nova Premier (Legacy since Mar 13, 2026, EOL Sep 14, 2026 — removed from the catalog): 1M context, 64K output, served only through inference profiles (us./eu./jp./global. — there is no bare in-region id). Nova 2 Pro/Omni are not on Bedrock.| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
global.xai.grok-4.7 | us.xai.grok-4.7 | 500K tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
global.xai.grok-4.6 | bedrock-grok-4.6, xai.grok-4.6, us.xai.grok-4.6 | 500K tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
global.amazon.nova-2-lite-v1:0 | nova-2-lite, amazon.nova-2-lite-v1:0, us./eu./jp.amazon.nova-2-lite-v1:0 | 1M tokens | 64K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
us./in. profiles (1M context, image input, 3/15 per MTok, cache read $0.30); GLM-5.3 through the us. profile (1M context, 128K output, text only — the AWS card publishes no price, so ChatCLI’s cost tracker uses the Z.AI list price).| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
global.moonshotai.kimi-k3 | us./in.moonshotai.kimi-k3 | 1M tokens | 131K tokens | 👁 Vision, 🔧 Tools |
global.zai.glm-5.3 | us.zai.glm-5.3 | 1M tokens | 128K tokens | 🔧 Tools |
/switch --model based on what your AWS account has access to. ChatCLI routes these models through the AWS Converse API (unified schema), so adding a new provider doesn’t require a release.Examples of IDs seen via ListFoundationModels (your actual list depends on account + region):| Provider | Example Model ID |
|---|---|
| Moonshot AI | moonshotai.kimi-k2.6, moonshotai.kimi-k2-thinking |
| MiniMax | minimax.m-2-5, minimax.m-2 |
| Z.AI | zai.glm-4-7, zai.glm-4-7-flash |
| Qwen | qwen.qwen3-32b, qwen.qwen3-coder-480b |
| Meta Llama | meta.llama3-70b-instruct-v1:0, us.meta.llama3-1-70b-... |
| Amazon Nova | amazon.nova-pro-v1:0, amazon.nova-lite-v1:0, global.amazon.nova-2-lite-v1:0 |
| Mistral | mistral.mistral-large-2407-v1:0 |
| DeepSeek | us.deepseek.r1-v1:0 |
google.gemma-3-27b-pt | |
| NVIDIA | nvidia.nemotron-nano-9b-v2 |
| TwelveLabs | twelvelabs.pegasus-v1.2 |
The dynamic listing (
/switch --model) merges bedrock:ListFoundationModels (filtered by ByOutputModality: TEXT + InferenceTypesSupported: ON_DEMAND) and bedrock:ListInferenceProfiles with the static catalog above. No allowlist — any Bedrock provider your account can access shows up automatically. Use the command to see what your AWS account can actually invoke in the configured region.amazon.titan-embed-text-v2:0 (default, 1024-dim, configurable 256/512/1024), amazon.titan-embed-text-v1 (1536-dim), Cohere cohere.embed-english-v3 / cohere.embed-multilingual-v3 (1024-dim), cohere.embed-v4:0 (1536-dim default, 128K context, also via us./eu./global. profiles), and amazon.nova-2-multimodal-embeddings-v1:0 (Nova MME, 3072-dim default, configurable 256/384/1024/3072). Enable with CHATCLI_EMBED_PROVIDER=bedrock. See RAG + HyDE.Advanced multimodal capabilities and massive context windows. Support streaming via SSE.
| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
gemini-3.8-flash | gemini-3.8-flash-latest | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
gemini-3.7-flash | gemini-3.7-flash-latest | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
gemini-3.6-flash | gemini-3.6-flash-latest | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
gemini-3.5-flash | gemini-3.5-flash-latest | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
gemini-3.5-flash-lite | — | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
gemini-3.1-pro-preview | gemini-3.1-pro | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
gemini-3.1-flash-lite | — | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
gemini-3-flash-preview | — | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
gemini-2.5-pro | gemini-2.5-pro-latest | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON, 💻 Code Exec |
gemini-2.5-flash | — | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
gemini-2.5-flash-lite | — | 1M tokens | 65K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
The whole Gemini 3.x generation shares the 1,048,576-token input window with a 65,536-token output cap (
gemini-3.8-flash, GA Sep 2 2026, is the current workhorse; gemini-3.7-flash stays served). gemini-3.1-pro without -preview is only an alias, not a real model code. Gemini 2.5 Flash Lite also supports Multimodal Live for real-time interactions; models with JSON Mode can return structured output via response_mime_type.Shut down by Google and removed from the catalog:
gemini-2.0-flash and gemini-2.0-flash-lite (Jun 1, 2026), and gemini-3 / gemini-3-pro / gemini-3-pro-preview (Mar 9, 2026 — use gemini-3.1-pro-preview). Next scheduled: gemini-3.1-flash-lite retires May 7, 2027 — migrate to gemini-3.5-flash-lite. The introductory price of 3.7/3.6 Flash (0.75/3.75) doubles on Jan 1, 2027.Real-time information integration and large context windows. Support streaming.
| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
grok-4.7 | grok-4.7-latest | 500K tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
grok-4.6 | grok-4.6-latest | 500K tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
grok-4.5 | grok-4.5-latest, grok-build-latest | 500K tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
grok-4.3 | grok-4.3-latest | 1M tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
grok-4.20-0309-reasoning | grok-4.20-reasoning, grok-4.20 | 1M tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
grok-4.20-0309-non-reasoning | grok-4.20-non-reasoning | 1M tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
grok-4.20-multi-agent-0309 | grok-4.20-multi-agent | 1M tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
grok-build-0.1 | grok-build, grok-code-fast-1 | 256K tokens | — | 👁 Vision, 🔧 Tools, 📋 JSON |
Grok models use the OpenAI-compatible API. xAI publishes no per-model output cap, so limits are managed by the provider.
grok-4.7 (Sep 21 2026) is the current flagship — 500K context, low|medium|high|xhigh effort, the same 2/6 tier as 4.6 (cached input $0.50; once the prompt reaches 200K xAI bills 2× on input and output, and ChatCLI’s cost tracker applies it per call). grok-4.6 stays served. The grok-4.7-fast variant is Cursor/Grok Build only, not on the public API, and therefore not in the catalog. The provider default (XAI_MODEL) is now grok-4.3.Retired by xAI on May 15, 2026 and removed from the catalog: grok-4-fast (+ grok-4-fast-reasoning*, grok-4-0709), grok-4-1 (+ grok-4-1-fast*), grok-3 and grok-3-mini. The API keeps those slugs alive but redirects them to grok-4.3 and bills grok-4.3 rates — pinned configs keep working, and ChatCLI’s cost tracker prices them accordingly. grok-code-fast-1 became an alias of grok-build-0.1. Cached-input pricing is tracked too: 0.50/MTokongrok−4.7/4.6,0.30 on grok-4.5, $0.20 on the 4.3/4.20 tier.Use models from the Copilot platform with your subscription (Individual, Business, Enterprise). Authenticate via
/auth login github-copilot.The table below shows models registered in the static catalog. With dynamic listing, ChatCLI queries the Copilot API and automatically discovers all models available for your account.| Model (ID) | Context |
|---|---|
gpt-6.1-sol | 1M tokens |
gpt-6-astra | 1M tokens |
gpt-6-sol | 1M tokens |
gpt-6-luna | 1M tokens |
claude-sonnet-5.5 | 1M tokens |
claude-haiku-5.5 | 1M tokens |
claude-opus-5.5 | 1M tokens |
gpt-4o | 128K tokens |
gpt-4o-mini | 128K tokens |
| + dynamic models | via API |
Available models vary depending on your plan and region. Use
/switch --model to see the full list fetched directly from the Copilot API.Retired by GitHub and removed from the static catalog:
claude-sonnet-4 (May 1, 2026) and gemini-2.0-flash (Oct 23, 2025). GitHub has also scheduled GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, Gemini 3.7 Flash and Grok 4.5 for retirement on Oct 19, 2026.Chinese AI models from Zhipu AI (z.ai) with strong multilingual and coding capabilities. OpenAI-compatible API with native tool calling support.
| Model (ID) | Context | Max Output | Capabilities |
|---|---|---|---|
glm-5.3 | 1M tokens | 128K tokens | 🔧 Tools, 📋 JSON |
glm-5.3-flashx | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
glm-5.3-flash | 1M tokens | 128K tokens | 👁 Vision, 🔧 Tools, 📋 JSON |
glm-5.2 | 1M tokens | 128K tokens | 🔧 Tools, 📋 JSON |
glm-5.1 | 200K tokens | 128K tokens | 👁 Vision, 🔧 Tools |
glm-5-turbo | 200K tokens | 128K tokens | 👁 Vision, 🔧 Tools |
glm-5 | 200K tokens | 128K tokens | 👁 Vision, 🔧 Tools |
glm-4.7 | 200K tokens | 128K tokens | 🔧 Tools |
glm-4.6 | 200K tokens | 128K tokens | 🔧 Tools |
glm-4.5 | 128K tokens | 96K tokens | 🔧 Tools |
glm-4.5-flash | 128K tokens | 16K tokens | 🔧 Tools |
glm-5v-turbo | 128K tokens | 16K tokens | 👁 Vision, 🔧 Tools |
glm-4.5v | 128K tokens | 16K tokens | 👁 Vision |
GLM-5.3 (released Aug 18, 2026) is Zhipu’s open-weight flagship: same base model as GLM-5.2 with scaled-up post-training focused on coding and long-horizon agentic work — 1M-token context, 128K output, reasoning always on, function calling and structured output (text-only input). GLM-5.3-Flash (Aug 26, 2026) is the natively multimodal MIT-licensed sibling: 1M context, 128K output, image/video/file input. List prices: GLM-5.3 1.40/4.40, GLM-5.3-Flash 0.15/0.50, GLM-5.3-FlashX (Sep 18 2026, the high-speed tier of Flash) 0.37/1.25 per MTok (GLM-5: 1.00/3.20) — ChatCLI’s cost tracker uses these rates, and it now prices the GLM-4.x line per tier as well (4.7 0.60/2.20, 4.7-flashx 0.07/0.40, 4.7-flash free, 4.5-air 0.20/1.10, 4.5-airx 1.10/4.50, 4.5-x 2.20/8.90, 4.5v 0.60/1.80, 4.5-flash free) instead of the old flat $0.50 fallback. The provider default model remains
glm-5; use /switch --model glm-5.3 to switch. codegeex-4 was removed — it is no longer served by the Z.AI international API.ZAI uses an OpenAI-compatible API at
https://api.z.ai/api/paas/v4/chat/completions. Authentication is via ZAI_API_KEY Bearer token. Model IDs are case-sensitive. Subscribers of the GLM Coding Plan can set ZAI_USE_CODING_PLAN=true to use the subscription endpoint (/api/coding/paas/v4) with the same key — usage draws from the plan instead of pay-as-you-go credits and /cost reports it at $0.Automatic JWT authentication: Keys in
id.secret format automatically enable JWT token rotation (HMAC-SHA256), cached for 30 minutes. Keys without ”.” work as traditional Bearer tokens. No additional configuration needed.High-performance models from MiniMax with large context windows and native tool calling. OpenAI-compatible API.
| Model (ID) | Context | Max Output | Capabilities |
|---|---|---|---|
MiniMax-M3 | 1M tokens | 131K tokens | 👁 Vision, 🔧 Tools |
MiniMax-M2.7 | 204K tokens | 131K tokens | 👁 Vision, 🔧 Tools |
MiniMax-M2.7-highspeed | 204K tokens | — | 👁 Vision, 🔧 Tools |
MiniMax-M2.5 | 196K tokens | 65K tokens | 👁 Vision, 🔧 Tools |
MiniMax-M2.5-highspeed | 196K tokens | — | 👁 Vision, 🔧 Tools |
MiniMax-Text-01 | 128K tokens | 2K tokens | 📋 JSON Mode |
MiniMax model IDs are case-sensitive (e.g.,
MiniMax-M2.7, not minimax-m2.7). The API uses a base_resp field for error handling with status_code and status_msg.Anthropic-compatible endpoint: Set
MINIMAX_API_COMPAT=anthropic to use https://api.minimax.io/anthropic/v1/messages with Anthropic Messages format. Native tool calling is disabled in this mode (falls back to XML). Model listing always uses the native endpoint.Moonshot AI’s Kimi family — the K3 flagship (Jul 2026) is a 2.8T-parameter MoE with 104B activated, 1M-token context via Kimi Delta Attention and multimodal input; K2.6 (1T/32B, 256K) remains fully supported, with the native MoonViT vision encoder and explicit “thinking” mode across the line. OpenAI-compatible API at
https://api.moonshot.ai/v1/chat/completions.| Model (ID) | Aliases | Context | Max Output | Capabilities |
|---|---|---|---|---|
kimi-k3 | kimi-k-3, k3, k-3 | 1M tokens | 131K tokens | 🔧 Tools, 👁 Vision, 🧠 Thinking, 📋 JSON Mode |
kimi-k2.7-code | kimi-k2.7, k2.7 | 256K tokens | 32K tokens | 🔧 Tools, 🧠 Thinking, 📋 JSON Mode |
kimi-k2.7-code-highspeed | kimi-k2-7-code-highspeed | 256K tokens | 32K tokens | 🔧 Tools, 🧠 Thinking, 📋 JSON Mode |
kimi-k2.6 | kimi-k2-6, k2.6, k2-6 | 256K tokens | 131K tokens | 🔧 Tools, 👁 Vision, 🧠 Thinking, 📋 JSON Mode |
Thinking mode: Set
MOONSHOT_THINKING=enabled|disabled|auto to toggle between Thinking (explicit reasoning, default for K3/K2.6) and Instant (direct response, cheaper). The default auto lets the model choose; models without the thinking capability ignore the flag.Public pricing (Aug 2026): kimi-k3 is 3.00/Minputtokens(cachemiss;cachehit0.30/M) and 15.00/Moutput.kimi−k2.7−codeandkimi−k2.6are0.95/M input (cache miss) and 4.00/Moutput;kimi−k2.7−code−highspeedis2×that(1.90/$8.00). Cache-hit input is cheaper on all of them, and the ChatCLI cost tracker prices the cached slice at each model’s rate (10% of input on K3, 20% on K2.7 Code, ~17% on K2.6). Retired by Moonshot and removed from the catalog (the API now returns 404 for them):
kimi-k2.5 and the moonshot-v1-128k/32k/8k series (Aug 31, 2026), kimi-k2-turbo-preview (May 25, 2026), kimi-latest (Jan 28, 2026) and kimi-thinking-preview (Nov 11, 2025). Migration target for all of them is kimi-k3; the provider default stays kimi-k2.6.Accepts all compatible models on the StackSpotAI platform, selected during Agent creation.
Multi-provider API gateway — access 200+ models from OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, and more through a single API key. Uses an OpenAI-compatible API at
https://openrouter.ai/api/v1/chat/completions.Models use the provider/model-name format:| Model (ID) | Provider | Capabilities |
|---|---|---|
openai/gpt-6.1-sol | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
openai/gpt-6-astra | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
openai/gpt-6-sol | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
openai/gpt-6-luna | OpenAI | 👁 Vision, 🔧 Tools, 📋 JSON |
openai/gpt-4o | OpenAI | 👁 Vision, 🔧 Tools |
openai/gpt-4o-mini | OpenAI | 👁 Vision, 🔧 Tools |
anthropic/claude-opus-5.5 | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
anthropic/claude-opus-5 | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
anthropic/claude-sonnet-5.5 | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
anthropic/claude-haiku-5.5 | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
anthropic/claude-sonnet-5 | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
anthropic/claude-fable-5.1 | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
anthropic/claude-fable-5 | Anthropic | 👁 Vision, 🔧 Tools, 📋 JSON |
anthropic/claude-sonnet-4 | Anthropic | 👁 Vision, 🔧 Tools |
anthropic/claude-opus-4 | Anthropic | 👁 Vision, 🔧 Tools |
google/gemini-2.5-pro | 👁 Vision, 🔧 Tools, 📋 JSON | |
google/gemini-2.5-flash | 👁 Vision, 🔧 Tools, 📋 JSON | |
meta-llama/llama-4-maverick | Meta | 🔧 Tools |
deepseek/deepseek-r1 | DeepSeek | 🔧 Tools |
mistralai/mistral-large | Mistral | 🔧 Tools |
The table above shows popular defaults. OpenRouter provides 200+ models — ChatCLI discovers them dynamically via the
/api/v1/models endpoint. Use /switch --model to browse the full list.OpenRouter supports native fallback routing via
OPENROUTER_FALLBACK_MODELS. If your primary model is unavailable, OpenRouter automatically routes to the next model in the list — handled server-side before ChatCLI’s own fallback chain kicks in.Served through the local Devin CLI wrapper — ChatCLI keeps the whole conversation and harness; Devin is only the transport. See Devin Provider. Family slugs use dots (
claude-sonnet-4.6, not 4-6); variant ids keep the CLI spelling (claude-opus-5-high, swe-1-6-fast). /switch --model lists what your account can invoke via devin models list --format json (tagged [api]); the table below is the static fallback.| Family | Models |
|---|---|
| Anthropic | claude-fable-5.1 · claude-fable-5 · claude-opus-5.5 · claude-opus-5 · claude-sonnet-5 · claude-opus-4.8 / 4.7 / 4.6 / 4.5 · claude-sonnet-4.6 / 4.5 / 4 · claude-haiku-4.5 |
| OpenAI | gpt-6.1-sol · gpt-6-astra / -sol / -luna · gpt-5.6-sol / -terra / -luna · gpt-5.5 · gpt-5.4 / -mini · gpt-5.3-codex · gpt-5.2 · gpt-5.1 · gpt-4.1 |
gemini-3.8-flash · gemini-3.7-flash · gemini-3.6-flash · gemini-3.5-flash · gemini-3.1-pro · gemini-3-flash | |
| xAI | grok-4.6 · grok-4.5 |
| Others | glm-5.3 / 5.2 · kimi-k3 / k2.7 / k2.6 · deepseek-v4-pro / v4-flash |
| Cognition (SWE) | swe-1.7-lightning · swe-1.7 · swe-1.6-fast · swe-1.6 |
Auth belongs to the binary (
devin auth login, corporate SSO) — no key in ChatCLI. Default model: claude-sonnet-4.6 (DEVIN_MODEL). The CLI reports no token usage, so cost tracking shows zero — cost lives in the Cognition subscription.Supports any local model via Ollama. Configure in Or switch interactively:
.env:OLLAMA_ENABLED=true
OLLAMA_MODEL="llama3"
/switch --model llama3Use ollama pull <model> to download new models.How model selection works
ChatCLI determines which model to use with the following priority (highest to lowest):--modelflag on the command line:chatcli --model gpt-5.4/switchcommand during a session:/switch --model claude-sonnet-4-6MODELenvironment variable: sets the default modelLLM_PROVIDERenvironment variable: determines the provider (openai, anthropic, google, xai, etc.)- Provider’s default model: each provider has a default model defined in the catalog
# Example: set provider and model via .env
LLM_PROVIDER=anthropic
MODEL=claude-sonnet-4-6
Model aliases
Each model has aliases for easier typing. ChatCLI automatically resolves aliases to the canonical model ID. For example:| Alias typed | Resolved model |
|---|---|
claude-4-5-sonnet | claude-sonnet-4-5 |
sonnet-4-5 | claude-sonnet-4-5 |
opus-4-6 | claude-opus-4-6 |
opus-4-7 | claude-opus-4-7 |
opus-4-8 | claude-opus-4-8 |
sonnet-5-5 / sonnet-5.5 | claude-sonnet-5-5 |
haiku-5-5 / haiku-5.5 | claude-haiku-5-5 |
sonnet-5 | claude-sonnet-5 |
opus-5-5 / opus-5.5 | claude-opus-5-5 |
opus-5 | claude-opus-5 |
fable | claude-fable-5-1 (tracks the newest Fable) |
fable-5-1 / fable-5.1 | claude-fable-5-1 |
fable-5 | claude-fable-5 (pinned) |
glm-5-3 | glm-5.3 |
glm-5-3-flashx | glm-5.3-flashx |
glm-5-3-flash | glm-5.3-flash |
glm-5-2 | glm-5.2 |
gpt-6.1 / gpt-6-1-sol | gpt-6.1-sol |
gpt-6 | gpt-6-astra (family flagship) |
gpt-5.6 | gpt-5.6-sol (family flagship) |
gpt-5.4-pro | gpt-5.4 (Responses-only tier) |
gpt-5-mini | gpt-5 (mini variant) |
gemini-3.1-pro | gemini-3.1-pro-preview |
grok-code-fast-1 | grok-build-0.1 |
grok-build-latest | grok-4.5 |
--model, /switch, and the MODEL variable.
Catalog system
Models are registered in thellm/catalog package with complete metadata. ChatCLI uses the catalog to automatically determine:
- API version — which endpoint and protocol version to use for each model
- Max tokens — context and output limits for managing prompts and responses
- Capabilities — which features are available (vision, tools, JSON mode, etc.)
- Provider-specific headers — for example, the
anthropic-versionheader varies per model
Dynamic model listing
ChatCLI fetches available models directly from each provider’s API, using the configured token or API key. This ensures you see exactly which models your account has access to — including new models not yet in the static catalog.How it works
- When ChatCLI starts or when you switch providers (via
/switch,/auth login, etc.), a background request queries the active provider’s models endpoint - Discovered models are cached for use in the
/switch --modelautocomplete - Each suggestion indicates its origin:
[API](dynamic) or[catalog](static)
Endpoints per provider
| Provider | Endpoint | Auth |
|---|---|---|
| OpenAI | GET /v1/models | API Key or OAuth |
| Anthropic | GET /v1/models | API Key or OAuth |
| Google AI | GET /v1beta/models | API Key |
| xAI | GET /v1/models | API Key |
| GitHub Copilot | GET /models | OAuth (Device Flow) |
| Ollama | GET /api/tags | No auth (local) |
| ZAI (Zhipu AI) | GET /models | API Key |
| MiniMax | GET /models | API Key |
| Moonshot (Kimi) | GET /v1/models | API Key (Bearer) |
| OpenRouter | GET /api/v1/models | API Key |
| StackSpot | — | Not supported (model fixed per agent) |
Smart autocomplete
When typing/switch --model and pressing Tab, ChatCLI suggests available models:
> /switch --model [Tab]
gpt-4o GPT-4o (Copilot) [API]
claude-sonnet-5.5 Claude Sonnet 5.5 (Copilot) [API]
o4-mini o4-mini (Copilot) [API]
> /switch --model [Tab]
gpt-4o GPT-4o (Copilot) [catalog]
gpt-4o-mini GPT-4o mini (Copilot) [catalog]
Pressing Enter with
/switch --model (no value) lists all available models with source indication (API or catalog).OAuth and dynamic listing
Dynamic listing works with both API key and OAuth:- Anthropic OAuth: uses
?beta=trueand Chrome-like headers, with automatic gzip decompression - OpenAI OAuth: queries the ChatGPT backend (
/backend-api/models) instead of the standard endpoint - GitHub Copilot OAuth: uses the Device Flow token to query
api.githubcopilot.com/models
/auth login, the model cache is automatically refreshed to reflect the new provider.
Anthropic API versioning
Claude models may use differentanthropic-version header values in API requests. The catalog manages this automatically:
- Newer models (claude-fable-5-1, claude-fable-5, claude-opus-5-5, claude-sonnet-5-5, claude-haiku-5-5, claude-opus-5, claude-sonnet-5, claude-opus-4-8, claude-opus-4-7, claude-sonnet-4-6) use the latest API version
- Every entry currently in the catalog ships with the default version; the retired Claude 3.x models that used older headers are no longer registered
- ChatCLI sends the correct header for each model without any user intervention