@model tool gives the agent/coder loop the power to route itself: inspect every configured provider’s models — enriched with a price-derived tier, cost per 1M tokens, context window and capabilities — and decide which model should serve the task at hand. It can switch the rest of the task to another model (the provider switches together with it), or delegate a self-contained subtask to a cheaper model without touching the main loop.
This is the in-loop counterpart of the per-call routing the MCP Server surface exposes to external clients (
provider/model params + list_providers): the same power, now available to the AI itself while it works. Sub-tasks that don’t need a frontier model stop paying frontier prices.How it works
model: frontmatter hints — with three guarantees:
- Task-scoped. The override is cleared at the start of every agent run. The AI’s routing decision never silently outlives the task it was made for.
- Non-invasive. It never mutates the session’s own provider/model/client — outside the task,
/modeland/switchremain authoritative. - Accounted. Cost tracking attributes every turn to the model that actually served it —
/costshows exactly what each routed model consumed.
@model use decision wins over a skill’s model: frontmatter hint — an explicit in-task decision outranks a static preference.
Subcommands
What list returns
Pricing tiers
So the model reasons over a label instead of raw prices,list derives a tier from the same pricing tables cost tracking uses:
The tool’s own description teaches the agent the routing policy: delegate mechanical subtasks to a
fast-cheap model; use use only at a clear phase change, never per call — model switches invalidate the provider’s prompt cache, and frequent flip-flopping costs more than it saves.
Qualified handles: PROVIDER:model
With 14 supported providers, the same model id can exist in several of them (claude-* on CLAUDEAI, Bedrock and OpenRouter; deepseek on Ollama and OpenRouter). The canonical, deterministic form is the qualified handle exactly as list prints it:
model: hints — active provider first, then catalog, then a family heuristic (sonnet→CLAUDEAI, gpt-*→OPENAI, glm-*→ZAI, …). The tool result always names the provider that was chosen, so the agent never operates blind.
Ollama tags are safe. The qualified form splits only when the prefix names a real provider —
qwen2.5:14b is not split (its prefix isn’t a provider), and OLLAMA:qwen2.5:14b keeps the tag colon inside the model part.delegate: the biggest token saver
use moves the whole loop — history included — to another model. delegate does something cheaper: it runs one prompt on the target model, with no session history attached, and returns the answer to the main loop as a tool result.
- The main loop’s provider prompt cache stays intact (nothing about its history changes).
- The fat agent history never travels to the cheap model — the delegated call pays only for the prompt you hand it.
- The delegated usage is recorded in cost tracking under the delegated model.
delegate to fast-cheap. Sustained phase change (e.g. a long mechanical migration after the design is settled) → use.
Safety & governance
- Kill switch:
CHATCLI_AGENT_MODEL_TOOL=falseunregisters the tool entirely — the AI cannot route models by itself. Surfaced in/configunder agent → token efficiency. - Capability guard: when the catalog knows the target model lacks native tool support,
usewarns that the loop will fall back to the text protocol for tool calls. - Permission model:
list/statusare read-only;use/resetmutate the loop’s routing anddelegatespends tokens — they are declared as such to the permission system.use/resetare also serialized (never run inside a parallel tool batch). - MoA isolation: Mixture-of-Agents participants run on a strict read-only whitelist —
@modelis unreachable from panel turns.
When the provider refuses a turn
A reply the provider’s safety classifier stops (Anthropicstop_reason: refusal, OpenAI finish_reason: content_filter) arrives with no text. It is not a network fault: retrying the same conversation on the same model mostly gets the same answer, and in the field Fable 5.1 refused half of a run’s turns. It used to surface as “no response received” and end the whole coder session.
Now the error names the cause, and the coder loop recovers on its own:
- It appends a note the model reads (the reply was stopped, nothing was executed, continue from the current step) to the turn and to the run’s history.
- It prints what the classifier flagged — the
stop_detailscategory (cyber,bio,frontier_llm,reasoning_extraction,general_harms) and the explanation when Anthropic gives one — and resends that turn on the model the provider recommends for that category when the refusal names one (stop_details.recommended_model); otherwise on a sibling model of the same provider, as a one-turn route override, walking a ladder: Fable/Mythos → Opus 5.5 → Opus 5 → Sonnet 5.5 → Sonnet 5; Opus → Sonnet 5.5 → Sonnet 5; Sonnet → Haiku 5.5 → Haiku 4.5; Haiku 5.5 → Haiku 4.5. Only models that provider really serves are candidates (exact id or alias match, dots and dashes count as equal). The next turn is back on your model. Areasoning_extractioncategory means the prompt asks for the internal reasoning in the text — remove the instruction rather than switch models. - A run refused more than twice keeps the fallback for the rest of the run instead of paying for a refused request every turn;
@model resetor a new run returns to your model.
via refusal fallback, and cost is attributed to the model that actually served the turn.
Configuration
Usage example
See also
- Cost Tracking — per-model attribution of every routed turn
- Token Efficiency — the other levers the agent uses to spend less
- MCP Server — the same routing power for external MCP clients
- Supported Models — the catalog behind tiers and capabilities
- Environment Variables → Model routing