Skip to main content
The @model tool gives the agent/coder loop the power to route itself: inspect every configured provider’s models — enriched with a price-derived tier, cost per 1M tokens, context window and capabilities — and decide which model should serve the task at hand. It can switch the rest of the task to another model (the provider switches together with it), or delegate a self-contained subtask to a cheaper model without touching the main loop.
This is the in-loop counterpart of the per-call routing the MCP Server surface exposes to external clients (provider/model params + list_providers): the same power, now available to the AI itself while it works. Sub-tasks that don’t need a frontier model stop paying frontier prices.

How it works

The routing decision lives in a route override honored per turn by the same mechanism as skill model: frontmatter hints — with three guarantees:
  1. Task-scoped. The override is cleared at the start of every agent run. The AI’s routing decision never silently outlives the task it was made for.
  2. Non-invasive. It never mutates the session’s own provider/model/client — outside the task, /model and /switch remain authoritative.
  3. Accounted. Cost tracking attributes every turn to the model that actually served it — /cost shows exactly what each routed model consumed.
When both are present, the AI’s @model use decision wins over a skill’s model: frontmatter hint — an explicit in-task decision outranks a static preference.

Subcommands

What list returns


Pricing tiers

So the model reasons over a label instead of raw prices, list derives a tier from the same pricing tables cost tracking uses: The tool’s own description teaches the agent the routing policy: delegate mechanical subtasks to a fast-cheap model; use use only at a clear phase change, never per call — model switches invalidate the provider’s prompt cache, and frequent flip-flopping costs more than it saves.

Qualified handles: PROVIDER:model

With 14 supported providers, the same model id can exist in several of them (claude-* on CLAUDEAI, Bedrock and OpenRouter; deepseek on Ollama and OpenRouter). The canonical, deterministic form is the qualified handle exactly as list prints it:
Bare model names are also accepted and resolved by the same pipeline as skill model: hints — active provider first, then catalog, then a family heuristic (sonnet→CLAUDEAI, gpt-*→OPENAI, glm-*→ZAI, …). The tool result always names the provider that was chosen, so the agent never operates blind.
Ollama tags are safe. The qualified form splits only when the prefix names a real provider — qwen2.5:14b is not split (its prefix isn’t a provider), and OLLAMA:qwen2.5:14b keeps the tag colon inside the model part.
Errors are actionable: asking for a provider without credentials returns “wanted X on PROVIDER but that provider is not configured (missing API key)” — the agent can pick another handle or surface the problem instead of retrying blindly.

delegate: the biggest token saver

use moves the whole loop — history included — to another model. delegate does something cheaper: it runs one prompt on the target model, with no session history attached, and returns the answer to the main loop as a tool result.
  • The main loop’s provider prompt cache stays intact (nothing about its history changes).
  • The fat agent history never travels to the cheap model — the delegated call pays only for the prompt you hand it.
  • The delegated usage is recorded in cost tracking under the delegated model.
Rule of thumb the tool teaches the agent: “summarize these files”, “extract this list”, “reformat this output” → delegate to fast-cheap. Sustained phase change (e.g. a long mechanical migration after the design is settled) → use.

Safety & governance

  • Kill switch: CHATCLI_AGENT_MODEL_TOOL=false unregisters the tool entirely — the AI cannot route models by itself. Surfaced in /config under agent → token efficiency.
  • Capability guard: when the catalog knows the target model lacks native tool support, use warns that the loop will fall back to the text protocol for tool calls.
  • Permission model: list/status are read-only; use/reset mutate the loop’s routing and delegate spends tokens — they are declared as such to the permission system. use/reset are also serialized (never run inside a parallel tool batch).
  • MoA isolation: Mixture-of-Agents participants run on a strict read-only whitelist — @model is unreachable from panel turns.

When the provider refuses a turn

A reply the provider’s safety classifier stops (Anthropic stop_reason: refusal, OpenAI finish_reason: content_filter) arrives with no text. It is not a network fault: retrying the same conversation on the same model mostly gets the same answer, and in the field Fable 5.1 refused half of a run’s turns. It used to surface as “no response received” and end the whole coder session. Now the error names the cause, and the coder loop recovers on its own:
  1. It appends a note the model reads (the reply was stopped, nothing was executed, continue from the current step) to the turn and to the run’s history.
  2. It prints what the classifier flagged — the stop_details category (cyber, bio, frontier_llm, reasoning_extraction, general_harms) and the explanation when Anthropic gives one — and resends that turn on the model the provider recommends for that category when the refusal names one (stop_details.recommended_model); otherwise on a sibling model of the same provider, as a one-turn route override, walking a ladder: Fable/Mythos → Opus 5.5 → Opus 5 → Sonnet 5.5 → Sonnet 5; Opus → Sonnet 5.5 → Sonnet 5; Sonnet → Haiku 5.5 → Haiku 4.5; Haiku 5.5 → Haiku 4.5. Only models that provider really serves are candidates (exact id or alias match, dots and dashes count as equal). The next turn is back on your model. A reasoning_extraction category means the prompt asks for the internal reasoning in the text — remove the instruction rather than switch models.
  3. A run refused more than twice keeps the fallback for the rest of the run instead of paying for a refused request every turn; @model reset or a new run returns to your model.
The switch is visible where every route change is: the session card and the feed of the live dashboard show via refusal fallback, and cost is attributed to the model that actually served the turn.

Configuration


Usage example


See also