> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Provider Fallback

> Configure automatic failover between LLM providers with intelligent error classification, exponential cooldown, and health monitoring.

ChatCLI supports an **automatic failover chain** between LLM providers. When the primary provider fails (rate limit, timeout, server error), the system automatically tries the next provider in the chain, completely transparently.

***

## How It Works

The fallback chain is an ordered list of providers. Each request traverses the list until it succeeds or all options are exhausted:

```text theme={"system"}
Request -> OpenAI (primary)
             | failed (rate limit)
           Claude (secondary)
             | failed (timeout)
           Google AI (tertiary)
             | success
           Response returned to user
```

***

## Configuration

<Tabs>
  <Tab title="Environment Variables">
    ```bash theme={"system"}
    # Ordered list of providers (first = highest priority)
    export CHATCLI_FALLBACK_PROVIDERS="OPENAI,CLAUDEAI,GOOGLEAI,ZAI,MINIMAX,MOONSHOT,OPENROUTER,COPILOT"

    # Specific model per provider (optional)
    export CHATCLI_FALLBACK_MODEL_OPENAI="gpt-6.1-sol"
    export CHATCLI_FALLBACK_MODEL_CLAUDEAI="claude-sonnet-5-5"
    export CHATCLI_FALLBACK_MODEL_GOOGLEAI="gemini-3.8-flash"
    export CHATCLI_FALLBACK_MODEL_ZAI="glm-5"
    export CHATCLI_FALLBACK_MODEL_MINIMAX="MiniMax-M2.7"
    export CHATCLI_FALLBACK_MODEL_OPENROUTER="anthropic/claude-sonnet-5"
    export CHATCLI_FALLBACK_MODEL_COPILOT="gpt-4o"

    # MiniMax in Anthropic-compatible mode (optional)
    export MINIMAX_API_COMPAT="anthropic"

    # Retry and cooldown control
    export CHATCLI_FALLBACK_MAX_RETRIES="2"       # attempts per provider
    export CHATCLI_FALLBACK_COOLDOWN_BASE="30s"    # base cooldown
    export CHATCLI_FALLBACK_COOLDOWN_MAX="5m"      # maximum cooldown
    ```
  </Tab>

  <Tab title="Server Flags">
    ```bash theme={"system"}
    chatcli server \
      --fallback-providers OPENAI,CLAUDEAI,GOOGLEAI,ZAI,MINIMAX,MOONSHOT,OPENROUTER,COPILOT \
      --fallback-max-retries 2 \
      --fallback-cooldown-base 30s \
      --fallback-cooldown-max 5m
    ```
  </Tab>

  <Tab title="Helm Chart">
    ```yaml theme={"system"}
    # values.yaml
    fallback:
      enabled: true
      providers:
        - name: OPENAI
          model: gpt-6.1-sol
        - name: CLAUDEAI
          model: claude-sonnet-5-5
        - name: GOOGLEAI
          model: gemini-3.8-flash
        - name: ZAI
          model: glm-5
        - name: MINIMAX
          model: MiniMax-M2.7
        - name: OPENROUTER
          model: anthropic/claude-sonnet-5
        - name: COPILOT
          model: gpt-4o
      maxRetries: 2
      cooldownBase: "30s"
      cooldownMax: "5m"
    ```
  </Tab>
</Tabs>

A non-empty `CHATCLI_FALLBACK_PROVIDERS` (or `--fallback-providers`) is the only switch; the server reads no separate enable variable, and neither the Helm chart nor the operator writes one. A provider without `CHATCLI_FALLBACK_MODEL_<PROVIDER>` runs its own default model (its model variable, such as `OPENAI_MODEL`, then the built-in default); only the primary provider keeps the server model. `maxRetries: 0` (or `CHATCLI_FALLBACK_MAX_RETRIES=0`) fails over to the next provider without retrying.

***

## Error Classification

The system automatically classifies each failure to decide the strategy:

| Class | Behavior | Examples |
| :- | :- | :- |
| `rate_limit` | Waits for backoff, then retries | HTTP 429, "too many requests" |
| `timeout` | Retries up to maxRetries | Deadline exceeded, connection timeout |
| `server_error` | Retries up to maxRetries | HTTP 500, 502, 503 |
| `auth_error` | **Attempts OAuth token refresh** and retries once; if it fails, advances in the chain | HTTP 401, 403, "invalid api key" |
| `model_not_found` | **Does not retry** — advances in the chain | HTTP 404, "model not found" |
| `context_too_long` | **Does not retry** — advances in the chain | "context length exceeded" |

***

## Exponential Cooldown

After consecutive failures, the provider enters cooldown with exponential backoff:

| Consecutive Failures | Cooldown |
| :-: | :-: |
| 1 | 30s |
| 2 | 60s |
| 3 | 120s |
| 4 | 240s |
| 5+ | 300s (max) |

<Note>In interactive CLI mode, authentication errors (401) automatically trigger an OAuth token refresh and retry the request. In server mode (fallback chain), authentication errors receive immediate maximum cooldown (5m). A successful request clears all cooldown for the provider. Use `ResetCooldowns()` to clear manually (e.g., after updating credentials). On the server the chain serves every request that names no provider or model and forwards no credential (`SendPrompt`, `StreamPrompt`, `InteractiveSession`, `AnalyzeIssue`, `AgenticStep`); a streamed request fails over only until the first chunk, never after text was shown, and the reply names the provider and model that answered.</Note>

***

## Health Monitoring

The chain tracks the state of each provider in real time:

```go theme={"system"}
health := chain.GetHealth()
for _, h := range health {
    fmt.Printf("Provider: %s, Available: %v, Fails: %d, Cooldown: %v\n",
        h.Name, h.Available, h.ConsecutiveFails, h.CooldownUntil)
}
```

Fields tracked per provider:

| Field | Description |
| :- | :- |
| `Available` | Whether the provider is available for requests |
| `ConsecutiveFails` | Number of consecutive failures |
| `LastErrorClass` | Type of the last failure |
| `CooldownUntil` | When the cooldown expires |
| `LastErrorAt` | Timestamp of the last failure |

***

## Tool Use with Fallback

The fallback chain also supports `SendPromptWithTools` for providers that implement the `ToolAwareClient` interface. Providers without native tool use support are automatically skipped in the tool call chain.

***

## Best Practices

<CardGroup cols={2}>
  <Card title="Order by cost-effectiveness" icon="ranking-star">
    Place the cheapest/fastest provider first in the chain.
  </Card>

  <Card title="Diversify providers" icon="shuffle">
    Mix providers from different companies for real resilience.
  </Card>

  <Card title="Configure models per provider" icon="sliders">
    Use models with equivalent capabilities to maintain quality.
  </Card>

  <Card title="Monitor health" icon="heart-pulse">
    Regularly check if any provider is in persistent cooldown.
  </Card>
</CardGroup>

<Tip>
  **OpenRouter native fallback:** In addition to ChatCLI's provider fallback chain, OpenRouter itself supports **server-side fallback routing** via `OPENROUTER_FALLBACK_MODELS`. When set (e.g., `anthropic/claude-sonnet-5,google/gemini-3.8-flash`), OpenRouter automatically routes to the next model if the primary is unavailable — before ChatCLI's own chain advances. These two mechanisms are complementary: OpenRouter handles model-level failover within its gateway, while ChatCLI handles provider-level failover across different APIs.
</Tip>

<Warning>Each provider in the chain needs its own configured API key. Make sure to configure the keys for all providers listed in `CHATCLI_FALLBACK_PROVIDERS`.</Warning>

<Note>
  Chain entries are identified by **provider name** — each provider can appear only once, and all entries of the same name share one key, endpoint and health state. To add a third-party OpenAI-compatible gateway as an entry **separate** from OpenAI, use the OpenRouter preset pointed at it: set `OPENROUTER_API_KEY` to the gateway key and `OPENROUTER_API_URL` to the gateway's full chat completions URL, then list `OPENROUTER` alongside `OPENAI` in the chain.
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.