> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Real-Time Streaming

> Real-time responses with character-by-character streaming and idle timeout watchdog

ChatCLI supports **real-time streaming** of LLM responses, displaying text character-by-character as it is generated by the API. This significantly improves the user experience by eliminating the wait for complete responses.

***

## StreamingClient Interface

Streaming is implemented as an optional interface that providers can adopt:

```go theme={"system"}
type StreamingClient interface {
    LLMClient

    SendPromptStream(ctx context.Context, prompt string,
        history []models.Message, maxTokens int) (<-chan StreamChunk, error)

    SupportsStreaming() bool
}
```

Detection is automatic via type assertion -- providers that implement `StreamingClient` receive streaming automatically:

```go theme={"system"}
if sc, ok := client.AsStreamingClient(c); ok {
    chunks, err := sc.SendPromptStream(ctx, prompt, history, maxTokens)
    // process chunks in real time
}
```

<Info>
  Providers that do not implement `StreamingClient` continue to work normally. ChatCLI falls back to `SendPrompt` (complete response) automatically.
</Info>

***

## Over the gRPC server

`chatcli server` streams the same way: `StreamPrompt` forwards each fragment as the provider emits it (no post-hoc splitting of a buffered reply) and the final message carries the usage and stop reason. The remote client behind `chatcli connect` implements `StreamingClient`, so a connected CLI streams like a local one and its `/cost` prices the reported tokens. See [server mode](/server/server-mode#streaming).

## StreamChunk

Each streaming chunk carries:

| Field | Type | Description |
| :- | :- | :- |
| `Text` | string | Incremental text in this chunk (may be empty) |
| `Done` | bool | `true` on the final chunk |
| `Usage` | \*UsageInfo | Token usage data (only on the final chunk) |
| `StopReason` | string | Stop reason: `end_turn`, `max_tokens`, `tool_use` |
| `Error` | error | Error during streaming (terminates the stream) |

### Streaming Contract

* The channel returns zero or more text chunks
* The final chunk has `Done=true` and may include `Usage` and `StopReason`
* If an error occurs, a chunk with `Error` is sent and the channel closes
* The channel closes after the final chunk or error
* The caller can cancel via context

***

## Supported Providers

| Provider | Streaming | Notes |
| :- | :- | :- |
| **Anthropic (API Key)** | Yes | Native streaming via Messages API |
| **Anthropic (OAuth)** | Yes | Streaming via OAuth token |
| **OpenAI** | Yes | Streaming via Chat Completions |
| **ZAI (Zhipu AI)** | Yes | OpenAI-compatible streaming |
| **MiniMax** | Yes | OpenAI-compatible streaming |
| **Moonshot (Kimi)** | Yes | OpenAI-compatible streaming |
| **OpenRouter** | Yes | Streaming via OpenAI-compatible API |
| Google (Gemini) | No | Fallback to complete response |
| xAI (Grok) | No | Fallback to complete response |
| Ollama | No | Fallback to complete response |

***

## Stream Watchdog

The **Stream Watchdog** monitors the stream to detect stalls (interruptions without data) and prevent ChatCLI from hanging indefinitely:

<Tabs>
  <Tab title="Timeouts">
    | Timer | Duration | Action |
    | :- | :- | :- |
    | **Warning** | 45 seconds | Logs a stall warning |
    | **Idle Timeout** | 90 seconds | Aborts stream and returns partial content |

    Both timers are reset on each received chunk. If the provider stops sending data for 90 seconds, the watchdog interrupts the stream and returns the accumulated text.
  </Tab>

  <Tab title="Result">
    The watchdog returns a `WatchdogResult`:

    | Field | Description |
    | :- | :- |
    | `Text` | Text accumulated so far |
    | `Usage` | Usage data (if stream completed) |
    | `StopReason` | Stop reason |
    | `WasStalled` | `true` if the watchdog triggered due to timeout |
    | `StallCount` | Number of stalls detected during the stream |
  </Tab>
</Tabs>

### Watchdog Configuration

| Environment Variable | Description | Default |
| :- | :- | :- |
| `CHATCLI_STREAM_IDLE_TIMEOUT_SECONDS` | Idle timeout in seconds | 90 |

<Tip>
  On slow networks or with providers that have high latency between chunks, increase the timeout to avoid premature interruptions. The default of 90 seconds is sufficient for most scenarios.
</Tip>

***

## Fallback to Non-Streaming

When streaming is not available (provider does not support it or connection error), ChatCLI falls back automatically:

```text theme={"system"}
1. Tries SendPromptStream()  -> real-time streaming
2. If not supported -> fallback to SendPrompt()
3. Complete response displayed at once
```

The `DrainStream` function allows converting a stream into a complete response when needed:

```go theme={"system"}
text, usage, stopReason, err := client.DrainStream(chunks)
```

***

## TUI Integration

In interactive mode (Bubble Tea), streaming integrates directly with the renderer:

* Each chunk is emitted as an event via `TUIEmitter`
* The Bubble Tea model updates the view incrementally
* Markdown is rendered progressively via Glamour
* The status bar shows the streaming state in real time

<Note>
  In one-shot mode (`-p`), streaming is disabled and `DrainStream` is used to collect the complete response before printing.
</Note>

***

## Next Steps

<CardGroup cols={2}>
  <Card title="Context Recovery" icon="life-ring" href="/context/context-recovery">
    What happens when max\_tokens is reached during streaming.
  </Card>

  <Card title="Provider Fallback" icon="arrow-right-arrow-left" href="/providers/provider-fallback">
    Fallback chain between providers with and without streaming.
  </Card>

  <Card title="Native Tool Use" icon="wrench" href="/agents/native-tool-use">
    Streaming with native tool calls.
  </Card>

  <Card title="Progress UI" icon="spinner" href="/coder/agent-progress-ui">
    Visual indicators during agent streaming.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.