> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Context Recovery

> Automatic 3-level recovery against context overflow, corporate-proxy payload limits and max output token escalation

ChatCLI implements an **automatic context recovery** system that handles three common failure types in long sessions: model context window overflow ("prompt too long"), corporate proxy/gateway payload limits (413 / WAF 403 / silent EOF), and output token limits. When the API rejects a request for any of these reasons, the system applies progressively more aggressive strategies to recover the session without losing the conversation.

***

## Context Overflow Recovery

When the API returns a "context too long" error, ChatCLI applies up to 3 recovery levels before giving up. Every loop recovers, not only the agent/coder one: the chat REPL (the prefix and the turn-context message are rebuilt over the recovered history), RPC chat behind MCP/ACP/the gateway, one-shot (`-p`, notice on stderr so stdout stays pipeable) and each MoA participant (its own thread is compacted; the panel no longer fails on one overflow). The shared helper is bounded (`CHATCLI_MAX_RECOVERY_ATTEMPTS`, default 3) and gives the same guarantees as a planned compaction: memory flushed first, hooks told with the `recovery` trigger, dropped messages archived to CCR, the cache rebuild accounted for.

The classifier that decides "this is an overflow" is shared by every loop and by the provider fallback chain, and covers each provider's phrasing — OpenAI chat and Responses ("context window", `context_length_exceeded`), Anthropic ("prompt is too long"), Gemini ("input token count … exceeds"), xAI ("maximum prompt length"), Mistral, Groq, Bedrock — plus a status-aware check on a 400/413 whose body pairs a token/length word with an excess word.

<Tabs>
  <Tab title="Level 1: Aggressive Budget">
    **First attempt**: halves the budget limits and cleans up misalignments.

    Actions:

    * Repairs tool result pairing (removes orphans, injects synthetics)
    * Applies 50% of `DefaultTurnBudgetChars` and `DefaultPerResultMaxChars` **for this session only** — the process-wide defaults are never mutated, so other sessions of a gateway or `rpcserve` process never observe halved limits
    * Applies budget enforcement with reduced limits
    * Truncates long assistant messages to 5,000 chars

    <Info>
      The original limits are restored after application. Only the current history is affected by the reduction.
    </Info>
  </Tab>

  <Tab title="Level 2: Emergency Truncation">
    **Second attempt**: keeps only system messages and the last N messages.

    Actions:

    * Preserves all system messages (system prompt, bootstrap, contexts)
    * Keeps the last 10 non-system messages (configurable)
    * Ensures the history starts with a `user` message (API requirement)
    * Validates tool result pairing in the truncated history
  </Tab>

  <Tab title="Level 3: Nuclear Truncation">
    **Third attempt**: keeps only the minimum needed to continue.

    Actions:

    * Preserves system messages
    * Keeps only the last 4 messages (2 user/assistant exchanges)
    * Injects a notice message explaining that the context was compacted

    ```text theme={"system"}
    [Context was automatically compacted due to size limits.
    Previous conversation history has been summarized.
    Continue from where you left off.]
    ```
  </Tab>
</Tabs>

### Error Detection — model overflow

The system recognizes multiple forms of overflow errors:

| Error Message | Provider |
| :- | :- |
| `context length exceeded` | Anthropic |
| `prompt is too long` | OpenAI |
| `request too large` | Various |
| `max_tokens exceed` | Various |
| `input too long` | Google |
| `token limit` | Generic |

***

## Corporate proxy / gateway recovery

Enterprise environments often sit behind a proxy or gateway that enforces **a POST body size cap** — typically 1-5 MB, completely independent of the model's context window. You can be well within Anthropic's 200K-token window (\~800 KB) and still take a mysterious rejection from the proxy. Worse: many proxies don't return a clean 413 — some send a WAF 403 (Cloudflare, Akamai, mod\_security), 431 (header too large), or simply drop the TCP connection mid-POST, surfacing as `EOF` / `connection reset` on the client.

ChatCLI detects all three patterns and funnels them through the same recovery flow as context overflow.

### Error Detection — proxy/gateway

| Pattern detected | Example | Function |
| :- | :- | :- |
| HTTP 413 + variants | `413 Payload Too Large`, `request entity too large`, `body too large`, `maximum request size`, `431 Request Header Fields Too Large` | `IsPayloadTooLargeError` |
| 403 with WAF/firewall signals | 403 with `firewall`, `waf`, `security policy`, `blocked by`, `cloudflare`, `cf-ray`, `mod_security`, `akamai`, `proxy denied`, `policy violation` | `IsProxyWAFRejection` |
| 403 with HTML body (Bedrock/AWS SDK) | `StatusCode: 403 ... deserialization failed ... invalid character '<' looking for beginning of value` (SDK trying to decode an HTML proxy block page as JSON) | `IsProxyWAFRejection` |
| EOF / reset with large history | `unexpected eof`, `connection reset`, `broken pipe`, `stream error` **and** history > 500 KB | `IsLikelyPayloadProblem` (heuristic) |

<Info>
  WAF detection is **conservative** — a 403 without firewall signals continues to be treated as an auth error (OAuth refresh + retry). Only when a 403 carries specific proxy/WAF signals is it reclassified as a recoverable payload failure. This prevents invalidating valid OAuth credentials when the real problem is on the network layer.
</Info>

<Info>
  **The corporate Bedrock case:** when the proxy/WAF intercepts the POST to Bedrock Runtime and returns an HTML block page with status 403, the AWS SDK tries to parse the body as JSON and fails with `"invalid character '<' looking for beginning of value"` and an empty `RequestID`. That pattern is an unambiguous middlebox fingerprint (a real AWS 403 returns well-formed JSON) — ChatCLI reclassifies it as a recoverable payload failure and triggers the same recovery ladder.
</Info>

<Info>
  EOF / connection-reset detection applies a **history-size threshold (500 KB)** before suspecting payload. Small requests that hit EOF keep being treated as transient network failures (normal retry). Only when the history is already suspiciously large is EOF reclassified as a probable body cap.
</Info>

### Pre-flight check

Every agent turn, history is measured before the request goes out. Two paths:

**With `CHATCLI_MAX_PAYLOAD` set:**
If history crosses 85% of the cap, `BudgetRatio` is forced to `0.40` up front — aggressive preventive compaction. The user sees:

```text theme={"system"}
ℹ pre-flight: history 4.2 MB ≈ 86% of configured cap (5.0 MB) — compacting
```

**Without a cap set:**
If history crosses 2.5 MB, a **one-shot warning per session** fires suggesting the env var. It will not re-trigger in the same run to avoid noise.

```text theme={"system"}
ℹ history 2.8 MB — if you are behind a proxy/gateway, export CHATCLI_MAX_PAYLOAD=5MB (adjust to the proxy limit)
```

### Adaptive learned cap

Bedrock providers annotate every transport error with the **exact size of the request the middlebox rejected**. When a payload rejection fires, ChatCLI derives the session cap directly from that observation — **¾ of the rejected size** — instead of guessing:

```text theme={"system"}
⚠ Recoverable failure (proxy/WAF rejection (403 + security signals)) — compacting and retrying
ℹ Gateway rejected a 126.4 KB request — payload cap set to 94.8 KB for this session (export CHATCLI_MAX_PAYLOAD to override)
```

This matters enormously in practice: real-world corporate gateways often cap bodies around **128 KB**, where a guessed multi-megabyte cap would change nothing. The learned cap only ever **tightens** — if `CHATCLI_MAX_PAYLOAD` is already set below it, your value wins; a later, smaller rejection tightens it further.

When the rejected size is *unknown* (providers without size annotation), the previous behavior remains as fallback: assume **4 MB** for the rest of the session.

```text theme={"system"}
ℹ Assuming 4 MB payload cap — export CHATCLI_MAX_PAYLOAD (e.g. 5MB, 512KB) to adjust
```

### Floor diagnosis — when compaction cannot help

The system prompt (agent charter, personas, skills, MCP tool docs) is **never compacted**. When it alone reaches the size the gateway just rejected, no amount of history compaction can produce an acceptable request — retrying identical payloads would only burn recovery attempts. ChatCLI detects this and **fails fast with an actionable message**:

```text theme={"system"}
✖ The non-compactable request base (system prompt: 131.2 KB — personas, skills, MCP tool docs)
  meets or exceeds the size the gateway rejected (126.4 KB). History compaction cannot help —
  reduce MCP servers, skills or personas, or raise the proxy/WAF body-size limit
```

At **¾ of the rejected size** a warning fires instead — the session continues, but you're near the edge:

```text theme={"system"}
⚠ System prompt alone is 96.1 KB — close to the gateway's rejection threshold;
  consider reducing MCP servers, skills or personas
```

<Tip>
  Three structural features shrink this floor: MCP tool descriptions are clamped to one line in the system-prompt index (full schema via `@tools describe`), injected skill bodies respect [`CHATCLI_SKILL_INJECT_BUDGET`](/extensions/skill-registry#skill-injection-budget) plus its 2× per-run cap, and mid-loop skill blocks age out to stubs after `CHATCLI_SKILL_AGE_TURNS` turns (recoverable via `@recall`). See [MCP Integration](/extensions/mcp-integration) and [Skill Registry](/extensions/skill-registry).
</Tip>

### Content shrinking — Level 3 that actually works

Whole-message dropping is a no-op when the history is short — agent sessions often hold just a system prompt plus a handful of huge tool results, and `MinKeepRecent` keeps exactly the messages that carry the bulk. Emergency truncation therefore has a final pass: it **shrinks the content of non-system messages, largest first**, until the payload budget is met. System messages are never touched.

When the [compression layer](/context/context-compression) is active, each message is **archived verbatim in the CCR store before its first cut** and the shrunk content carries a `<<ccr:KEY>>` marker — the model can recover the original at any time with `@recall`. Nuclear truncation (level 3 of the recovery ladder) additionally hard-caps every kept non-system message at 4,000 chars.

### What every level guarantees

The same guarantees hold whichever loop compacts (chat, agent/coder, one-shot, `/compact`) and in overflow recovery:

* **Nothing the model asked to keep is touched.** Messages flagged `PreserveVerbatim` (an `@recall` result the model explicitly requested in full) are excluded from the summarized segment and re-appended right after the summary at Level 2, kept whole at Level 3, and skipped by content shrinking. A verbatim native tool result re-appended without its owning tool call becomes a user-role message, so the tool-result pairing repair never deletes it. Verbatim messages may hold at most a quarter of the budget: past that, the oldest are archived to CCR again, stubbed with an `@recall` marker and lose the flag, so compaction always converges.
* **A compaction that changes nothing is not a compaction.** A no-op, a rejected summary or a failure pops the undo snapshot (so `/rewind compact` never "restores" an identical history) and fires its paired `PostCompact` hook with `outcome=skipped`; restoring a `/rewind` checkpoint uses the same bookkeeping as `/rewind compact` (journal rewrite, expected cache rebuild, undo stack cleared).
* **Level 3 is no longer lossy.** The messages it drops are archived to the CCR store first and the truncation notice carries an `@recall` marker for the whole segment — the same recoverability Levels 1 and 2 already had. Overflow recovery flushes pending memory, archives the messages it removes and appends the marker to the recovered history.
* **A bad summary never replaces the segment.** A quality gate rejects refusals and answers too short for the segment (floor scales with its size, capped at 80 characters); the summarizer is retried once and, failing that, the pipeline falls to Level 3 with its archive.
* **The summarizer is not paid for nothing.** When a summary does not bring the history under budget, the next two compactions skip Level 2 and go straight to Level 3 instead of paying a summarizer call per turn.
* **Hooks see every compaction.** `PreCompact`/`PostCompact` fire with a `trigger` of `auto`, `manual` or `recovery` ([hooks](/extensions/hooks-system#available-events)).
* **It is accounted.** `/cost` shows the number of compactions, how many landed at Level 3 and what the summarizer cost; the counters persist with the session ([cost tracking](/providers/cost-tracking)).

Guided `/compact <instruction>` follows the same rules: it uses the configured summarizer route (`CHATCLI_COMPACT_MODEL`) when there is one, gives the summary its own 10-minute allowance instead of the old 60-second command deadline, keeps `PreserveVerbatim` messages, archives the segment and runs the quality gate.

### System notice injected in history

After a payload-limit-triggered recovery, ChatCLI injects a `user` message before the retry instructing the model to prefer smaller reads going forward. This breaks the model's loop of trying to re-read the same huge file that caused the 413 in the first place. The notice is injected **at most once per session** — recovery detects an existing copy anywhere in the history (even folded into a compaction summary) and never stacks a second one, since every extra byte works against the very limit being recovered from:

```text theme={"system"}
[SYSTEM NOTICE — PAYLOAD LIMIT HIT] A proxy/gateway rejected the previous
request due to body size. History was compacted to recover. Going forward:
(1) When reading files, prefer targeted reads with line ranges
    (e.g. sed -n '100,200p' file, or read_file with offset+limit) instead
    of reading entire files.
(2) Prefer grep/ripgrep with specific patterns over full-file reads.
(3) If you previously read a large file, its full content is persisted at
    the path shown in the tool-result preview — re-read specific ranges
    from that file rather than repeating the original read.
(4) Summarize findings incrementally rather than accumulating raw tool output.
```

<Tip>
  This hint is intentionally injected **in English**. The AI follows English instructions much more faithfully even when the user is on pt-BR, and this message is never shown to the user — it only enters the history sent to the model.
</Tip>

***

## Max Output Token Escalation

When the model stops generating because it hit the `max_tokens` limit, ChatCLI can automatically escalate:

| Attempt | Action |
| :- | :- |
| 1st | Doubles the current `max_tokens` (up to the provider's cap) |
| 2nd | Doubles again (up to the provider's cap) |
| 3rd+ | Stops escalating, returns partial content |

### Continuation Message

When the model is interrupted by a token limit, ChatCLI injects a continuation message:

```text theme={"system"}
Your response was cut off at the token limit.
Resume DIRECTLY from where you stopped -- do not repeat any content.
Continue the implementation or explanation from the exact point of interruption.
```

<Tip>
  The message instructs the model to continue from where it left off, avoiding repetition of already generated content.
</Tip>

***

## Configuration

| Environment Variable | Description | Default |
| :- | :- | :- |
| `CHATCLI_CONTEXT_WINDOW` | **Global** context-window override (in tokens), for any provider/model. Takes precedence over the catalog. Use it when your gateway/agent's real window differs from what ChatCLI assumes — the compaction budget derives from this value. | *(auto from catalog)* |
| `CHATCLI_MAX_RECOVERY_ATTEMPTS` | Maximum context recovery attempts | `3` |
| `CHATCLI_MAX_TOKEN_ESCALATIONS` | Maximum max\_tokens escalations | `2` |
| `CHATCLI_EMERGENCY_KEEP_MESSAGES` | Messages kept in emergency truncation | `10` |
| `CHATCLI_MAX_PAYLOAD` | **Human-friendly** ceiling for POST body size (e.g. `5MB`, `512KB`, `2.5MB`, `5`=5MB). When set, the compactor respects this ceiling as an extra cap on top of the model's context window, and pre-flight forces compaction on crossing 85% of it. | *(unset — no cap)* |

### Live feedback during compaction

Since this release, the terminal never "freezes" during a long compaction anymore. `HistoryCompactor` emits status at each pipeline phase via `SetStatusCallback`:

```text theme={"system"}
│ 📦 Compacting history (23 msgs, 4.2 MB → target 2.9 MB)
│ 🧹 Trim: stripping reasoning/dedup (no LLM)…
│ 🧠 Summarizing old messages via LLM (may take 30-90s — ESC cancels)…
│ ✓ Summary applied (23 → 9 msgs, 4.2 MB → 1.8 MB)
```

**Cancellation**: the summarization LLM call now derives its context from the turn — `Ctrl+C` / `ESC` propagates correctly and aborts compaction without corrupting history (returns `ctx.Err()` instead of blindly falling through to emergency truncation).

### Microcompact (pre-budget)

The budget check itself is **payload-honest**: it weighs message text plus native tool-call arguments, image payloads and system-part excess, so tool- or vision-heavy histories cross the threshold when the real request does — not after a proxy already rejected it. Before `NeedsCompaction` checks whether history exceeds budget, the agent loop applies `ApplyMicrocompact` — a **pure-Go, no-LLM, no-network** pass that progressively truncates/summarizes old tool results (2+ turns old → head+tail preview; 4+ turns old → one-line summary). In most cases this keeps history inside budget without triggering the (expensive) Level 2.

The **summarizer's input is budgeted** against the summarizer model's own window (half of it, floor 20K chars) instead of a fixed head/tail cut per message: every message of the segment gets a fair allowance, a message that had already been folded into a `<<ccr:KEY>>` stub is **restored from the CCR archive** when the original fits, over-long messages keep head and tail in a 3:1 proportion, and native tool calls are named so the summary can list the commands executed. The same rendering serves `/compact <instruction>`.

With the [compression layer](/context/context-compression) active, microcompact is **lossless**: the original tool result is archived in the CCR store before the cut and the preview/summary stub embeds a `<<ccr:KEY>>` marker. Markers carry forward across levels — when a truncated preview later degrades to a one-line summary, the marker survives on the summary — so the model can always expand the original with `@recall`.

### Memory flush before compaction

The memory worker distills facts from the live history a couple of turns behind the conversation; compaction replaces the middle of that history with a summary, so anything the worker had not reached yet used to be distilled from the summary at best. Every compaction site (auto-compact in chat, agent and coder, explicit and guided `/compact`, one-shot) now hands the not-yet-extracted segment to the worker's durable queue **before** the rewrite, so the original messages reach long-term memory verbatim regardless of what the summary keeps.

### Repeated reads (pre-budget)

At the same turn boundary, `DedupRepeatedReads` removes the copies a coder session accumulates by reading the same file over and over: when a later `@coder read` of a path covers the same or a wider line range, every earlier read of that path is replaced by a one-line stub (archived in the CCR store first, so `@recall` can restore it) and the newest read stays verbatim. Narrower later reads never supersede a wider earlier one, `@recall` output is never touched, and each pass is reported in the turn line. The model keeps exactly one current view per file instead of five.

```text theme={"system"}
🗜 microcompact: 3 truncated, 2 summarized, 1.7 MB freed
```

Configurable via env:

| Environment Variable | Description | Default |
| :- | :- | :- |
| `CHATCLI_MICROCOMPACT_TRUNCATE_TURNS` | Age (in turns) at which tool results start getting truncated | `2` |
| `CHATCLI_MICROCOMPACT_SUMMARIZE_TURNS` | Age (in turns) at which tool results are replaced by a one-line summary | `4` |
| `CHATCLI_COMPACT_MODEL` | Dedicated `PROVIDER:model` for the Level 2 structured summary (cheaper/faster than the session model) | *(session client)* |

### Aggressive Budget Ratio

At level 1, the tool result budget limits are multiplied by `0.5` (50%). This means:

| Parameter | Normal | Level 1 Recovery |
| :- | :- | :- |
| Budget per turn | 200,000 chars | 100,000 chars |
| Max per result | 20,000 chars | 10,000 chars |

***

## Recovery Flow

```text theme={"system"}
API returns recoverable error (context overflow | 413 | WAF 403 | EOF w/ large history)
  │
  ├─ Payload-related: cap learned from the rejected size (¾) ────────┐
  │   ├─ System prompt ≥ rejected size → FAIL FAST with diagnosis    │
  │   └─ System notice injected into history (once per session)      │
  │                                                                   │
  ├─ Attempt 1: Aggressive budget (50%) + pairing cleanup             │
  │   └─ Resend to API                                                │
  │       ├─ Success → continues normally                             │
  │       └─ Failure → next attempt                                   │
  │                                                                   │
  ├─ Attempt 2: Emergency truncate (system + last 10 msgs)            │
  │   └─ Resend to API                                                │
  │       ├─ Success → continues with reduced history                 │
  │       └─ Failure → next attempt                                   │
  │                                                                   │
  └─ Attempt 3: Nuclear truncate (system + last 4 msgs,               │
      each capped at 4,000 chars, CCR-archived first)                 │
      └─ Resend to API                                                │
          ├─ Success → continues with minimal history                 │
          └─ Failure → error reported to user                         │
                                                                      │
  Rejected size unknown: CHATCLI_MAX_PAYLOAD auto-set to 4MB ◄────────┘
                         if not configured
```

<Warning>
  After nuclear truncation (level 3), the model's working context drops to the last 2 exchanges. With the [compression layer](/context/context-compression) active the dropped content stays recoverable — stubs carry `<<ccr:KEY>>` markers the model can expand with `@recall` — but proactive `/compact` remains the better experience: a structured summary beats a pile of recall markers.
</Warning>

***

## Interaction with Other Systems

Context recovery works in conjunction with:

<CardGroup cols={2}>
  <Card title="Tool Result Budget" icon="database" href="/context/tool-result-management">
    The result budget is the first line of defense. Recovery activates when the budget was not sufficient.
  </Card>

  <Card title="Microcompaction" icon="compress" href="/context/tool-result-management">
    Progressive compaction reduces context growth over time.
  </Card>

  <Card title="Conversation Control" icon="clock-rotate-left" href="/usage/conversation-control">
    The `/compact` command is the proactive way to prevent overflow.
  </Card>

  <Card title="Cost Tracking" icon="money-bill-trend-up" href="/providers/cost-tracking">
    Monitor context usage to anticipate when /compact will be needed.
  </Card>
</CardGroup>

## External context engine (MCP)

`CHATCLI_CONTEXT_ENGINE=mcp:<server>` hands the Level 2 summary to a configured MCP server exposing `context_compact(segment, budget_chars, instruction)`: the rendered conversation segment (the same budgeted rendering the embedded summarizer receives), the character budget and, for guided `/compact <instruction>`, the user's instruction. Its answer replaces the segment; the CCR archive, checkpoints and `@recall` keep working exactly as before. An error, a timeout (3 min) or an empty answer falls back to the embedded summarizer for that compaction. Shown in `/config compression`.

`CHATCLI_CONTEXT_ENGINE=provider` adds the model's own server-side context editing on top of the local pipeline. On Anthropic (API key or OAuth) every tool-loop request carries the `context-management-2025-06-27` beta and a `context_management` block with one `clear_tool_uses_20250919` edit: once the prompt passes 100 K input tokens the API clears the oldest tool results itself, keeping the five most recent tool uses and freeing at least 20 K tokens per edit, before those tokens are billed. The applied edits come back in the response and are mirrored locally: the oldest tool results the server cleared are stubbed in the local history (the originals archived to CCR, recoverable with `@recall`), the chars/token calibration skips that turn, the next cache write is booked as an expected rebuild and `/context status` shows how many edits, tool results and tokens the provider cleared. Local compaction, the CCR archive and `/rewind compact` keep working unchanged, so nothing is lost that was not already recoverable. Providers without a documented server-side equivalent ignore the setting.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.