> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent Squad

> A coordinated team of agents: live observability of every run, a shared kanban work board, agent-to-agent mail, and an orchestrator playbook that drives a single prompt to a validated delivery — autonomously.

The **Agent Squad** layer turns [Multi-Agent Orchestration](/agents/multi-agent-orchestration) into a real team. You give one prompt; the orchestrator plans the work into cards, dispatches specialist workers, watches where each one is, routes review feedback back to the coder, schedules follow-ups, and only delivers after validation — without you babysitting anything.

It is built from four pieces, each usable on its own:

| Piece | Human surface | AI surface |
| - | - | - |
| Live run registry | `/agents` | `@agents` |
| Work board (kanban) | `/board` | `@board` |
| Squad mail | `/mail` | `@mail` + worker `send_mail` |
| Delivery playbook | — | orchestrator system prompt |

***

## When does the squad activate?

There is no "squad switch" in the code — the orchestrator decides **per turn**, guided by a decision ladder taught in its system prompt (always present in `/coder` and `/agent`, since multi-agent mode is on by default). Knowing the ladder avoids the classic confusion of "I asked something and no cards appeared":

| Your request looks like… | What the orchestrator does |
| - | - |
| 1–2 reads or one small edit ("fix this typo") | Direct tool calls. No workers, no cards — the squad would be pure overhead. |
| 3+ independent operations, multi-file change, codebase-wide search | Dispatches parallel workers via `agent_call` — still no board, it is faster than card bureaucracy. |
| A multi-step goal with a deliverable ("implement OAuth with tests, reviewed") | **Full squad**: cards on the board, assigned workers, review loop, validated delivery. The playbook forbids ending with cards outside `done`/`blocked`. |
| Continuous or deferred work ("monitor X", "every day at 9am") | Squad + [scheduler](/tools/scheduler): the job is linked to its card, which waits in `blocked` with a note. |

Two deterministic triggers also exist: high-complexity tasks fire [Plan-and-Solve](/agents/harness/plan-and-solve) before the loop (a structured plan naturally becomes cards), and any `[SQUAD MAIL]` in the inbox is injected at the turn boundary with instructions to react before continuing.

<Tip>You can always force it: `/plan <goal>` plans first, and an explicit instruction in the prompt — "create cards on the board and deliver this reviewed" — beats the heuristic. The middle ground is model judgment by design: the playbook gives a ruler, not a rigid if/else.</Tip>

<Note>For an **approved plan of 5+ genuinely independent tasks**, there is a stricter layer above the playbook: the [Task Graph](/agents/task-graph) — a persisted DAG where the engine runs each task's validation commands itself and an independent reviewer issues the verdict before anything counts as done.</Note>

<Note>A dispatched worker can be granted session plugins for its task — `<agent_call agent="coder" task="..." tools="@browser,@websearch" />` — gated by the same security policy. See [Worker capabilities](/agents/task-graph#worker-capabilities).</Note>

***

## Live run registry — where is each agent?

Every agent execution registers itself in a process-wide registry: the orchestrator loop, each dispatched worker, delegate subagents, Mixture-of-Agents panel members and scheduler headless runs. Runs form a parent → child tree, and each one reports its current ReAct turn and the action in flight.

The live dispatch panel uses it to show real progress per agent:

```text theme={"system"}
⠋ [gpt-5.6-sol] [1m02s] [████████░░░░░░░░░░░░] 2/4 agents (50%)
  ✓ [file] Read engine files ─ concluido (12.4s)
  ⠋ [coder] Implement /foo command ─ turno 7/30 · patch cli/foo.go
      ↳ [subagent] analyze metrics endpoint — turno 3/15 · read
  ⠋ [reviewer] Review the diff ─ turno 2/30 · git-diff
  ○ [tester] Run the test suite ─ pendente
```

### `/agents`

```bash theme={"system"}
/agents              # tree of live runs + recent history
/agents show run-3   # one run in full: turn, action, tool calls, children
/agents cancel run-3 # cancel a stuck run (propagates to its subagents)
```

### `@agents` (for the AI)

The orchestrator uses the same registry through the `@agents` tool (`list`, `show`, `cancel`) — so it can check what is already running before dispatching duplicate work, and kill a worker that is looping without progress.

Cancellation is per-run: it cancels that run's context and everything it spawned, never the whole batch.

### Cross-process runs — see the gateway daemon from your REPL

When the [Conversation Hub](/gateway/conversation-hub) is enabled, each ChatCLI process mirrors its live runs into the shared hub database. `/agents` (and `@agents`) then list runs owned by **other processes** — a squad executing inside the gateway daemon, a scheduler headless run — in a dedicated section, tagged with their origin:

```text theme={"system"}
Runs in other processes (via hub) (2)
  ▶ run-3f9a2b01-2 [worker/coder] · turno 5/30 · patch api.go · 42s — build the customer API [gateway]
  ▶ run-77ab01c4-1 [headless/scheduler] · 12s — nightly dependency scan [scheduler]
```

Every mirrored run carries a liveness heartbeat; if its process dies without finalizing, the entry is flagged (`⚠ no heartbeat`) instead of lying forever as "running". `/agents cancel` works cross-process too: the request is flagged in the hub and the owning process honors it on its next sync tick (about a second).

Run IDs embed a per-process instance token (`run-<inst>-<n>`), so two processes can never mint colliding IDs.

***

## Driving the squad while it runs

The terminal is not locked while a squad (or any `/coder` run) is working: type `/agents`, `/board`, `/mail`, `/jobs`, `/taskgraph` or [`/dash`](/usage/live-dashboard) at any moment and the command executes right away — the live progress panel pauses, the output prints, the panel resumes.

On macOS/Linux you also **see what you type**: the line being composed renders live on its own row **below** the panel or the turn spinner as `❯ your text▌`, with backspace working — ChatCLI owns the line editing while the loop runs, so the spinner no longer eats your keystrokes. (Windows keeps the classic behavior: the line is delivered on Enter.)

```text theme={"system"}
⠋ [claude-sonnet-5] [1m02s] [████████░░░░] 2/4 agents (50%)
/board                     ← typed mid-run
  ⚡ /board
  📋 doing: card-2 Implement API (coder, run-…-3)
  ...panel resumes...
```

Two rules keep this safe:

* These four commands never reach the model or a pending security prompt — they are intercepted before the type-ahead queue, so typing `/board` at the wrong moment can never be read as a prompt answer or become an instruction to the LLM.
* If the terminal is momentarily owned by a security confirmation, the command queues and runs at the next turn boundary (you'll see `⚡ Applying N command(s) queued during the run`).

`/mail send <agent> <text>` mid-run is the steering wheel: the directive lands in the recipient's inbox and is delivered at its next ReAct turn — including agents running in another process, via the hub.

Anything else you type mid-run keeps the existing behavior: it queues as type-ahead and is injected into the conversation as your next instruction at the turn boundary.

***

## Work board — the squad's kanban

The board is the shared unit of work: cards flowing across **backlog → doing → review → blocked → done**. A card carries an assignee (worker agent type), timestamped notes (review verdicts, delivery summaries), linked agent run IDs and scheduler job IDs, and its full transition history.

```bash theme={"system"}
/board                       # kanban grouped by column
/board show card-3           # description, notes, history, linked runs/jobs
/board create Fix login bug  # you can add work too
/board move card-3 review
/board assign card-3 reviewer
/board note card-3 needs tests for the OAuth path
/board archive 24h           # clean old done cards
```

The AI manages the same board via `@board` (`create`, `list`, `show`, `move`, `assign`, `note`, `link`, `archive`). Agent results include their `run_id`, so the orchestrator links every execution to its card for traceability.

Persistence is a single JSON document written atomically under `~/.chatcli/board.json` (override with `CHATCLI_BOARD_PATH`). A corrupt file surfaces an error — it is never silently wiped.

***

## Squad mail — agents talking to each other

Squad mail is a directed message bus between agents. Messages are injected into the recipient's context at its **next turn boundary** — the only point that cannot split a native tool call and its result.

* **Workers** get a universal native `send_mail` tool (regardless of their command allowlist): a reviewer can hand its verdict straight to the coder mid-flight.
* **The orchestrator** drains its own inbox every turn (delivered as `[SQUAD MAIL]` blocks) and uses `@mail` (`send`, `inbox`, `history`).
* **You** can redirect any agent without interrupting it:

```bash theme={"system"}
/mail send coder prioritize the login fix over the refactor
/mail list      # recent traffic (who told what to whom)
/mail pending   # queued messages per recipient
```

### Durable and cross-process

Squad mail is persisted through the [Conversation Hub](/gateway/conversation-hub) SQLite store (WAL mode): messages survive restarts and flow between processes. A directive typed in your REPL reaches agents running inside the [Chat Gateway](/gateway/chat-gateway) daemon, and vice versa. Delivery acks stop other processes (and post-restart hydration) from redelivering consumed messages.

***

## The delivery playbook

With observability, a board and messaging in place, the orchestrator system prompt teaches the full autonomous cycle:

1. **Plan** — break the goal into cards (`@board create`, one per deliverable, assignee = agent type).
2. **Develop** — move the card to `doing`, dispatch the assigned workers, link the `run_id`.
3. **Review** — move to `review`, dispatch reviewer/tester, record the verdict as a card note. Failed review → findings go back to the coder (new dispatch or `@mail send coder`), card returns to `doing`.
4. **Deliver** — validate for real (build + tests, red → green), then move to `done` with a delivery note.
5. **Loop** — repeat until no card sits outside `done`. Continuous or deferred work is scheduled via [`@scheduler`](/tools/scheduler) with the job linked to its card.

The orchestrator is instructed to never end a run with unfinished cards unless they are in `blocked` with a note explaining the blocker.

***

## Structured gateway telemetry

When the squad runs inside the [Chat Gateway](/gateway/chat-gateway) (Telegram, Slack, …), progress is emitted from **typed agent events** instead of scraping the rendered terminal output: reasoning lines, tool start/end with duration, plan counters, and one line per worker state change (current turn and action, then terminal status).

```text theme={"system"}
🧠 Reading the config to understand defaults
▸ Reading: cli/foo.go
✓ Reading: cli/foo.go (120ms)
🤖 [reviewer] turn 4/30 · git-diff
✓ [reviewer] completed (41s)
📋 2/5 · apply patch
```

The legacy stdout-scraping path remains available with `CHATCLI_GATEWAY_STRUCTURED_PROGRESS=false`.

***

## Window management for workers

Every worker shares the orchestrator's window management instead of running with only the L0 microcompact: at each turn boundary the session compactor (Level 1 trim, Level 2 summary with the configured summarizer, CCR archive through the active tenant's layer) runs when the worker's history crosses the session budget; a turn that fails with a context-overflow error is compacted with the same bounded recovery levels the main loop uses and retried instead of failing the task; and each worker turn is journaled under the parent session, tagged with the worker's name, so the record is complete while the parent's own history is never polluted by worker events. MoA participants recover per thread the same way.

## Configuration

| Variable | Default | Purpose |
| - | - | - |
| `CHATCLI_AGENT_RUNS_HISTORY` | `200` | Finished runs retained for `/agents` and `@agents` |
| `CHATCLI_HUB_RUNS` | `true` | Mirror runs through the hub for cross-process `/agents` (set `false` to disable) |
| `CHATCLI_BOARD_PATH` | `~/.chatcli/board.json` | Board file location |
| `CHATCLI_GATEWAY_STRUCTURED_PROGRESS` | `true` | Structured gateway telemetry (set `false` for legacy scraping) |
| `CHATCLI_HUB_POLL_MS` | `1000` | Cross-process mail poll cadence (shared with hub sync) |

All are surfaced in `/config agent` and `/config gateway`.

<Info>The squad tools are exposed to the LLM both in the prompt tool catalog and as native function-calling definitions (`agents_runs`, `board_cards`, `squad_mail`), so orchestration works on every provider.</Info>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.