/coder, with its tools (files, shell, web, MCP) and your memory, while the progress and the answer come back to the conversation.
Connect a channel
Pick a channel for its setup, variables and limits.Start and stop
The gateway runs as a detached daemon:/gateway start launches chatcli gateway in the background and gives the prompt back at once, so you keep using ChatCLI as usual.
- Each channel starts only when its required variables are set, so configure just the ones you want. With none configured,
/gateway startsays so and nothing runs. - The daemon keeps its pid in
~/.chatcli/gateway.pid(a secondstartwhile it runs is refused) and writes to~/.chatcli/gateway.log, which is the first place to look when a channel stays quiet. chatcli gatewayruns the same daemon in the foreground, for a service manager or a container; it stops onCtrl+CorSIGTERM.- With
CHATCLI_GATEWAY_IN_SERVER=true, the gateway runs insidechatcli serverinstead and shares its Conversation Hub with connected clients.
How a message is handled
- The adapter receives the message through the platform’s own HTTP API (no third-party SDKs), downloads any voice note or image, and hands it to the runner. Each chat is one conversation, keyed by platform and chat.
- A voice note is transcribed first, so the agent always gets text; an image is attached to the turn for the model to see.
- The agent runs the request with its tools and without confirmations, in the gateway’s conversational voice and in the language of the message.
- While it works, the person sees that it is on it: the native “typing…” on Telegram, or one short notice (
gateway.thinking, ”🤔 Got it — working on that…”) on the other channels, sent only when the reply takes more than about 2 seconds. Progress from the agent’s tool calls is merged and sent as new messages at most every 3 seconds. - The reply is the agent’s final answer, written as a chat message; ”✅” is sent only if the agent finished without one. An error comes back as “⚠️” followed by the error.
Concurrency and context
Messages are received in parallel (up to 64 queued, 4 workers) but agent runs happen one at a time across all channels, because the agent uses shared ChatCLI state. Each turn gets the last 12 turns of the sender’s Conversation Hub thread as context; durable state lives in the files the agent edits and in named sessions.Limits shared by every channel
- A reply is cut at 3500 characters and ends with
…. - One voice note and one image per message are used; downloads are capped by
CHATCLI_GATEWAY_MAX_AUDIO_BYTESandCHATCLI_GATEWAY_MAX_IMAGE_BYTES(20 MB each by default).
Runtime model
The gateway mirrors the model (and provider) your interactive session is using — not the.env default. Switching model or provider in the REPL with /switch, /model, or /max-tokens propagates to the daemon: it re-reads the choice before every message, so a conversation already running on Telegram starts answering with the new model without restarting the gateway.
Because the daemon runs as a separate process, the sync goes through a small state file at ~/.chatcli/runtime_model.json that the interactive session writes and the daemon reads. This covers both cases: starting the gateway after a model switch, and switching while the gateway is already running.
When you switch provider, the daemon adopts the new provider with its correct model — as long as that provider’s credentials are present in the environment the daemon inherited (usually your
.env). Adjustments that live only in memory, such as StackSpot’s /switch --realm / --agent-id, do not propagate through this file; set them via environment variable or restart the gateway.Conversational replies (not “coder tone”)
The gateway runs the same engine as/coder — all the tools (read/edit files, shell, web, MCP) are still available — but with its own conversational voice. The final reply is the message the person reads in chat, not a technical commit summary: direct, natural text, with no tables, banners, ASCII art, or long code blocks (unless code is asked for). Under the hood this is a dedicated gateway system prompt used in place of the coder prompt, preserving the tool-use mechanics.
Dynamic language (follows the sender)
Replies come out in the language of the user’s message, detected every turn — not pinned to the daemon locale. Portuguese → answers in Portuguese; Spanish → Spanish; and so on. The dynamic language directive is applied on every gateway path (including with an active persona), so the reply is never statically stuck in one language. In the interactive CLI, the fixed per-locale directive (CHATCLI_LANG) still applies — only the gateway changes.
Usage example
111111111 sends:
“list the Go files changed in the last commit and summarize the diff”The bot shows “typing…”, sends progress as the agent runs
git and reads the files, and ends with the summary written as a chat message. A message from anyone who is not in CHATCLI_TELEGRAM_ALLOWED_USERS gets no answer.
Voice messages (transcription)
The gateway accepts voice notes and audio on every channel. The message is transcribed to text before the pipeline — so it works with any of the 14 chat providers (they only ever see text; no multimodal model or message redesign needed). The adapter downloads the media, transcribes it, and the agent treats it as a normal text request — the transcript is even recorded in the Conversation Hub.Transcription backend (zero-config, local-first, keyless)
Selection is local/keyless first — and since v1.135 it has an embedded floor: with nothing configured, the gateway uses the embedded Whisper (multilingual, via sherpa-onnx — the same engine as the Kokoro TTS), no API key and no cgo. The daemon pre-downloads engine + model at startup, so the first voice note arrives with everything ready.CHATCLI_TRANSCRIPTION_CMD— your own local STT command (any wrapper). Reads the transcript from stdout, or from the.txtit writes into{output_dir}.CHATCLI_TRANSCRIPTION_URL— a self-hosted OpenAI-compatible endpoint (whisper.cppwhisper-server, faster-whisper, Speaches). Keyless (unlessCHATCLI_TRANSCRIPTION_KEY).- Embedded Whisper already provisioned — when the cache (
~/.cache/chatcli/stt/) holds engine + model, it wins over any cloud key. - A whisper CLI on PATH — if
whisper(openai-whisper) orwhisper-cli(whisper.cpp) is installed, it’s used automatically, zero config. The ggml model is downloaded once to the cache (~/.cache/chatcli/whisper/), like faster-whisper does. GROQ_API_KEY→ Groq Whisper (free tier).OPENAI_API_KEY→ OpenAI Whisper.- Nothing configured → embedded Whisper: one-time download (engine ~25MB +
basemodel ~200MB) when the daemon starts. Only platforms without a prebuilt engine (outside Linux/macOS/Windows x64/arm64) fall back to the configuration hint.
CHATCLI_TRANSCRIPTION_PROVIDER pins a backend (embedded|command|url|groq|openai) — =embedded forces the embedded engine even with whisper/keys present. CHATCLI_TRANSCRIPTION_MODEL picks the embedded model size (tiny|base|small|medium|large-v3, default base) or the cloud model; _LANG pins the language (default: auto-detect the spoken language); CHATCLI_TRANSCRIPTION_CACHE_DIR relocates the cache (absolute path — useful for air-gapped pre-seeding); CHATCLI_GATEWAY_MAX_AUDIO_BYTES caps the download size (default 20MB). The active backend is shown in /config integrations.
Voice notes are OGG/Opus. The embedded engine decodes WAV and OGG/Opus by itself, so Telegram, WhatsApp and Discord voice notes work with nothing else installed; it needs ffmpeg only for MP3, M4A/AAC, FLAC and WMA. A local whisper.cpp cannot decode Opus and needs ffmpeg, which the gateway then uses to convert to 16 kHz WAV automatically. Cloud and self-hosted backends decode on the server. The language is detected from the audio, so the transcript, and the reply, follow the spoken language.
Quick setup
Zero-config (embedded Whisper — recommended):/gateway stop && /gateway start) and send a voice message.
Voice replies
On Telegram the way back speaks too: by default (CHATCLI_GATEWAY_VOICE_REPLY=auto) a voice note gets a voice note back and text gets text, with any TTS backend, including the embedded Kokoro engine (offline, no API key). Each conversation can turn it on or off in plain words (“answer me in audio”, “stop sending audio”) through the @voice tool, and the choice is kept per conversation. The other channels answer in text: the gateway does not synthesize audio for them. Details in Voice replies.
Images
A photo or image file in a message is attached to the turn for the model to see, with native vision when the model supports it; an image sent without text gets a default instruction to describe it. When the agent generates an image during the turn, the first one goes back with the reply on every channel; setCHATCLI_GATEWAY_IMAGE_REPLY=never to send text only.
The gateway also treats the user’s memory index as real knowledge: personal questions (“what do you know about me?”) consult persistent memory through @memory recall before any “I don’t know”.
Cross-channel continuity
When the Conversation Hub is active (the default), the gateway shares the conversation with the chatcli on your notebook: a topic started on Telegram continues in the terminal and vice-versa. Each incoming message resolves the sender’s principal, reads recent context, and records the turn in the hub — so what you said on the notebook shows up as context on Telegram, with zero configuration (single-user mode). For real-time push to a connected CLI, run the gateway inside the server withCHATCLI_GATEWAY_IN_SERVER=true. Multi-user bots use CHATCLI_HUB_ISOLATE=true + bindings. Details in Conversation Hub.
Session binding from the channel
On top of the ephemeral hub, channel users can bind their conversation to a named saved session — the durable, cross-surface layer — by sending/session commands in the chat:
While bound, the turn’s context comes from the named session file — which carries the turns other surfaces (terminal REPL, MCP/ACP server, another channel) wrote through — and each completed gateway turn is appended to that same file. So
/session attach projeto-x on WhatsApp continues the exact conversation you started with /session attach projeto-x in the terminal or the IDE. The binding is per principal (sender) and persisted in the hub’s runtime settings, so it survives daemon restarts.
/session delete is deliberately not exposed to channels: a gateway conversation may be multi-user, and destroying store state stays an operator/REPL decision. And a message that merely starts with / but isn’t a session command flows to the model as normal user text — input is never hijacked.See also
Conversation Hub
One conversation across channels and your terminal.
Proactive messaging
Let the agent message a channel first with
@send.Session management
Named sessions shared by every surface.
Environment variables
Every gateway setting in one table.