Skip to main content
chatcli server (alias chatcli serve) runs ChatCLI as a gRPC service. The API keys live on the server; clients connect with chatcli connect, the operator or any gRPC client. The same binary runs on a laptop, in a container (Docker) and in Kubernetes (Helm chart).

Quick start

1

Start a local server

With no credential the server only listens on loopback:
2

Connect from another terminal

The listener is plaintext, and the client dials TLS unless you allow plaintext explicitly:
3

Make it reachable from other machines

Bind every interface and set a credential, or the server refuses to start (see Bind address and the credential rule):

Bind address and the credential rule

The listen address comes from CHATCLI_BIND_ADDRESS. When it is unset the server picks: An explicit CHATCLI_BIND_ADDRESS always wins. There is no --bind flag. On a loopback address (127.0.0.1, ::1, localhost) the server may run with no credential: the machine is the trust boundary. On any other address a server with no credential would admit every caller as an administrator, so it refuses to start and exits with status 1:
Any one of these satisfies the rule: JWT material that is set but cannot be loaded is not a credential. When JWT is the only credential configured, the server refuses to start with refusing to start: CHATCLI_JWT_PUBLIC_KEY is set but no RSA public key could be loaded: ... (or the matching CHATCLI_JWT_SECRET message). When a token or mTLS is also configured, the server starts, logs the JWT error, and accepts only the working credentials.
The refusal is written to the structured log. In a container, a systemd unit or a pipe the server also writes that log to stderr, so docker logs/kubectl logs show it with no extra setting; on an interactive terminal it goes to the log file only. See Logs.

Flags

A flag given on the command line wins over its env var. chatcli server --help prints the flag list.

Environment variables

Everything the local CLI reads (provider keys such as ANTHROPIC_API_KEY, OPENAI_API_KEY, *_MODEL, *_MAX_TOKENS, CHATCLI_CA_BUNDLE) also configures the server. These are specific to server mode:

Server Authentication

Callers send the credential as gRPC metadata authorization: Bearer <token>. chatcli connect --token sets it for both the shared token and a JWT.
Every caller with the token is the same principal, legacy-token, with role admin. That also means they share one rate-limit bucket.

Roles

Where the server enforces them:

Failed-authentication limiter

Every client host has a budget of failed bearer/JWT authentications: a burst of 5, refilled at one every 12 seconds; the table is cleared every 5 minutes. Only a failed authentication spends from it (a missing, malformed, wrong or expired credential); valid credentials are never throttled, so a busy client, the operator’s probes and a scripted loop of one-shot calls pass at full speed. Once a host has exhausted its budget, its calls are answered with Unauthenticated: authentication failed before the credential is checked, until a slot refills, and the server log records auth failure rate limit exceeded. Callers identified by a client certificate alone do not pass through it.

TLS

The server enables TLS when both --tls-cert and --tls-key are set; the minimum version is TLS 1.3. Neither means plaintext on purpose. One without the other is refused before anything starts, instead of falling back to a plaintext listener: FATAL: --tls-cert is set (server.crt) but --tls-key is not: refusing to start a plaintext listener. Set both, or neither for plaintext (and the mirror message for a key without a certificate). A certificate that cannot be loaded is fatal and printed to stderr: FATAL: TLS certificate load failed: ... (cert=..., key=...). The certificate must be valid for the name clients dial. A private CA for testing:
The certificate is read once at startup: a renewed certificate takes effect after a restart.

LLM credentials

The server calls the model with its own keys unless the request brings credentials. chatcli connect options: The server image does not include the Devin CLI, so DEVIN is not available as a server provider.

Request routing

Every prompt RPC (SendPrompt, StreamPrompt, InteractiveSession, AnalyzeIssue, AgenticStep) resolves its model client the same way:
  • A request that forwards a credential (--llm-key, --use-local-auth, StackSpot fields, an Ollama URL) or names a provider or model gets a dedicated client for that request.
  • A request that names nothing and forwards nothing goes through the fallback chain when one is installed, else the server’s default provider and model.
  • max_tokens is the request value when set, else the provider’s *_MAX_TOKENS override (ANTHROPIC_MAX_TOKENS, OPENAI_MAX_TOKENS, BEDROCK_MAX_TOKENS, …), else the catalog ceiling of the model.
  • A 401/403 from the provider on the server’s credentials rebuilds the providers once and retries; rebuilds are throttled to one per 30 seconds process-wide. Caller-forwarded credentials are never refreshed.
  • A reply stopped by a safety classifier (stop_reason: refusal) is resent once on the sibling model of the same provider; on a stream, only while no text has been sent yet.
Requests run concurrently, each on its own client.

Fallback chain

  • The chain is built from --fallback-providers / CHATCLI_FALLBACK_PROVIDERS alone, in that order. The server does not add --provider to it: list the primary first yourself. The Helm chart and the operator’s spec.fallback put the primary in front for you.
  • Each entry uses CHATCLI_FALLBACK_MODEL_<PROVIDER> when set. Without it, the primary provider (--provider) keeps the server model (--model), and any other provider runs its own default model: its model variable (for example OPENAI_MODEL), then the built-in default. A fallback provider never receives the primary’s model id.
  • A provider whose client cannot be built (no key) is skipped with a warning. The chain is installed only when at least two entries remain; the log then shows Fallback chain initialized.
  • It serves only requests that bring no credentials and name no provider or model. The response names the provider and model that actually answered.
  • A non-empty CHATCLI_FALLBACK_PROVIDERS is the only switch. CHATCLI_FALLBACK_ENABLED is not read by the server, and neither the Helm chart nor the operator sets it; /config marks it as not read when it is set.
See Provider fallback for error classification and cooldowns.

gRPC API

Service chatcli.v1.ChatCLIService (proto: proto/chatcli/v1/chatcli.proto in the repository), plus the standard grpc.health.v1.Health:

Streaming

StreamPrompt forwards each fragment as the provider emits it. The final message (done: true) carries the provider’s usage and stop_reason. When the selected route cannot stream, the reply is produced in one call and sent as a single chunk.

Reply attribution and token usage

SendPrompt, AnalyzeIssue, AgenticStep and the final StreamPrompt message carry a TokenUsage, a stop_reason, and the provider and model that answered (a fallback chain may answer with another provider than the default):
chatcli connect feeds this usage to its cost tracker, so /cost over a remote connection prices real tokens.

Pipeline RPCs

SendPrompt is a model proxy: with chatcli connect, memory, /context attachments, skills, knowledge and compaction run in your CLI and the server only calls the model. The pipeline RPCs make the server own the whole turn, with the same engine the stdio MCP and ACP servers expose:
  • The engine reads its default provider and model from LLM_PROVIDER / LLM_MODEL, not from --provider / --model.
  • Sessions are namespaced by the authenticated principal (<subject>/<session>), so two callers never share a conversation by picking the same id.
  • Without CHATCLI_SERVER_PIPELINE=true the RPCs return Unavailable: the pipeline RPCs are not enabled on this server (start it with CHATCLI_SERVER_PIPELINE=true).
  • GetServerInfo.pipeline_enabled reports whether they are served.
The engine is one ChatCLI inside the server process: its workers (memory, MCP servers, scheduler) run there and its turns are serialized, so concurrent callers queue. It cannot share a process with the co-located gateway: with both CHATCLI_SERVER_PIPELINE=true and CHATCLI_GATEWAY_IN_SERVER=true the gateway runs, the pipeline stays off, and the log says Pipeline RPCs disabled: CHATCLI_SERVER_PIPELINE and CHATCLI_GATEWAY_IN_SERVER cannot share one process; run the gateway separately.
Helm: pipeline.enabled: true on the server chart. Operator: spec.pipeline.enabled: true on the Instance.

AIOps RPCs

GetAlerts, StreamAlerts, AnalyzeIssue and AgenticStep feed the operator’s remediation pipeline. Alerts are the K8s watcher’s:

StreamAlerts

  • With include_current the stream starts with what GetAlerts would return, then only new alerts follow. The watcher deduplicates by type and object within its window, so each alert is pushed once.
  • Heartbeats every 15 seconds tell a quiet watcher from a dead connection. Without a watcher the stream carries heartbeats only.
  • A subscriber that falls behind its bounded buffer is dropped with ABORTED; reopen with include_current to resync.

AnalyzeIssue

Resource discovery

ListRemotePlugins, ExecuteRemotePlugin and DownloadPlugin expose the plugins installed on the server; ListRemoteAgents, GetAgentDefinition, ListRemoteSkills and GetSkillContent expose its agents and skills. The in-REPL /connect command registers the server’s plugins in your session, see Remote Connection.

Health checks

The health RPCs skip the bearer check, but with mTLS the TLS handshake still needs a client certificate (grpc-health-probe -tls-client-cert … -tls-client-key …). The container image ships grpc-health-probe at /usr/local/bin/grpc-health-probe. A kubelet grpc probe cannot do TLS, which is why the Helm chart probes /healthz and a TCP connect instead.

Metrics

--metrics-port (default 9090) serves Prometheus metrics at /metrics (OpenMetrics negotiation supported) and /healthz. The metrics listener binds every interface regardless of CHATCLI_BIND_ADDRESS and has no authentication: firewall it or set --metrics-port 0. Go runtime and process collectors (go_*, process_*) are registered too.

OpenTelemetry export (OTLP)

The ChatCLI engine can push its session counters to an OpenTelemetry collector over OTLP/HTTP (JSON), configured by the standard OTel variables only:
In server mode this covers the turns that run on an engine inside the process, that is the pipeline RPCs and the co-located gateway. Plain SendPrompt/StreamPrompt proxy traffic is measured by the Prometheus metrics above. Exported sums: chatcli.llm.tokens, chatcli.llm.cost, chatcli.context.compactions, chatcli.context.compaction_cost, chatcli.cache.requests, chatcli.cache.storage_cost and, only with OTEL_RESOURCE_ATTRIBUTES containing chatcli.session=attr, chatcli.session.cost.

Limits and keepalive

The rate limiter is a token bucket per caller that runs after authentication. It keys on the caller’s subject: JWT sub, mtls:<name>, legacy-token for the shared token (so every shared-token client shares one bucket), system when no credential is configured; the health RPCs are keyed by peer address. Over the limit a unary call or a stream fails with ResourceExhausted: rate limit exceeded, retry after N seconds and a retry-after header with the same N: the time one token takes to refill, rounded up and never below 1 second (1 at the default 10 rps). Keepalive: the server accepts client pings every 20 seconds or slower (even without active streams) and pings idle connections every 60 seconds, closing them after 10 seconds without an answer. chatcli connect and the operator ping every 30 seconds.

SSRF protection

Provider URLs that a caller supplies (for example --ollama-url, fields base_url, api_base, endpoint, url, host, server_url, realm_url) are checked before use:
  • only https://; http:// requires CHATCLI_ALLOW_HTTP_PROVIDERS=true on the server;
  • the host (or every address it resolves to) must not be in 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16, 100.64.0.0/10, ::1, fc00::/7, fe80::/10, ff00::/8 or IPv4-mapped ranges; there is no override for private addresses;
  • cloud metadata hostnames (metadata.google.internal, metadata.goog, instance-data) are refused.
The server’s own configuration (for example OLLAMA_BASE_URL) is not subject to this check.

Audit log

Every RPC appends one JSON line to a hash-chained, file-locked trail (the same trail the LLM request auditor writes, so both kinds of entries interleave in one verifiable chain). Fields: timestamp, kind (grpc), request_id, action (RPC name), actor (user:<subject> or anonymous), role, ip, client_id, method, resource, result (success, error, denied), duration. A relative path disables the file with an error in the log. Entries also go to the structured log under the audit logger.

Logs

The server uses a structured JSON logger:
  • Every entry goes to the rotating file LOG_FILE (default ~/.chatcli/app.log; CHATCLI_LOG_FILE is an alias).
  • chatcli server and chatcli gateway also write the same entries as JSON lines on stderr whenever stderr is not a terminal — a container, a systemd unit, a pipe — so docker logs, kubectl logs and log collectors get the startup refusals and everything after them with no extra setting. On an interactive terminal the file is the only destination. CHATCLI_LOG_STDERR=true|false forces it either way. Stdout stays free for the stdio transports.
  • CHATCLI_ENV=dev switches to the development console (colored, on stdout) plus the file; the stderr JSON lines are then not duplicated.
  • LOG_LEVEL sets the level. Rotation: LOG_MAX_SIZE (or CHATCLI_LOG_MAX_SIZE_MB, in MB), CHATCLI_LOG_MAX_BACKUPS (default 3), CHATCLI_LOG_MAX_AGE_DAYS (default 28), CHATCLI_LOG_COMPRESS (default true, gzip). These are what the operator’s Instance spec.features.logRotation sets; the server Helm chart sets them from its logging block (20 MB, 3 backups by default, so the log fits the pod’s 200Mi data volume; see Logs, health and metrics).
In the container images the home directory is ephemeral and the distroless image has no shell to read a file with; the stderr stream is what to read there.

gRPC reflection

Reflection needs both the --enable-reflection flag and CHATCLI_GRPC_REFLECTION=true. The flag defaults to the value of CHATCLI_GRPC_REFLECTION, so the env var alone turns it on (this is what the chart’s server.grpcReflection and the Instance’s spec.server.security.enableReflection set). Reflection calls need the same credential as any RPC:
Keep it off in production.

Multiple replicas

gRPC keeps one HTTP/2 connection open, so a ClusterIP Service pins each client to one pod. With more than one replica, use a headless Service: the clients resolve dns:/// to every pod address and balance round-robin. Helm: service.headless: true; the operator switches to headless automatically when spec.replicas > 1. Sessions, the hub database and the audit trail are per pod unless their storage is shared. With the Helm chart, several replicas on the default ReadWriteOnce sessions volume work only on one node; see Rollouts on the sessions volume.

Operating the server

The token is read at startup. Change it and restart:
  • binary: set the new CHATCLI_SERVER_TOKEN and restart the process;
  • Helm with server.token: helm upgrade … --reset-then-reuse-values --set server.token="$(openssl rand -hex 32)"; the Secret checksum annotation rolls the pods;
  • Helm with secrets.existingSecret: update the Secret, then kubectl -n chatcli rollout restart deploy/chatcli.
Clients keep failing with authentication failed until they use the new token. To avoid a hard cut-over, add JWTs first (both credentials work at the same time) and retire the shared token later.
  • HS256: one secret signs and verifies; changing CHATCLI_JWT_SECRET invalidates every issued token at the restart. Keep tokens short-lived.
  • RS256: CHATCLI_JWT_PUBLIC_KEY may hold several PEM blocks, and a token verified by any of them is accepted. Add the new public key next to the old one, restart, switch your issuer to the new private key, and remove the old key once the old tokens have expired.
The certificate and the client CA bundle are loaded at startup. After renewing the files (or the Secret), restart the server. The operator rolls Instance pods when a referenced Secret changes; with the Helm chart run kubectl rollout restart.
  • Binary: /update inside ChatCLI or your package manager, then restart the server.
  • Image: pin the tag (ghcr.io/diillson/chatcli:1.214.0) and change it deliberately; latest moves.
  • Helm: helm upgrade chatcli oci://ghcr.io/diillson/charts/chatcli --version 1.214.0 -n chatcli --reset-then-reuse-values.
GetServerInfo and chatcli_server_info{version=…} report the running version.

Troubleshooting

Next steps

Remote Connection

Connect to the server

Docker & Kubernetes

Run the server in a container or with Helm

K8s Watcher

Kubernetes context for every prompt

K8s Operator

Managed Instances and AIOps