chatcli server (alias chatcli serve) runs ChatCLI as a gRPC service. The API keys live on the server; clients connect with chatcli connect, the operator or any gRPC client. The same binary runs on a laptop, in a container (Docker) and in Kubernetes (Helm chart).
Quick start
1
Start a local server
With no credential the server only listens on loopback:
2
Connect from another terminal
The listener is plaintext, and the client dials TLS unless you allow plaintext explicitly:
3
Make it reachable from other machines
Bind every interface and set a credential, or the server refuses to start (see Bind address and the credential rule):
Bind address and the credential rule
The listen address comes fromCHATCLI_BIND_ADDRESS. When it is unset the server picks:
An explicit
CHATCLI_BIND_ADDRESS always wins. There is no --bind flag.
On a loopback address (127.0.0.1, ::1, localhost) the server may run with no credential: the machine is the trust boundary. On any other address a server with no credential would admit every caller as an administrator, so it refuses to start and exits with status 1:
JWT material that is set but cannot be loaded is not a credential. When JWT is the only credential configured, the server refuses to start with
refusing to start: CHATCLI_JWT_PUBLIC_KEY is set but no RSA public key could be loaded: ... (or the matching CHATCLI_JWT_SECRET message). When a token or mTLS is also configured, the server starts, logs the JWT error, and accepts only the working credentials.
Flags
A flag given on the command line wins over its env var.
chatcli server --help prints the flag list.
Environment variables
Everything the local CLI reads (provider keys such asANTHROPIC_API_KEY, OPENAI_API_KEY, *_MODEL, *_MAX_TOKENS, CHATCLI_CA_BUNDLE) also configures the server. These are specific to server mode:
Server Authentication
Callers send the credential as gRPC metadataauthorization: Bearer <token>. chatcli connect --token sets it for both the shared token and a JWT.
- JWT (HS256 or RS256)
- Mutual TLS
Roles
Where the server enforces them:
Failed-authentication limiter
Every client host has a budget of failed bearer/JWT authentications: a burst of 5, refilled at one every 12 seconds; the table is cleared every 5 minutes. Only a failed authentication spends from it (a missing, malformed, wrong or expired credential); valid credentials are never throttled, so a busy client, the operator’s probes and a scripted loop of one-shot calls pass at full speed. Once a host has exhausted its budget, its calls are answered withUnauthenticated: authentication failed before the credential is checked, until a slot refills, and the server log records auth failure rate limit exceeded. Callers identified by a client certificate alone do not pass through it.
TLS
The server enables TLS when both--tls-cert and --tls-key are set; the minimum version is TLS 1.3. Neither means plaintext on purpose. One without the other is refused before anything starts, instead of falling back to a plaintext listener: FATAL: --tls-cert is set (server.crt) but --tls-key is not: refusing to start a plaintext listener. Set both, or neither for plaintext (and the mirror message for a key without a certificate). A certificate that cannot be loaded is fatal and printed to stderr: FATAL: TLS certificate load failed: ... (cert=..., key=...).
The certificate must be valid for the name clients dial. A private CA for testing:
LLM credentials
The server calls the model with its own keys unless the request brings credentials.chatcli connect options:
The server image does not include the Devin CLI, so
DEVIN is not available as a server provider.
Request routing
Every prompt RPC (SendPrompt, StreamPrompt, InteractiveSession, AnalyzeIssue, AgenticStep) resolves its model client the same way:
- A request that forwards a credential (
--llm-key,--use-local-auth, StackSpot fields, an Ollama URL) or names a provider or model gets a dedicated client for that request. - A request that names nothing and forwards nothing goes through the fallback chain when one is installed, else the server’s default provider and model.
max_tokensis the request value when set, else the provider’s*_MAX_TOKENSoverride (ANTHROPIC_MAX_TOKENS,OPENAI_MAX_TOKENS,BEDROCK_MAX_TOKENS, …), else the catalog ceiling of the model.- A 401/403 from the provider on the server’s credentials rebuilds the providers once and retries; rebuilds are throttled to one per 30 seconds process-wide. Caller-forwarded credentials are never refreshed.
- A reply stopped by a safety classifier (
stop_reason: refusal) is resent once on the sibling model of the same provider; on a stream, only while no text has been sent yet.
Fallback chain
- The chain is built from
--fallback-providers/CHATCLI_FALLBACK_PROVIDERSalone, in that order. The server does not add--providerto it: list the primary first yourself. The Helm chart and the operator’sspec.fallbackput the primary in front for you. - Each entry uses
CHATCLI_FALLBACK_MODEL_<PROVIDER>when set. Without it, the primary provider (--provider) keeps the server model (--model), and any other provider runs its own default model: its model variable (for exampleOPENAI_MODEL), then the built-in default. A fallback provider never receives the primary’s model id. - A provider whose client cannot be built (no key) is skipped with a warning. The chain is installed only when at least two entries remain; the log then shows
Fallback chain initialized. - It serves only requests that bring no credentials and name no provider or model. The response names the provider and model that actually answered.
- A non-empty
CHATCLI_FALLBACK_PROVIDERSis the only switch.CHATCLI_FALLBACK_ENABLEDis not read by the server, and neither the Helm chart nor the operator sets it;/configmarks it as not read when it is set.
gRPC API
Servicechatcli.v1.ChatCLIService (proto: proto/chatcli/v1/chatcli.proto in the repository), plus the standard grpc.health.v1.Health:
Streaming
StreamPrompt forwards each fragment as the provider emits it. The final message (done: true) carries the provider’s usage and stop_reason. When the selected route cannot stream, the reply is produced in one call and sent as a single chunk.
Reply attribution and token usage
SendPrompt, AnalyzeIssue, AgenticStep and the final StreamPrompt message carry a TokenUsage, a stop_reason, and the provider and model that answered (a fallback chain may answer with another provider than the default):
chatcli connect feeds this usage to its cost tracker, so /cost over a remote connection prices real tokens.
Pipeline RPCs
SendPrompt is a model proxy: with chatcli connect, memory, /context attachments, skills, knowledge and compaction run in your CLI and the server only calls the model. The pipeline RPCs make the server own the whole turn, with the same engine the stdio MCP and ACP servers expose:
- The engine reads its default provider and model from
LLM_PROVIDER/LLM_MODEL, not from--provider/--model. - Sessions are namespaced by the authenticated principal (
<subject>/<session>), so two callers never share a conversation by picking the same id. - Without
CHATCLI_SERVER_PIPELINE=truethe RPCs returnUnavailable: the pipeline RPCs are not enabled on this server (start it with CHATCLI_SERVER_PIPELINE=true). GetServerInfo.pipeline_enabledreports whether they are served.
pipeline.enabled: true on the server chart. Operator: spec.pipeline.enabled: true on the Instance.
AIOps RPCs
GetAlerts, StreamAlerts, AnalyzeIssue and AgenticStep feed the operator’s remediation pipeline. Alerts are the K8s watcher’s:
StreamAlerts
- With
include_currentthe stream starts with whatGetAlertswould return, then only new alerts follow. The watcher deduplicates by type and object within its window, so each alert is pushed once. - Heartbeats every 15 seconds tell a quiet watcher from a dead connection. Without a watcher the stream carries heartbeats only.
- A subscriber that falls behind its bounded buffer is dropped with
ABORTED; reopen withinclude_currentto resync.
AnalyzeIssue
Resource discovery
ListRemotePlugins, ExecuteRemotePlugin and DownloadPlugin expose the plugins installed on the server; ListRemoteAgents, GetAgentDefinition, ListRemoteSkills and GetSkillContent expose its agents and skills. The in-REPL /connect command registers the server’s plugins in your session, see Remote Connection.
Health checks
grpc-health-probe -tls-client-cert … -tls-client-key …). The container image ships grpc-health-probe at /usr/local/bin/grpc-health-probe. A kubelet grpc probe cannot do TLS, which is why the Helm chart probes /healthz and a TCP connect instead.
Metrics
--metrics-port (default 9090) serves Prometheus metrics at /metrics (OpenMetrics negotiation supported) and /healthz. The metrics listener binds every interface regardless of CHATCLI_BIND_ADDRESS and has no authentication: firewall it or set --metrics-port 0.
Go runtime and process collectors (
go_*, process_*) are registered too.
OpenTelemetry export (OTLP)
The ChatCLI engine can push its session counters to an OpenTelemetry collector over OTLP/HTTP (JSON), configured by the standard OTel variables only:SendPrompt/StreamPrompt proxy traffic is measured by the Prometheus metrics above. Exported sums: chatcli.llm.tokens, chatcli.llm.cost, chatcli.context.compactions, chatcli.context.compaction_cost, chatcli.cache.requests, chatcli.cache.storage_cost and, only with OTEL_RESOURCE_ATTRIBUTES containing chatcli.session=attr, chatcli.session.cost.
Limits and keepalive
The rate limiter is a token bucket per caller that runs after authentication. It keys on the caller’s subject: JWT
sub, mtls:<name>, legacy-token for the shared token (so every shared-token client shares one bucket), system when no credential is configured; the health RPCs are keyed by peer address. Over the limit a unary call or a stream fails with ResourceExhausted: rate limit exceeded, retry after N seconds and a retry-after header with the same N: the time one token takes to refill, rounded up and never below 1 second (1 at the default 10 rps).
Keepalive: the server accepts client pings every 20 seconds or slower (even without active streams) and pings idle connections every 60 seconds, closing them after 10 seconds without an answer. chatcli connect and the operator ping every 30 seconds.
SSRF protection
Provider URLs that a caller supplies (for example--ollama-url, fields base_url, api_base, endpoint, url, host, server_url, realm_url) are checked before use:
- only
https://;http://requiresCHATCLI_ALLOW_HTTP_PROVIDERS=trueon the server; - the host (or every address it resolves to) must not be in
10.0.0.0/8,172.16.0.0/12,192.168.0.0/16,127.0.0.0/8,169.254.0.0/16,100.64.0.0/10,::1,fc00::/7,fe80::/10,ff00::/8or IPv4-mapped ranges; there is no override for private addresses; - cloud metadata hostnames (
metadata.google.internal,metadata.goog,instance-data) are refused.
OLLAMA_BASE_URL) is not subject to this check.
Audit log
timestamp, kind (grpc), request_id, action (RPC name), actor (user:<subject> or anonymous), role, ip, client_id, method, resource, result (success, error, denied), duration. A relative path disables the file with an error in the log. Entries also go to the structured log under the audit logger.
Logs
The server uses a structured JSON logger:- Every entry goes to the rotating file
LOG_FILE(default~/.chatcli/app.log;CHATCLI_LOG_FILEis an alias). chatcli serverandchatcli gatewayalso write the same entries as JSON lines on stderr whenever stderr is not a terminal — a container, a systemd unit, a pipe — sodocker logs,kubectl logsand log collectors get the startup refusals and everything after them with no extra setting. On an interactive terminal the file is the only destination.CHATCLI_LOG_STDERR=true|falseforces it either way. Stdout stays free for the stdio transports.CHATCLI_ENV=devswitches to the development console (colored, on stdout) plus the file; the stderr JSON lines are then not duplicated.LOG_LEVELsets the level. Rotation:LOG_MAX_SIZE(orCHATCLI_LOG_MAX_SIZE_MB, in MB),CHATCLI_LOG_MAX_BACKUPS(default 3),CHATCLI_LOG_MAX_AGE_DAYS(default 28),CHATCLI_LOG_COMPRESS(defaulttrue, gzip). These are what the operator’s Instancespec.features.logRotationsets; the server Helm chart sets them from itsloggingblock (20 MB, 3 backups by default, so the log fits the pod’s 200Mi data volume; see Logs, health and metrics).
gRPC reflection
Reflection needs both the--enable-reflection flag and CHATCLI_GRPC_REFLECTION=true. The flag defaults to the value of CHATCLI_GRPC_REFLECTION, so the env var alone turns it on (this is what the chart’s server.grpcReflection and the Instance’s spec.server.security.enableReflection set). Reflection calls need the same credential as any RPC:
Multiple replicas
gRPC keeps one HTTP/2 connection open, so a ClusterIP Service pins each client to one pod. With more than one replica, use a headless Service: the clients resolvedns:/// to every pod address and balance round-robin. Helm: service.headless: true; the operator switches to headless automatically when spec.replicas > 1. Sessions, the hub database and the audit trail are per pod unless their storage is shared. With the Helm chart, several replicas on the default ReadWriteOnce sessions volume work only on one node; see Rollouts on the sessions volume.
Operating the server
Rotate JWT keys
Rotate JWT keys
- HS256: one secret signs and verifies; changing
CHATCLI_JWT_SECRETinvalidates every issued token at the restart. Keep tokens short-lived. - RS256:
CHATCLI_JWT_PUBLIC_KEYmay hold several PEM blocks, and a token verified by any of them is accepted. Add the new public key next to the old one, restart, switch your issuer to the new private key, and remove the old key once the old tokens have expired.
Renew TLS certificates
Renew TLS certificates
The certificate and the client CA bundle are loaded at startup. After renewing the files (or the Secret), restart the server. The operator rolls Instance pods when a referenced Secret changes; with the Helm chart run
kubectl rollout restart.Upgrade
Upgrade
- Binary:
/updateinside ChatCLI or your package manager, then restart the server. - Image: pin the tag (
ghcr.io/diillson/chatcli:1.214.0) and change it deliberately;latestmoves. - Helm:
helm upgrade chatcli oci://ghcr.io/diillson/charts/chatcli --version 1.214.0 -n chatcli --reset-then-reuse-values.
GetServerInfo and chatcli_server_info{version=…} report the running version.Troubleshooting
Next steps
Remote Connection
Connect to the server
Docker & Kubernetes
Run the server in a container or with Helm
K8s Watcher
Kubernetes context for every prompt
K8s Operator
Managed Instances and AIOps