Skip to main content
This page runs the ChatCLI server in a container and on Kubernetes with the server Helm chart. For operator-managed servers (Instance resources and the AIOps pipeline) see K8s Operator.
Every recipe below sets a credential. A server that listens beyond loopback (any container with a published port, every Kubernetes pod) refuses to start without a shared token, JWT material or a client CA, and by default says so only in its log file. See the credential rule.

Images

Both are multi-arch (linux/amd64, linux/arm64), carry SBOM and provenance attestations and are signed with cosign (keyless):
The server image contains /usr/local/bin/chatcli (stamped with the release version, which GetServerInfo and the operator report), /usr/local/bin/grpc-health-probe, ENTRYPOINT ["chatcli", "server"], EXPOSE 50051, and a HEALTHCHECK that runs grpc-health-probe -addr=:50051 in plaintext every 30 seconds. Arguments after the image name are chatcli server flags. The Devin CLI is not in the image, so DEVIN is not a server provider.

Build locally

Without --build-arg VERSION=… the binary reports version dev.

Docker

The server binds 127.0.0.1 outside Kubernetes, which inside a container means nothing outside the container reaches it. Publishing a port therefore needs CHATCLI_BIND_ADDRESS=0.0.0.0, and that bind needs a credential.
In a container the server writes its log to stderr as JSON lines, so docker logs carries startup errors and everything after them with no extra setting (CHATCLI_ENV=dev switches to the colored development console instead). Check it and connect (the listener is plaintext, so the client needs CHATCLI_ALLOW_INSECURE=true):
For a quick test without a credential, keep the server on loopback inside the container and use it from there only: omit CHATCLI_BIND_ADDRESS and the -p flags.

TLS in Docker

The built-in HEALTHCHECK probes plaintext, so with TLS it reports unhealthy. The image has no shell, and docker run --health-cmd always runs through one, so either disable it (--no-healthcheck) or use an exec-form check in Compose (below). The server certificate needs localhost in its SANs for the client above; see TLS for a certificate recipe.

Docker Compose

The repository’s docker-compose.yml builds the image from source and runs it hardened (read-only root filesystem, no-new-privileges, a 100 MB /tmp tmpfs, 2 CPUs / 1 GB). It binds 0.0.0.0 and reads every provider variable from your shell, so export a credential first:
The compose file sets HOME=/home/nonroot and keeps the whole home in one named volume, chatcli-home, mounted at /home/nonroot: ~/.chatcli (sessions, the conversation hub, memory, plugins, the log file) survives restarts, and the volume is writable under the read-only root filesystem because Docker seeds it with the ownership of the image’s home. Drop plugin binaries into ~/.chatcli/plugins inside that volume. The image’s HEALTHCHECK probes plaintext. With TLS on the server, replace it in an override; Compose merges docker-compose.override.yml automatically:
CHATCLI_FALLBACK_PROVIDERS alone turns on the fallback chain (there is no separate enable switch), and it must list the primary provider first. The compose file also passes CHATCLI_FALLBACK_MAX_RETRIES (default 2), CHATCLI_FALLBACK_COOLDOWN_BASE (30s) and CHATCLI_FALLBACK_COOLDOWN_MAX (5m) from your shell.

Kubernetes (Helm)

The server chart is published as an OCI artifact at oci://ghcr.io/diillson/charts/chatcli. Chart version and server image version are the same number, and the chart’s appVersion pins the image.

Prerequisites

  • Kubernetes 1.30+, kubectl pointed at the cluster
  • Helm 3.8+ (OCI support), Helm 4 included
  • An LLM provider key (or IRSA / Workload Identity for Bedrock)

Install

1

Install with a credential

Inside a pod the server binds 0.0.0.0, so server.token (or JWT / mTLS, below) is required. The chart stores the token in its Secret and delivers it as CHATCLI_SERVER_TOKEN through secretKeyRef, never as a command-line argument. In a pod the server also writes its log to stderr, so kubectl logs shows it.
2

Wait for it

3

Connect

Object names follow the release: a release named other than chatcli produces <release>-chatcli (Deployment, Service and Secret). helm install prints the exact commands in its notes, and warns when no credential is set.
What the chart creates: a Deployment (non-root UID 1000, read-only root filesystem, all capabilities dropped, RuntimeDefault seccomp), a Service, a ConfigMap and a Secret loaded with envFrom, a ServiceAccount and RBAC for the watcher, a 1 Gi PVC for sessions (persistence.enabled, default on), the 17 AIOps CRDs plus a pre-install/pre-upgrade hook that re-applies them (crdUpgrade.enabled), and optionally an Ingress, HPA, PDB, NetworkPolicy and ServiceMonitor.

Probes

A kubelet grpc probe cannot do TLS, so the chart probes: With server.metricsPort: 0 startup and liveness fall back to a TCP check on the gRPC port. These probes work with and without TLS.

Credentials

security.jwtIssuer / security.jwtAudience add iss / aud checks; security.mtlsRole sets the role of certificate-only callers (viewer/readonly, user/operator, admin; default user). Prefer the *Ref values: inline ones land in the Deployment’s environment.

TLS

With tls.existingSecret, certFile and keyFile default to /etc/chatcli/tls/tls.crt and /etc/chatcli/tls/tls.key; set them only when the files in your Secret have other names. Without a Secret, set both paths (to files you mount with extraVolumes / extraVolumeMounts).
tls.enabled: true with neither tls.existingSecret nor both certFile and keyFile fails helm install / helm upgrade (tls.enabled=true needs tls.existingSecret ... or both tls.certFile and tls.keyFile), instead of starting a server that would listen in plaintext.
For mTLS with a client CA that lives in the same Secret use security.tlsClientCA: /etc/chatcli/tls/ca.crt; for a separate CA Secret mount it with extraVolumes / extraVolumeMounts (example in Security values). Through kubectl port-forward the certificate needs localhost in its SANs.

Values reference

Server

LLM

With secrets.existingSecret, create the Secret with variable names as keys, the credential included:

Provider fallback

The chart puts llm.provider (with llm.model) first unless you list it yourself, because the server builds the chain from the list alone and installs it only with two or more working providers. Put every provider’s key in the Secret. fallback.enabled only decides whether the chart writes CHATCLI_FALLBACK_PROVIDERS; the chart writes no separate enable variable, since the server reads none.

MCP

K8s watcher

A target without kind is a Deployment; a target with a namespace omitted watches default. Targets outside the release namespace, or in several namespaces, switch the chart’s RBAC to a ClusterRole automatically, and so does a single-target watcher.namespace other than the release namespace (an empty one counts as default).

Storage, memory, pipeline and resources

The chart mounts an emptyDir at /home/chatcli/.chatcli and /tmp and sets HOME=/home/chatcli, so the read-only root filesystem works. The server reads the MCP, agents, skills and bootstrap ConfigMaps only at startup. Each one the chart renders (from mcp.servers or *.definitions) is hashed into a checksum/<name> pod annotation, so a helm upgrade that edits it rolls the pods. A ConfigMap you manage yourself (*.existingConfigMap) cannot be hashed by Helm: after editing one, run kubectl -n chatcli rollout restart deploy/chatcli. agents, skills or bootstrap enabled with neither definitions nor existingConfigMap mount nothing, and the server finds an empty directory (earlier charts mounted a ConfigMap that did not exist, leaving the pod in ContainerCreating).

Rollouts on the sessions volume

With persistence.enabled and no ReadWriteMany access mode, a rollout stops the old pod before starting the new one: the chart renders RollingUpdate with maxSurge: 0 and maxUnavailable: 1. Under the API server default (one surge pod, no unavailable pod), a new pod scheduled on another node could not attach the ReadWriteOnce volume while the old pod held it, and the old pod was only stopped once the new one was ready, so the rollout hung. Stopping the old pod first means a short gap in service on each rollout instead. Without persistence, or with ReadWriteMany, nothing is rendered and the API server default applies. An explicit strategy wins in every case; type: Recreate without rollingUpdate is rendered with rollingUpdate: null. A fresh install, a Helm 3 release, and a release already upgraded once on the chart’s default can switch to Recreate directly. A release installed by Helm 4 on an earlier chart (up to 1.211.x, which rendered no strategy) needs one upgrade on the default first, because Helm 4’s server-side apply cannot remove the defaulted rollingUpdate block (the upgrade fails with spec.strategy.rollingUpdate: Forbidden), or this patch:
A ReadWriteOnce volume attaches to one node at a time, so several replicas on it work only while they all run on that node; pods scheduled elsewhere stay in ContainerCreating. Run one replica, or use a ReadWriteMany storage class for several. ReadWriteOncePod admits a single pod: the render fails when replicaCount (or autoscaling.maxReplicas with the HPA on) is above 1.

Memory on the sessions PVC

With memory.enabled and persistence, the sessions stay at the PVC root (mounted at ~/.chatcli/sessions) and memory lives in the memory.subPath directory of the same PVC (mounted at ~/.chatcli/memory). kubelet creates that directory on the first mount, and the default podSecurityContext.fsGroup makes it writable. Earlier charts (up to 1.211.x) mounted memory at the PVC root too, so its JSON stores sat among the session files, where the session list showed them and session expiry could delete them. Upgrading an install that ran with memory and persistence migrates on its own. No session moves. The chart sets CHATCLI_MEMORY_LEGACY_DIR=/home/chatcli/.chatcli/sessions (the PVC root as the server sees it) whenever persistence, memory and a memory.subPath are all on, and the first time the server opens memory on the new layout it copies memory’s own files from there into the memory directory:
  • only memory’s files (MEMORY.md and its backups, memory_index.json, memory_tombstones.json, episodes.json, user_profile.json, topics.json, projects.json, usage_stats.json, graph.json, vector_index.json, memory_archive.json, compactor_state.json, their .corrupt quarantines, the YYYYMM/ daily notes, weekly/, monthly/ and pending/); the session files stay where they are;
  • a copy only: nothing at the root is moved, deleted or rewritten, and a file already in the memory directory is never overwritten; each file lands atomically with mode 0600;
  • once: it runs only while the memory directory holds none of memory’s files, and writes the marker .migrated-from-legacy there when done;
  • an unreadable file is skipped with a warning; a file sealed with CHATCLI_ENCRYPTION_KEY is resealed for its new path, and while the key is missing or wrong the copy stays pending (.migrating-from-legacy) and resumes on the next start;
  • the log records it as memory: adopted the legacy memory directory.
memory.subPath: "" keeps the old shared layout instead (no copy, no CHATCLI_MEMORY_LEGACY_DIR). An upgrade with plain --reuse-values carries no memory.subPath key and keeps the old layout too; --reset-then-reuse-values picks up the new default.

Security

Service, ingress, network policy

Scaling and monitoring

The values schema rejects unknown keys, so a misspelled value fails the install instead of being ignored.

Exposing the server

With className: nginx the chart sets backend-protocol: GRPC (plaintext to the pod) and ssl-redirect: "true" as defaults. A key you set in ingress.annotations replaces the default instead of being rendered twice, so for end-to-end TLS through nginx (with tls.* on the server) add nginx.ingress.kubernetes.io/backend-protocol: GRPCS.

Logs, health and metrics

In a pod the server writes every log entry to stderr as JSON lines, which is what kubectl logs and your log collector read; the rotating file at /home/chatcli/.chatcli/app.log (an emptyDir, no shell to read it with) is a copy that dies with the pod. CHATCLI_LOG_STDERR=false in extraEnv turns the stderr copy off. With the default read-only root filesystem, /home/chatcli/.chatcli is an emptyDir limited to 200Mi, and a volume past its limit gets the pod evicted. The server’s own rotation defaults (100 MB, 3 backups: up to 400 MB) do not fit, so the chart sets a rotation that does, at most 80 MB of log: While that emptyDir is in use, the render fails when maxSizeMB × (maxBackups + 1) exceeds 100 MB. A field set to null renders no variable (the server default applies and counts as such in the check). An extraEnv entry of the same name wins over the chart’s value, and one for the size or the backups skips the check.
  • A changed chart-managed Secret or ConfigMap rolls the pods (checksum annotations), the MCP, agents, skills and bootstrap ConfigMaps included. A changed secrets.existingSecret or *.existingConfigMap does not: run kubectl -n chatcli rollout restart deploy/chatcli.
  • With persistence on a volume without ReadWriteMany, each rollout stops the old pod before the new one starts: expect a short gap (see Rollouts on the sessions volume).
  • If both the server chart and the chatcli-operator chart are installed, keep them on the same version: each re-applies its own copy of the CRDs.
  • helm uninstall deletes the sessions PVC; back it up first. CRDs stay; deleting them deletes every resource of those kinds.

Troubleshooting

See also the server troubleshooting table.

Next steps

Server Mode

Flags, auth, limits, operations

Remote Connection

Connect to the server

K8s Watcher

Monitor workloads