Security Overview
The table below summarizes all active protections across every layer of the stack.Authentication and Authorization
JWT Authentication (Recommended)
The gRPC server verifies JWTs with configurable issuer, audience, and either a shared secret (HS256) or an RSA public key (RS256). JWTs carry a role claim that maps to an RBAC level.- RS256 (RSA public key)
- Via Helm
iss and aud, which means a token minted for a different service by the same issuer is accepted — set both wherever a signing key is shared.RBAC Roles
Three role levels exist. The server resolves every caller to one of them and checks it in the handlers that need it:viewer and readonly grant the same thing.
Health RPC and the standard grpc.health.v1.Health service answer without authentication, for load balancers, grpc-health-probe and kubelet gRPC probes. Under mTLS the TLS handshake still requires a client certificate, and a kubelet gRPC probe cannot speak TLS, so probe /healthz on the metrics port instead.CHATCLI_MTLS_ROLE (default user). A server with no credential at all (loopback only) treats every caller as admin.
Legacy Bearer Token
For simpler deployments, the server supports static bearer token authentication with constant-time comparison (crypto/subtle.ConstantTimeCompare), preventing timing attacks. Every holder of the token is the same caller: subject legacy-token, role admin, one shared rate-limit bucket. Prefer the environment variable (or a Kubernetes Secret) over the flag, which is visible in the process list. Clients send it with chatcli connect --token, or CHATCLI_REMOTE_TOKEN.
- Via flag
- Via environment variable
OAuth 2.0 + PKCE
ChatCLI supports OAuth 2.0 with PKCE for the following providers:Encryption and Data Protection
AES-256-GCM Credential Encryption
All OAuth credentials are encrypted at rest using AES-256-GCM in~/.chatcli/auth-profiles.json. The encryption key is automatically generated and stored with strict permissions.
Encryption at Rest (sessions, memory, contexts, archives, costs)
CHATCLI_ENCRYPTION_KEY is set./config security reseal, which now fsyncs. Tenant roots carry a 16-byte digest (roots created with the older short digest keep being used).
Encryption at rest is an explicit opt-in: it is active while CHATCLI_ENCRYPTION_KEY is set in the process environment. When it is, every store that embeds conversation content is sealed before it touches disk and opened transparently on read:
- saved sessions (
/session save,/session attachwrite-through) - exit autosaves and MCP/ACP session mirrors (
autosave-*,mcp-*) - agent park snapshots (
/park,/resume) - the transcript journal (line by line) and
/memory exportfiles - long-term memory JSON stores (
facts,episodes,profile,topics,projects,patterns, graph cache, compactor state — daily notes and rollups stay human-editable Markdown) - knowledge contexts (
~/.chatcli/contexts/*.json;/context exportfiles stay plaintext on purpose) - the CCR archive (
~/.chatcli/ccr/*.ccr, the originals behind@recall) - cost snapshots (
~/.chatcli/costs/*.json)
CHATCLI_ENC_v1 header + 12-byte nonce + AES-256-GCM ciphertext. The key is derived with HKDF-SHA256 from SHA-256(CHATCLI_ENCRYPTION_KEY); the secret itself is never written anywhere.
CHATCLI_ENCRYPTION_KEY — it is never silently treated as empty or corrupt.Persisted redaction and the REPL history. Secret redaction always runs on the LLM path; under CHATCLI_ENV_REDACT_MODE=strict it also masks what ChatCLI persists for itself — session files, the transcript journal, CCR archives, hub mirrors — always on a copy, never the live history (the permissive default keeps stores verbatim so /rewind and exports stay faithful). The redactor covers Slack tokens and webhooks, PEM private keys, connection strings with credentials, AWS secret keys, GCP service-account key ids and Azure account/SAS keys. With the at-rest key set, the REPL prompt history (.chatcli_history) is sealed line by line. Retention re-runs every 6 h in the gateway daemon and expires tenant parks, sessions and queued memory segments past the session window (your own named sessions are never touched).Read-only latch. A memory store whose sealed file this process cannot open (key unset, wrong, or retired without CHATCLI_ENCRYPTION_KEY_PREVIOUS) is locked: it loads empty in memory, logs an error, refuses every write, and is listed under “Locked stores” in /config security and /memory stats. A gateway daemon, a cron job or a shell started without the key can therefore never overwrite your memory. Daily notes and rollups stay plain Markdown by contract; the memory worker’s pending queue is redacted and sealed like the other stores.Key rotation
- Set the new secret in
CHATCLI_ENCRYPTION_KEYand list the retired one inCHATCLI_ENCRYPTION_KEY_PREVIOUS(comma-separated when several). Reads try the current key first, then the retired ones; writes always use the current key. - Run
/config security reseal: every store file under the state root (sessions, transcripts, memory, contexts, CCR, costs — per tenant under the gateway) is rewritten with the current key; plaintext files get sealed on the way. The command reports how many files changed and the key fingerprint. - Unset
CHATCLI_ENCRYPTION_KEY_PREVIOUS.
/config security shows whether encryption is on, the current key’s fingerprint, how many retired keys are configured and what the seal covers.
Tamper-evident audit trail
Every line of the audit trail (CHATCLI_AUDIT_LOG_PATH) carries seq, prev_hash, chain_v and hash — hash = SHA-256(prev_hash ‖ canonical sorted-key JSON of the entry), a form independent of which process wrote it. An edited, removed or reordered line breaks the chain from that point on. Several writers share one file safely: every append takes an exclusive file lock, re-reads the tail when the file changed under it (another writer appended, or the file rotated) and only then links the new line — the REPL, a gateway daemon and the gRPC server (kind: "grpc") form one chain. The file rotates at 64 MiB; the first line of the new file names the file it continues (rotated_from) and links to its last hash, so verification follows the boundary, and the retention pass removes rotated files past the session window (the live file is never touched). With encryption at rest enabled every line is sealed on disk (enc: prefix) and opened transparently on verify. A torn last line (a crash mid-write) is reported as such, never as tampering, and the next entry continues from the last complete line. /config security verify-audit [path] re-hashes the trail and reports the first broken line, the sealed count, the rotation origin, a torn tail and the rotated siblings; trails written before the shared chain still verify with their original hash.
TLS 1.3 Transport Security
- Server TLS
- Mutual TLS (mTLS)
- Development (no TLS)
kubectl logs even if the structured logger cannot flush before the crash.Secret Redaction on the LLM Path
Content the model receives without the user retyping it passes through one redaction chokepoint before it leaves the process: tool outputs in agent/coder mode (file reads, exec, plugins, MCP), squad worker tool outputs, the@file/@git/@env context assembled in chat, and the conversation segment handed to the memory extractor — so a secret pasted into a conversation is neither sent out again nor distilled into a persisted fact.
Two layers compose. KEY=VALUE lines (env dumps, .env files, compose/CI logs) are judged by name — AWS_SECRET_ACCESS_KEY, DATABASE_URL, anything ending in _TOKEN, _PASSWORD, _KEY — plus value heuristics (known prefixes, long hex). Free text is scanned for the value shapes providers hand out (sk-…, ghp_…, AKIA…, JWTs, bearer headers, credential fields in JSON).
PostToolUse hooks still receive it; only the model’s copy is redacted.
OS Keychain Integration
The key that encrypts stored OAuth credentials (~/.chatcli/auth-profiles.json) can live in the OS keychain instead of a file:
file— the file, always. The keychain is never consulted.keychain— the keychain. A key already on disk is migrated into it once, and the file is removed only after the keychain has handed that key back. A write that appeared to succeed and a read that returned nothing would otherwise leave credentials no future process can decrypt.auto(default) — an existing file key keeps being used, untouched. Only a key being created for the first time goes to the keychain, and only where one is available. Relocating a working installation’s key without being asked is not a default’s business.
/config server prints the backend actually in effect, which is not always the one requested.
CredReadW/CredWriteW/CredDeleteW), because cmdkey can create and list credentials but never reveals a secret. Entries are stored per machine rather than roaming.Agent Mode Security
Command Allowlist (Strict Mode)
Strict mode is the default. Only commands on the allowlist run — and the rule applies to every command on the line, not just the first one. A line is a sequence of invocations, so checking only the leading word would make any allowed command a passphrase for the rest of it:File operations
File operations
Text processing
Text processing
Development tools
Development tools
git is on the list, not git status separately. What limits a git push is the denylist and the coder policy, not the allowlist.Containers and infrastructure
Containers and infrastructure
Network
Network
System information
System information
Editors and viewers
Editors and viewers
Custom Allowlist
Extend the allowlist with your own commands. Commas and semicolons both work:Denylist Patterns
The denylist is not limited to permissive mode: it runs on every command in both modes, as the layer that actually refuses destructive work. Around 50 patterns:python -c, perl -e, ruby -e, node -e and php -r are no longer blocked by a regex on the invocation — the inline source is analysed and only high-risk code is refused, so python -c "print(1)" runs and python -c "import os; os.system(...)" does not. The @coder path keeps the stricter regex form and refuses the whole family.Read Path Blocking
In strict workspace mode, the agent can only read files within the current workspace directory:~/.ssh, ~/.gnupg, ~/.aws, ~/.azure, ~/.gcloud and ~/.config/gcloud (everything under them); ~/.kube/config (unless CHATCLI_AGENT_ALLOW_KUBECONFIG=true); ~/.netrc, ~/.npmrc, ~/.docker/config.json, ~/.pypirc, ~/.gem/credentials, ~/.m2/settings.xml and ~/.gradle/gradle.properties; /etc/shadow, /etc/gshadow, /etc/master.passwd and /proc/*/environ; and key material (.pem, .key, .p12, .pfx, .jks, .keystore, .p8, .der) outside the home directory. The refusal names the reason.
Shell Configuration Sourcing
By default, shell configuration files (~/.bashrc, ~/.zshrc) are not sourced during agent command execution to prevent malicious aliases and functions:
Input guard — typeahead protection in security prompts
When a security box appears (coder/agent mode), three layers defend against accidental typing being consumed as a y/n response:- Flush kernel TTY —
TCIFLUSH(Linux) /TIOCFLUSH(BSD/Darwin) /FlushConsoleInputBuffer(Windows) discards bytes in the kernel queue before the box renders. - Drain channel — empties the centralized non-blocking stdin channel (the 10-line buffer the reader goroutine uses).
- Intent debounce — discards any input that arrives in the first 250ms after the box is drawn (minimum human reaction window).
also update the changelog) never answers the prompt, but it is no longer thrown away: the drain re-queues it and it reaches the model at the next turn boundary. From the drain, only lines shaped like a prompt answer (y, n, yes, no, sim, a, always, d, deny, or a bare Enter) stay discarded.
Additionally, at the start of every agent turn, ChatCLI runs stty sane on the controlling /dev/tty to recover from a prior go-prompt teardown that may have left the terminal in raw mode (echo off). Without this reset, you type and don’t see characters on screen — even though the kernel is capturing them.
Output Sanitizer
The stdout and stderr of every agent command pass through a regex redaction of secret shapes (API keys, tokens, credentials in connection strings) before they are stored in the result the model receives, and then through the LLM-path redaction like any other tool output. When the agent hands a command result back to the model (thec<N> and ac<N> continuations), stdout and stderr are also fenced as data in a <COMMAND_OUTPUT cmd="..."> block, prefixed with a warning when prompt-injection phrases are detected, and capped at CHATCLI_MAX_COMMAND_OUTPUT bytes (default 102400, cut on a character boundary and marked [TRUNCATED: output exceeded N bytes]). The terminal shows the full output; only the copy sent to the model is capped. Coder tool results are sized by CHATCLI_TOOL_RESULT_MAX_CHARS instead.
EDITOR Validation
When the user edits commands in agent mode, theEDITOR variable is validated against an allowlist of known editors:
Kubeconfig Access Control
Control whether agent commands can access kubeconfig:Shell Injection Protection
All code paths where dynamic values are interpolated into shell commands use theutils.ShellQuote() function, which applies POSIX quoting with single quotes:
- Quote injection:
'; rm -rf /; echo ' - Command substitution:
$(malicious)or`malicious` - Variable expansion:
$HOME,${PATH} - Pipe/redirection:
| cat /etc/passwd,> /etc/crontab
Binary Resolution via LookPath
Thestty binary (used to restore the terminal) is resolved once at startup via exec.LookPath("stty"), returning the absolute path. This prevents an attacker from placing a malicious stty in the PATH.
Plugin Security
Ed25519 Signature Verification
Plugin binaries are verified with Ed25519 signatures. A plugin is signed with a developer’s private key, and the matching public key must be registered on every machine that installs it.Generate a signing key pair
Sign the plugin
Register the public key on machines that install it
Distribute and check
.sig file next to the binary, and covers the binary’s SHA-256. There is no plugin manifest: replacing the binary breaks the signature, which is what matters. The plugins directory and ~/.chatcli/trusted-keys/ are created 0700; the private key is written 0600, the public key and the .sig 0644.What happens to an unsigned plugin
Quarantine for Unsigned Plugins
A newly seen unsigned plugin can be held out of the runtime for a window — the gap between a binary appearing in the plugins directory and that binary running with ChatCLI’s permissions./plugin quarantine [release <name>] does the same inside the REPL.
- It applies only to unsigned plugins, and only where
CHATCLI_ALLOW_UNSIGNED_PLUGINS=truealready tolerates them. A verified signature is a stronger statement than any waiting period. - Replacing a binary restarts its wait. The review was of the bytes, not of the filename.
- A release records that a human vouched for those exact bytes; state survives a restart, so waiting periods do not reset when ChatCLI starts.
What plugin sandboxing does not do
gRPC Server Security
SSRF Prevention
Provider URLs a client supplies are checked before the server will use them, so a caller cannot point the server at an internal address. The web-fetching tools apply their own dial-time check, including redirects and DNS rebinding. Blocked ranges:10.0.0.0/8,172.16.0.0/12,192.168.0.0/16(RFC 1918)127.0.0.0/8(loopback)169.254.0.0/16(link-local, including cloud metadata endpoints)100.64.0.0/10(shared address space)::1/128,fc00::/7,fd00::/8,fe80::/10,ff00::/8,::ffff:0:0/96(IPv6 loopback, private, link-local, multicast, IPv4-mapped)- the hostnames
metadata.google.internal,metadata.googandinstance-data
https URLs are accepted by default:
Rate Limiting
Token-bucket rate limiting protects against abuse and DoS:sub claim, or the certificate principal under mTLS — rather than on the address it shares with every other tenant behind the same NAT or ingress. Every holder of the shared token is the single subject legacy-token, and a server with no credential (loopback) names every caller system, so each of those shapes shares one bucket; issue JWTs to give callers buckets of their own. Bearer and JWT callers also pass a per-host failure limiter inside authentication: each client host may fail authentication 5 times in a burst, then once every 12 seconds; while its budget is exhausted every call from that host gets Unauthenticated before the credential is checked (the log says auth failure rate limit exceeded). Only failed authentications spend from it — valid credentials are never throttled, however many calls a host or an ingress makes — so it bounds credential guessing without touching CHATCLI_RATE_LIMIT_RPS. The table of hosts is cleared every 5 minutes; callers identified by a client certificate alone skip it. A call over the per-subject limit gets ResourceExhausted.
Message Size Limits
Prevent memory exhaustion from oversized messages:Input Validation
Every RPC the service exposes has a validator, and a test walks the generated service descriptor to keep it that way — a new RPC cannot ship without someone deciding how its request is bounded.- String and byte-length limits on every text field
- Repeated fields bounded (conversation history, insight recommendations, metadata maps)
- Enum validation for severity
- Kubernetes naming rules (RFC 1123) for namespaces, object names and kinds
- Numeric ranges for risk scores, step counters and page limits
Audit Logging
All sensitive operations are recorded in structured JSON audit logs:result distinguishes success, error and denied — a call authentication refused is a security event and does not read like a handler failure. Streaming RPCs produce an entry too, with details.stream and the number of messages received.
LLM request trail (every surface, every provider)
The gRPC audit above records transport metadata. With the sameCHATCLI_AUDIT_LOG_PATH, the CLI also records every LLM request on every surface — REPL, one-shot, gateway, MCP/ACP server, squad workers — as kind: "llm" lines in the same file, one on send and one on receive. The sink hangs off the observability chokepoint all fifteen provider adapters pass through, so a new provider is audited the day it ships. A line carries when, provider and model, payload size, history length, cache markers, outcome, latency, the token usage the provider reported and the running total of secrets the LLM-path redactor rewrote in this process — never the prompt content.
fields also carries caller, the id of that run, on both the send and the receive line: the trail says not only that forty requests were made but which agent made each one. A chat turn or a background job has no caller and the field is absent. The same id is what the Live Dashboard uses to hang a request from its agent.
The path must be absolute; a relative value disables the trail with a logged error. The file is created 0600 and appended by every process that shares the path.
Bind Address
Control which network interface the server listens on:Interceptor Chain
All requests pass through a chain of gRPC interceptors:Validation
Audit
Metrics
0.Recovery
Logging
Auth
Rate Limiting
sub or certificate principal), or on the address for anonymous callers. It runs after auth on purpose, so tenants behind one ingress do not share a bucket.gRPC Reflection (Disabled by Default)
gRPC reflection exposes the full service schema, allowing tools likegrpcurl and grpcui to discover and call all RPCs. In production, this can facilitate reconnaissance by attackers.
spec.server.security.enableReflection and the server chart value server.grpcReflection set that variable.
Kubernetes Operator Security
Fail-Closed Authentication
The operator REST API (port8090, which also serves the dashboard) uses fail-closed authentication: with no API keys loaded every /api/ call gets 401 (“no API keys configured”) unless dev mode is set explicitly, and a request without a known key in the X-API-Key header is denied. API keys are the only mechanism; there is no OIDC, SSO or user login. The dashboard page itself loads without a key and keeps the key you enter in the browser’s localStorage. /healthz and /readyz on that port answer without a key.
Roles are viewer (read), operator (acknowledge, snooze, resolve, approve, reject, runbook edits) and admin (everything, including runbook deletion). Any other role string grants nothing.
Each key entry also accepts an optional name: the identity recorded on the approval decisions that key takes (it falls back to description, then to a key-<hash> fingerprint). An approve or reject through the REST API or the dashboard is recorded as <typed name> (api-key: <identity>), and a quorum counts each key once, so give every approver their own key: a shared key, or dev mode, cannot satisfy a rule that requires two approvers.
The API serves on every operator replica. Requests are rate-limited before authentication: a request without a valid key is limited per client host to 30 per minute (behind an Ingress or proxy every client shares the proxy’s host, and forwarding headers are not trusted), and a valid key is limited to 600 per minute. Excess requests get 429 with Retry-After.
API keys are hot-reloaded every 30 seconds with the following priority order:
- Secret
chatcli-operator-secrets(priority) —api-keysfield containing a YAML list of{key, role, name, description}entries (nameis optional). The operator chart renders it withapiKeys.create: trueandapiKeys.entries; the keys are then stored in the Helm release, so prefer creating the Secret yourself (kubectl, External Secrets, Vault). - ConfigMap
chatcli-operator-config(fallback) — sameapi-keysfield - Reject the request (or accept in dev-mode if
CHATCLI_OPERATOR_DEV_MODE=true)
Resource Type Allowlist
TheApplyManifest remediation action (which applies a manifest stored in a ConfigMap, in the target’s namespace only) creates or updates only resource kinds on an allowlist. The other remediation actions act on the target workload directly and are governed by approval policies, not by this list. The default is wider than a single workload type, because remediation that cannot touch a Service or an HPA is remediation that escalates to a human for routine work:
CHATCLI_ALLOWED_RESOURCE_TYPES also needs a matching ClusterRole rule.
A second list names kinds that are called out as dangerous, and the refusal names the reason — ClusterRole and ClusterRoleBinding (cluster-wide escalation), Role, RoleBinding, Namespace, Node, PersistentVolume, StorageClass, Secret, ServiceAccount, NetworkPolicy, PodSecurityPolicy, the mutating and validating webhook configurations, CustomResourceDefinition, PriorityClass, ResourceQuota and LimitRange. ApplyManifest refuses them, and anything outside both lists, and the remediation attempt fails with that error; there is no approval path that lets such a manifest through.
Log Scrubbing
Before the operator sends an incident’s enrichment context (pod logs, events, metrics, source snippets) to the LLM, it replaces sensitive values with[REDACTED:<type>]. Eighteen built-in patterns cover AWS access keys and secrets, JWTs, bearer tokens, api_key=/password=/token=/secret= assignments, database URIs with credentials, Kubernetes service-account tokens, GitHub tokens and fine-grained PATs, Slack tokens, sk- API keys, private key headers, IPv4 addresses, e-mail addresses, long base64 strings and long hex strings. The operator’s own log output is not scrubbed.
security.logScrubPatterns.
CORS Policy
The operator REST API is deny-all until an origin is named: with none configured, no CORS headers are written and a browser blocks every cross-origin call.Vary: Origin, because the header carries one value and echoing an unmatched origin would turn the list into “any site”. "*" is accepted; combined with credentials it echoes the origin instead, since browsers reject the literal star in that combination. The operator logs which policy took effect at startup.
RBAC and NetworkPolicy
Operator. The operator chart (rbac.create: true) grants the operator a cluster-wide ClusterRole: full access to the ChatCLI CRDs; get/list/watch/create/update/patch on Secrets and ConfigMaps in every namespace; workloads, Services, PVCs, Jobs and RBAC objects it provisions for Instances; pods get/list/watch/create/update/delete (pod remediations and the chaos stress pods); pods/eviction create (DrainNode evicts through the Eviction API, so PodDisruptionBudgets hold); node update (cordon/drain remediations); core and events.k8s.io Events create/patch; create/update on every kind the ApplyManifest allowlist admits, the Prometheus Operator (servicemonitors, podmonitors, prometheusrules) and Istio (serviceentries, virtualservices, destinationrules) kinds included (rules for an API group that is not installed are inert); read-only ReplicaSets; leases for leader election. The same rules are in operator/config/rbac/role.yaml. There is no namespace-scoped mode for the operator. It also pre-provisions the chatcli-watcher ClusterRole (read access to the watched workloads, Jobs and CronJobs included) and the chatcli-role-* ClusterRoles, and may bind only those.
Server chart. The standalone server chart is namespace-scoped by default:
- Namespace-Scoped RBAC (Default)
- Cluster-Wide RBAC
ingressFrom narrows who may connect; egress is a separate choice:
egress: restricted narrows outbound traffic to DNS, HTTPS and the Kubernetes API:
restricted allows DNS, HTTPS, the Kubernetes API, the Instances’ gRPC port and, when prometheusUrl names one, the Prometheus port. The raw manifests carry the same policy in operator/config/network-policy/network-policy.yaml (make deploy-network-policy).
Operator-managed Instances get no NetworkPolicy; write one for their pods (label app.kubernetes.io/instance: <instance name>) that admits the operator namespace on the gRPC port and your Prometheus on the metrics port.
Pod SecurityContext
The server chart defines a restrictive SecurityContext by default (the operator chart does the same without fixing a UID; the operator image runs as65532):
securityContext.readOnlyRootFilesystem is true, the chart automatically mounts an emptyDir volume at /tmp (limited to 100Mi) so the application can write temporary files.runAsNonRoot, UID 1000, RuntimeDefault seccomp, no privilege escalation, read-only root filesystem and every capability dropped, including on the plugin-loader init container, so they pass the restricted Pod Security Standard. spec.securityContext replaces the pod-level part. Nothing labels namespaces for Pod Security Admission; add pod-security.kubernetes.io/enforce: restricted yourself.
Operator Dev Mode
For local development, the operator can run in dev mode:admin instead of refusing it. Once keys are loaded they are enforced as usual. It does not touch TLS or the operator’s connection to the servers.
Operator TLS
The operator has two TLS surfaces. REST API and dashboard (port 8090). HTTP by default; TLS 1.3 when both paths are set:spec.server.tls.enabled: true with a secretName whose certificate is valid for <instance>.<namespace>.svc.cluster.local. The trust root is the ca.crt key of that Secret, else the system CAs. The credential it presents comes from the Instance (spec.server.token, security.operatorTokenRef, short-lived HS256 JWTs it mints from security.jwtSecretRef, or the client certificate in security.operatorClientCertSecretName). Operator-wide fallbacks, mounted the same way through extraVolumes:
TLSConfigured is False (and no Deployment is created) when TLS is enabled without a Secret name, AuthenticationConfigured is False when a reachable server has no credential, OperatorCredentialConfigured says which credential the operator presents, and ServerReachable reports the last probe, repeated every five minutes, or every 30 seconds after a failed probe (see Instance conditions). Rotating any Secret an Instance references rolls its pods.
Container Security (Docker)
The developmentdocker-compose.yml includes the following hardening measures (it also binds 0.0.0.0 inside the container, so it needs CHATCLI_SERVER_TOKEN or JWT material to start):
gcr.io/distroless/static-debian12:nonroot, UID 65532) and carries grpc-health-probe for its HEALTHCHECK. The operator image is Alpine-based and runs as 65532.
CI/CD Security
ChatCLI’s CI/CD pipeline includes multiple security checks:govulncheck
gosec
Trivy
Dependabot
.github/dependabot.yml.Cosign Image Signing
Coder Mode Governance (Policy Manager)
Word Boundary Matching
The policy system uses word boundary matching to prevent permission escalation by prefix. Example:/, =, etc.) and not a word continuation (letter, digit, -, _). This ensures that read does not match readlink.
Default Rules
Read commands are allowed; execution always asks:0600 at ~/.chatcli/coder_policy.json; a coder_policy.json in the working directory overrides it for that project.
The dangerous-command guard sits below the policy
Every@coder subcommand that runs a shell line — exec and test — is checked against the dangerous-pattern list regardless of what the policy says, including after an “allow always”. A command that matches is refused and the model is told not to retry it.
--allow-unsafe and --allow-sudo exist on both subcommands for the cases that genuinely need them, and they lift only the engine’s own check — the guard above still applies in agent and coder mode.Managed Configuration (organization defaults and locked policies)
An operator can ship amanaged.env with the machine image or the MDM profile and have every ChatCLI process on that machine honor it — REPL, one-shot, gateway, MCP/ACP server alike:
.env → managed default → code default. /config managed shows the file, its entries and which are locked; every /config section tags values that came from it as (managed) or (managed · locked). An unreadable file is reported once at boot and ignored (never a crash); a missing file changes nothing.
Security Environment Variables Reference
Complete reference of all security-related environment variables:Server Security
Agent Security
Plugin and Auth Security
Operator Security
Version Check
ChatCLI automatically checks for newer versions on GitHub. To disable (e.g., air-gapped environments or CI/CD):Production Best Practices
Use JWT authentication with RBAC
Enable TLS in production
spec.server.tls.enabled: true with a Secret holding tls.crt, tls.key and ca.crt.Use strict agent security mode
Require plugin signatures
CHATCLI_ALLOW_UNSIGNED_PLUGINS as false (default), sign your plugins and register the public key:CHATCLI_PLUGIN_QUARANTINE=24h so a binary nobody installed on purpose does not run the moment it appears.Require client certificates
Turn on encryption at rest
/config security verify-audit.Sandbox coder execution
Configure rate limiting
Enable audit logging
Keep gRPC reflection disabled
--enable-reflection or set CHATCLI_GRPC_REFLECTION=true in production. Use only for local debugging.Use namespace-scoped RBAC for the server chart
rbac.clusterWide: false (default) unless you need to monitor multiple namespaces. The operator always needs its cluster-wide ClusterRole; review it before installing.Fence the unauthenticated ports
9090, operator 8080) have no authentication. Enable networkPolicy in both charts, restrict metricsIngressFrom / ingressFrom to your monitoring namespace and apiIngressFrom to whatever fronts the dashboard, and write a NetworkPolicy for operator-managed Instance pods.Manage dashboard API keys as Secrets
chatcli-operator-secrets in the operator namespace with one key per team and the lowest role that works, generate keys with openssl rand -hex 32, and rotate by editing the Secret (a removed entry stops working within 30 seconds). Never run with security.devMode: true.Set resource limits
spec.resources sets them:Enable environment variable redaction
Use the OS keychain for the credential key
/config server afterwards: it prints the backend that actually took effect, which is the file wherever no keychain is available.Monitor the audit log
Keep ChatCLI updated
CHATCLI_DISABLE_VERSION_CHECK, check periodically: