Instance resources and, on top of them, an AIOps pipeline that turns watcher alerts into incidents, asks the LLM for a root cause and executes remediation. This page is the operator’s journey: install it, bring up a first Instance end to end, connect to it, open the dashboard, and harden it for production.
The internals of the AIOps pipeline (correlation, analysis, remediation actions, post-mortems) are covered in AIOps Platform and Incident Lifecycle. This page covers everything you need to operate it.
How it fits together
What you need to know before you start:- An
Instanceis one ChatCLI gRPC server (chatcli server) that the operator deploys and keeps in shape: a Deployment, a Service, ConfigMaps, a ServiceAccount and optionally a PVC and watcher RBAC. - The operator always dials the server over TLS 1.3. There is no plaintext path from the operator to an Instance. The address it dials is
<name>.<namespace>.svc.cluster.local:<port>, so the server certificate must be valid for that name. - Every Instance needs a credential. Inside a cluster the server listens on every interface and refuses to start without a shared token, JWT material or a client CA. The operator checks this up front and does not create the Deployment when the spec has none.
- One Instance per cluster drives AIOps. The WatcherBridge lists every Instance in the cluster and attaches to the first ready one it finds. You can run more Instances for chat traffic, but only one feeds the incident pipeline, and which one is not something to rely on.
- The REST API and web dashboard are served by the operator, not by the Instance: Service
chatcli-operator, port8090, authenticated with API keys.
API group and CRDs
All resources are namespaced and live inplatform.chatcli.io/v1alpha1. The operator ships 17 CRDs:
Each AIOps kind has its own page under AIOps Platform in the sidebar, for example Notifications and escalation, SLOs and SLAs and Approval workflow.
Prerequisites
Install the operator
- Helm (recommended)
- Raw manifests (make deploy)
The chart is published as an OCI artifact on GHCR; no repository clone is needed:What it installs: the 17 CRDs, the operator Deployment (image
ghcr.io/diillson/chatcli-operator, tag = chart appVersion), the Service chatcli-operator (ports metrics 8080, health 8081, api 8090), the operator’s ClusterRole and binding, and the pre-provisioned ClusterRoles chatcli-watcher and chatcli-role-{viewer,operator,admin,superadmin}.CRDs on upgrade. Helm installs files from crds/ only on the first install and never updates them. The chart therefore runs a pre-install/pre-upgrade hook Job (crdUpgrade.enabled: true, image registry.k8s.io/kubectl:v1.31.10) that re-applies every CRD, so the schema always matches the controller. Disable it only if you manage CRDs out of band; then apply crds/ from the new chart yourself before upgrading.The Service name follows the Helm release name: with the release
chatcli-operator it is chatcli-operator. Another release name, for example aiops, gives aiops-chatcli-operator. The commands on this page assume the release chatcli-operator in chatcli-system.Verify the install
/readyz (port 8081) answers. Until you create API keys, the log shows SECURITY: no API keys ConfigMap found and CHATCLI_OPERATOR_DEV_MODE is not set and every dashboard call is rejected: that is expected.
Operator chart values
Operator chart values
Your first Instance
The walkthrough creates an Instance namedchatcli in the namespace chatcli. The operator will dial it at chatcli.chatcli.svc.cluster.local:50051; if you pick other names, replace them everywhere, certificate included.
1
Create the namespace
2
Store the LLM API key
The Secret is loaded whole into the server container (Other providers read
envFrom), so its keys are the provider variables:OPENAI_API_KEY, GOOGLEAI_API_KEY, XAI_API_KEY, ZAI_API_KEY, MINIMAX_API_KEY, MOONSHOT_API_KEY, OPENROUTER_API_KEY or GITHUB_COPILOT_TOKEN. Bedrock takes AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY (and optionally AWS_SESSION_TOKEN) plus AWS_REGION or BEDROCK_REGION, or no keys at all with IRSA (see serviceAccount.annotations). Put the key of every provider of a fallback chain in this Secret.This is not the dashboard key Secret.
chatcli-api-keys (any name, referenced by spec.apiKeys.name) lives next to the Instance and holds LLM keys; chatcli-operator-secrets lives in the operator namespace and holds dashboard API keys.3
Create the server token
authorization: Bearer <token>; a caller with the shared token is an administrator on the server. Other credential options are in Authentication options.4
Issue the TLS certificate
The certificate must be valid for the name the operator dials. Include the short names for in-cluster clients and
localhost / 127.0.0.1 so chatcli connect works through kubectl port-forward (the CLI has no server-name override). Add your external DNS name if you expose the server outside the cluster.- openssl (private CA)
- cert-manager
ca.crt is the trust root the operator uses for this Instance. Without it the operator falls back to the system CAs (or security.grpcTLS.caFile), which only works for a certificate from a publicly trusted CA.5
Create the Instance
spec.image.tag is omitted on purpose: the operator runs the server image that matches its own release (ghcr.io/diillson/chatcli:1.214.0).6
Check that it is up and reachable
READY means the Deployment has all its replicas ready. VERSION is what the running server reported to the operator’s own probe: it is only filled once the operator reached the server over TLS with its credential. Read every condition at once:Instance conditions
A ready Instance is probed again every five minutes, so
ServerReachable and VERSION follow a server that was upgraded, restarted or lost its credential. status.serverProbeTime moves when the outcome changes and otherwise at most once per interval (no status write loop).
A ready Instance whose last probe failed is probed again after 30 seconds instead, so a False recorded while the pods were being replaced clears shortly after the rollout. Any change to the Instance, an annotation included, also reconciles and probes it at once without rolling the pods:
Logs worth reading
Connect the CLI
The Instance Service isClusterIP (headless with more than one replica). From a workstation, forward it and connect with TLS:
- Client certificate (mTLS): there is no flag. Set
CHATCLI_TLS_CLIENT_CERTandCHATCLI_TLS_CLIENT_KEY; they are used only together with--tls. - Plaintext: without
--tlsthe CLI still dials TLS with the system CAs unlessCHATCLI_ALLOW_INSECURE=trueis set. A plaintext server is only useful for local development; the operator cannot reach one. - The full client reference is in Remote Connect.
Exposing the server outside the cluster
The operator creates only the ClusterIP Service. To expose an Instance, add your own resource in front of it, and include the external DNS name in the certificate SANs:- L4 load balancer (for example an AWS NLB): create a
Serviceof typeLoadBalancerselectingapp.kubernetes.io/name: chatcliandapp.kubernetes.io/instance: chatcli, port 50051 totargetPort: grpc, with TCP passthrough. TLS stays end to end, which is also the only way client certificates (mTLS) reach the server. - ingress-nginx: the backend speaks TLS, so use
nginx.ingress.kubernetes.io/backend-protocol: "GRPCS", or TLS passthrough (nginx.ingress.kubernetes.io/ssl-passthrough: "true", which requires the controller flag--enable-ssl-passthrough). A terminating ingress cannot carry mTLS client certificates to the server. - Clients then connect with
chatcli connect chatcli.example.com:443 --tls --token "$TOKEN"(add--ca-certfor a private CA).
Dashboard and REST API
The operator serves the web dashboard at/ and the REST API under /api/v1/ on port 8090 of every replica. The API requires an X-API-Key header. Keys are read from the Secret chatcli-operator-secrets (key api-keys) in the operator namespace, falling back to the ConfigMap chatcli-operator-config (same key), and re-read every 30 seconds: adding, rotating or removing a key needs no restart.
- kubectl
- YAML
- Helm values
viewer (read everything), operator (acknowledge, snooze, resolve, approve, review post-mortems, write runbooks), admin (everything, including deleting runbooks). Any other role string grants nothing. An optional name on an entry is the identity recorded on the approval decisions that key takes; a quorum counts each key once, so give each approver their own key.
How the operator applies a change, at startup and on every 30-second poll alike: a Secret without an api-keys entry falls back to the ConfigMap; when neither object provides one (both deleted, or neither holds the entry), every key is revoked within about 30 seconds (401 from then on). An api-keys entry that is not valid YAML keeps the last valid key set in force and is logged once per version, so a typo does not lock everyone out; revoke a key by removing its entry, not by breaking the YAML. A read error other than not found also keeps the keys in force.
- Publishing it: put an Ingress in front of the Service at the root path of a host of its own. The page calls
/api/v1/...with an absolute path, so a sub-path (/chatcli, with a rewrite or strip-path) loads the page and breaks every panel. Examples and controller notes: Web Dashboard: Accessing and publishing. - Rate limits: 30 requests per minute per client host for requests without a valid key (what bounds key guessing), 600 per minute per valid key. Exceeding them returns
429withRetry-After. - TLS: set
security.apiTLS.certFile/keyFileand mount the Secret withextraVolumes/extraVolumeMounts; the API then serves HTTPS with TLS 1.3 only. - Dev mode:
security.devMode: true(CHATCLI_OPERATOR_DEV_MODE, read as a boolean:true,TRUE,1ort) admits every call as admin when no keys are configured; the startup log reports the same mode the API applies. It exists for local experiments; never enable it on a shared cluster. - Endpoints, error codes and pagination: REST API reference. Dashboard tour: Web Dashboard.
Turn on AIOps
The watcher runs inside the Instance pod: it collects pod status, events, logs and optional Prometheus metrics for the targets you list and raises alerts. The operator’s WatcherBridge turns those alerts into the incident pipeline.- Detection. The WatcherBridge (on the operator leader) keeps the server’s
StreamAlertsstream open (heartbeat every 15s; a server without the RPC, older than 1.211.0, is polled withGetAlertsevery 30s, and the stream is retried every 10 minutes or as soon as that server goes away;alertTransport: pollforces polling) and createsAnomalyresources, deduplicated foraiops.dedupTTLMinutes(default 30). - Correlation. Anomalies on the same resource are grouped into an
Issuewith a risk score and severity. - Analysis. An
AIInsightis created and the server’sAnalyzeIssueRPC asks the LLM for a root cause, enriched with Kubernetes context, logs, metrics, GitOps state and linked source code. - Remediation. A
RemediationPlanruns a matching Runbook, an AI-generated runbook, or an agentic observe-decide-act loop (AgenticStepRPC), with 54 typed actions (plusCustom, which the safety checks reject), a pre-action snapshot and automatic rollback. ApprovalPolicies, the decision engine and the cluster tier can park a plan for a human. - Resolution. The Issue resolves or escalates after
aiops.maxRemediationAttempts, and aPostMortemis generated. Notification and escalation are driven by NotificationPolicy and EscalationPolicy.
ExecDiagnostic Allowlist
ExecDiagnostic runs a command inside a pod only when it matches, character for character, an entry of a read-only allowlist of 100 built-in commands (process and filesystem inspection, cgroup v1/v2 files, /proc, network and DNS introspection, curl/wget against localhost health, metrics and pprof endpoints on the common ports, nc -zv to the API server and DNS). Anything else is rejected with command "..." not in approved diagnostic commands whitelist.
The list is read by the operator process, so extend it on the operator chart, not on the Instance. Commas separate entries, so escape them for --set, or put the value in your values file (operator-values.yaml here):
Source repositories
ASourceRepository links a workload to its git repository, so incident analysis sees the recent commits and the code around stack traces. Create it and its Secret in the workload’s namespace. The operator makes a shallow clone under its writable /tmp emptyDir (chart value tmpVolume.sizeLimit, default 1Gi) and re-syncs every syncIntervalMinutes (default 30). The operator image ships git and openssh-client.
Credentials are handed to git for each command and never written to the clone’s
.git/config. A clone made by an earlier version, which kept the token in the origin URL, has the token stripped from origin and its reflog expired on the next sync. Without known_hosts an SSH sync fails unless spec.sshHostKeyPolicy: acceptNew trusts the key the host presents on first contact (and rejects a different key afterwards, for the life of the operator pod); the default is strict.
Watching the pipeline
Authentication options
The operator picks its own credential in this order:
server.token, then security.operatorTokenRef, then security.jwtSecretRef (minted tokens), then security.operatorClientCertSecretName. A client certificate, when configured, is presented on every connection on top of any bearer credential.
JWT HS256 (the operator mints its tokens)
JWT HS256 (the operator mints its tokens)
exp (a token without it is rejected), are checked with 30 seconds of clock skew, and honor nbf; iss and aud are checked when configured. The role claim maps admin to admin, operator and user to user, viewer and readonly to read-only; no claim means user. The operator mints its own tokens (sub: chatcli-operator, role: operator, one hour, renewed after 45 minutes) stamped with jwtIssuer and jwtAudience. How to issue tokens for people: Server Mode authentication.JWT RS256 (your identity provider signs)
JWT RS256 (your identity provider signs)
operatorTokenRef (or a client certificate). That token must carry exp like any other, so keep the Secret refreshed before it expires (for example with External Secrets); the operator reconnects when the Secret changes, without rolling the server pods.mTLS (client certificates)
mTLS (client certificates)
mtlsRole (viewer, readonly, user, operator or admin; server default user); a bearer token, when also sent, takes precedence. Issue client certificates with extendedKeyUsage = clientAuth from the same CA as in the TLS step.security.bindAddress: "127.0.0.1") may run without a credential, but then nothing outside the pod reaches them, the operator included: ServerReachable stays False. It only makes sense for a server you reach from a sidecar.
Instance spec reference
Onlyspec.provider is required. The CRD schema is the only validation: there is no admission webhook.
Top level
spec.server
spec.server.security
spec.watcher
targets[]: name (the resource name; deployment is its deprecated alias), kind (Deployment, StatefulSet, DaemonSet, Job, CronJob; default Deployment), namespace (required), metricsPort (0 = no metrics), metricsPath (server default /metrics), metricsFilter (glob patterns).
spec.fallback
The operator puts
spec.provider (with spec.model) first unless the list already names it, in which case your order is kept. The server builds the chain only when at least two of its providers have credentials, and uses it for requests that bring no credentials or explicit provider of their own. Put every provider’s key in the apiKeys Secret. Behavior details: Provider Fallback.
spec.aiops
The AIOps pipeline reads these settings from the Instance the WatcherBridge streams alerts from (labels platform.chatcli.io/instance and platform.chatcli.io/instance-namespace on each Anomaly and Issue), falling back to the first Ready Instance.
spec.features and spec.pipeline
spec.mcp, spec.agents, spec.plugins
Status
What the operator creates
Everything is named after the Instance, labeledapp.kubernetes.io/name: chatcli, app.kubernetes.io/instance: <name>, app.kubernetes.io/managed-by: chatcli-operator, and owned by the Instance (deleted with it).
The Deployment runs
chatcli server --port <port> --metrics-port <metricsPort> --provider ... [--model ...] [--tls-cert ... --tls-key ... [--tls-client-ca ...]] [--mcp-config ...] [--watch-config ... | --watch-deployment ...] with:
- strategy
Recreatewhenpersistence.enabled(the sessions PVC is ReadWriteOnce, so the old pod releases it before the new one starts; each rollout has a brief downtime), otherwiseRollingUpdate; - probes on
GET /healthzof themetricsport (plain HTTP, no credential): startup every 5s up to 5 minutes, readiness every 10s (out of the Service after 30s of failures), liveness every 20s (restart after 2 minutes); - pod security context
runAsNonRoot, UID 1000,fsGroup: 1000withfsGroupChangePolicy: OnRootMismatch(so a fresh sessions volume is writable whatever the provisioner does),seccompProfile: RuntimeDefault(unlessspec.securityContext), and for every container, theplugin-loaderinit container included,allowPrivilegeEscalation: false,readOnlyRootFilesystem: true, all capabilities dropped: the pod passes therestrictedPod Security Standard; - writable emptyDirs
/tmp(100Mi) and/home/chatcli/.chatcli(200Mi); anything outside the sessions PVC is lost when the pod restarts; HOME=/home/chatcli, the<name>ConfigMap and theapiKeysSecret asenvFrom, thenextraEnv, then the typed credential and feature variables. Secret references are optional, so a missing Secret does not block the pod: it starts without the value.
Rollout triggers
The pod reads its configuration at startup, so the operator stamps hashes on the pod template and a change rolls the pods:
Changes to the pod template itself (image, resources, env, scheduling) roll the pods as with any Deployment. Every Secret the Instance references is watched, so creating or rotating one reconciles the Instance right away.
What refreshes only the operator:
operatorTokenRef and operatorClientCertSecretName are read by the operator, not by the pod. Rotating them does not restart the server; the operator rebuilds its connection with the new material (the status probe dials fresh every time).
What does not roll the pods: a new plugin image under the same tag. Run kubectl -n <ns> rollout restart deploy/<name> after pushing one.
Running in production
Checklist
- The TLS certificate covers
<name>.<namespace>.svc.cluster.local(plus the short names, and your external name if exposed), and the Secret hasca.crtfor a private CA. -
TLSConfigured,AuthenticationConfigured,OperatorCredentialConfigured,AvailableandServerReachableare allTrue. - A credential per audience: a shared token only for automation you trust with admin, JWTs (HS256 or RS256) or mTLS for people.
-
spec.resourcesrequests and limits are set on every Instance. - Dashboard API keys live in
chatcli-operator-secrets, one per team with the least role it needs, andsecurity.devModeis off. - NetworkPolicies restrict the Instance and operator pods, including the plain-HTTP metrics ports.
- Images are pinned (operator chart version;
spec.image.tagonly if you want to decouple the server from the operator release). - Only one Instance in the cluster is meant to drive AIOps.
- The ExecDiagnostic allowlist covers your workloads’ health ports.
- Custom resources and referenced Secrets are backed up.
TLS and RBAC
- TLS everywhere the operator talks. The operator dials Instances with TLS 1.3 only, verifying the certificate against the Instance’s
ca.crt(or the system CAs, orsecurity.grpcTLS.caFile). The server itself accepts TLS 1.3 only. Rotate with cert-manager or by replacing the Secret: the pods roll (chatcli.io/tls-hash) and the operator trusts the newca.crton its next dial. - Operator RBAC. The operator’s ClusterRole is broad by necessity: full access to its 17 CRDs; Deployments, Services, ConfigMaps, ServiceAccounts and PVCs; read and write on Secrets cluster-wide (it reads referenced Secrets and executes
RotateSecret); pods (including create, for chaos stress pods), pod eviction (DrainNodeuses the Eviction API), logs, and core andevents.k8s.ioevents; nodes (cordon/drain); create and update on every kind theApplyManifestallowlist admits (workloads, Jobs, CronJobs, Services, ConfigMaps, HPAs, PDBs, Ingresses, and the Prometheus Operator and Istio kinds; rules for an API group that is not installed are inert), plus NetworkPolicies, for remediation; read-only ReplicaSets; leases for leader election. - No runtime RBAC escalation. The operator never creates or modifies ClusterRoles. It may only
bindthe pre-provisionedchatcli-watcherandchatcli-role-{viewer,operator,admin,superadmin}(restricted withresourceNames), so a compromised operator cannot bind a more privileged ClusterRole. Thechatcli-role-*ClusterRoles are there for you to bind to people; the operator binds nothing to users. - Remediation guardrails.
ApplyManifestonly creates or updates the 16 allowed kinds (Deployment, StatefulSet, DaemonSet, Service, ConfigMap, HorizontalPodAutoscaler, PodDisruptionBudget, Ingress, CronJob, Job, ServiceMonitor, PrometheusRule, PodMonitor, ServiceEntry, VirtualService, DestinationRule; extend withsecurity.allowedResourceTypes) in the Issue’s namespace. The operator RBAC (chart and kustomize) grants create and update on each of them, the ServiceMonitor, PodMonitor, PrometheusRule, ServiceEntry, VirtualService and DestinationRule kinds included; a kind you add withsecurity.allowedResourceTypesalso needs a ClusterRole rule you add. ReplicaSet is not on the list: its Deployment owns it and would revert a direct write.ExecDiagnosticis limited to the allowlist. Logs pass through 18 scrub patterns (cloud and API keys, JWTs, bearer tokens, passwords, connection strings, private keys, IPv4 addresses, e-mail addresses, long base64/hex secrets) before they reach the LLM.
Credentials and rotation
Rotating the shared token rolls the server and switches the operator at the same time, but every CLI user needs the new value. For zero-downtime user rotation, prefer JWTs.
Resources
The operator applies no requests or limits to Instances, which leaves the pods BestEffort: first to be evicted under node pressure. Start fromrequests: {cpu: 250m, memory: 256Mi} and limits: {cpu: "1", memory: 1Gi} and adjust to your traffic; the watcher, MCP servers and the pipeline RPCs add to the footprint. The operator itself defaults to 100m/128Mi requests and 500m/256Mi limits; raise them on clusters with many Issues.
High availability
- Operator: run
replicaCount: 2or more withleaderElect: true. The controllers and the WatcherBridge run on the leader only; the REST API and dashboard are served by every replica, so the Service keeps answering during a failover. The chart creates no PodDisruptionBudget for the operator; add one if you drain nodes often. - Instances:
replicas> 1 turns the Service headless and the operator balances round-robin across pods. Each pod is an independent server (its own memory, watcher and emptyDir state), and the alert stream attaches to one pod at a time. For an AIOps Instance, one replica is usually right. - Persistence and replicas: the sessions PVC is ReadWriteOnce, so with
persistence.enabledthe Deployment uses theRecreatestrategy: the old pod stops and releases the volume before the new one starts, and each rollout (image, configuration or credential change) has a brief downtime. Without persistence the strategy isRollingUpdate. A second replica scheduled on another node still cannot attach the volume (Multi-Attach error): keep one replica with persistence and pin it withscheduling.affinityif your storage is zonal, or run without persistence when you scale out. - The operator creates no PodDisruptionBudget, HorizontalPodAutoscaler or NetworkPolicy for Instances.
Network policies
-
Operator: enable
networkPolicy.enabledin the chart (ormake deploy-network-policyfor the raw manifests). Ingress is allowed on the API port (narrow it withnetworkPolicy.apiIngressFrom), the metrics port (metricsIngressFrom) and the health port. Withegress: restricted, egress is limited to DNS, 443, the Kubernetes API port (kubernetesApiPort, 6443), the Instances’ gRPC port (instanceGrpcPort, 50051) and the Prometheus port taken fromprometheusUrl; addegressExtraPortsfor SMTP (587/465), git over SSH (22) or webhooks on other ports. -
Instances: the operator creates none. A starting point:
Add a
fromfor your own clients on 50051, and egress to the watcher targets’metricsPortwhen you collect their metrics. Kubelet probe traffic is exempt from NetworkPolicy on most CNIs but not all, which is why the example leaves 9090 open to any source.
Pod Security
Instance pods and the operator pod meet therestricted Pod Security Standard with the defaults, so you can label their namespaces pod-security.kubernetes.io/enforce: restricted. If you set spec.securityContext, it replaces the default entirely: keep runAsNonRoot: true and seccompProfile: {type: RuntimeDefault} in it. The default also sets fsGroup: 1000 with fsGroupChangePolicy: OnRootMismatch, so a fresh sessions volume is writable whatever the provisioner does; keep fsGroup: 1000 in your own context when you enable persistence.
Metrics and monitoring
-
The server exposes
/metrics(and/healthz) on the Instance Service portmetrics(9090); the operator exposes controller and AIOps metrics (chatcli_operator_*) on 8080 (serviceMonitor.enabled: truein the chart). -
Both are plain HTTP without authentication, and the server’s metrics listener binds every interface regardless of
bindAddress. Restrict them with NetworkPolicy. -
For an Instance, create your own ServiceMonitor:
-
Metric names, Grafana dashboards and useful queries: Web Dashboard. Alert on
chatcli_operator_instance_ready == 0and on theServerReachablecondition.
Images, versions and upgrades
-
Operator and server images are published as
ghcr.io/diillson/chatcli-operator:<version>andghcr.io/diillson/chatcli:<version>(pluslatest), multi-arch (amd64, arm64), with SBOM and provenance, signed keyless with cosign: -
Pinning. The chart pins the operator image to its
appVersionand passesCHATCLI_OPERATOR_APP_VERSION, so an Instance withoutspec.image.tagruns the same release and ahelm upgradeof the operator upgrades those servers too. Setspec.image.tagonly to decouple an Instance from the operator release. Avoidlatestin production. - Upgrading and uninstalling: see Upgrade and uninstall.
Backup
- Custom resources, for example
kubectl get instances,runbooks,notificationpolicies,escalationpolicies,approvalpolicies,incidentslas,servicelevelobjectives,sourcerepositories,clusterregistrations -A -o yaml; addissues,postmortems,auditeventsif you keep incident history in the cluster. Tools such as Velero back up namespaces with their custom resources. - The Secrets the Instances reference, and
chatcli-operator-secrets, from your secret store. - The sessions PVC holds server sessions. It is owned by the Instance: deleting the Instance deletes the claim (the volume follows its StorageClass reclaim policy).
- LLM spend booked by the pipeline is kept in ConfigMaps
chatcli-cost-ledgerin the namespaces where Issues occur.
What does not exist
- No admission or validating webhook: the CRD schema is the only validation, and the operator reports the rest as conditions.
- No plaintext connection from the operator to an Instance.
- No gRPC reflection unless you set
spec.server.security.enableReflection. - No PodDisruptionBudget, HPA, NetworkPolicy, ServiceMonitor, Ingress or LoadBalancer created for Instances; the Service is always ClusterIP or headless.
- No TLS or authentication on the metrics endpoints.
- No SSO/OIDC for the dashboard: API keys only.
- No file audit log in the operator: its actions are AuditEvent resources. The server has its own audit file (
security.auditLogPath). - No automatic RBAC for people: bind the
chatcli-role-*ClusterRoles yourself.
Upgrade and uninstall
An upgrade replaces the operator and, for every Instance that follows the operator release (nospec.image.tag), its server too. You can go straight from any earlier release to this one: the CRD hook re-applies the complete schema and the new operator re-renders every Instance, so there are no intermediate versions to install. Read Crossing these releases first when you come from an older one.
Before you start
- Back up the custom resources and Secrets (see Backup).
-
Save the values the release runs with, and keep that file as the source of your settings from now on:
A release installed without overrides, like the install command on this page, prints
null; the file works as it is. - Check Crossing these releases for anything that needs action before the upgrade.
Run the upgrade
- Helm
- Raw manifests
--reset-then-reuse-values (Helm 3.14+) starts from the new chart’s defaults and re-applies your previous overrides. Do not use plain --reuse-values: it skips the defaults of keys added by the new chart and can fail to render (nil pointer evaluating).What happens, in order
- CRDs. The pre-upgrade hook Job re-applies the 17 CRDs (
kubectl apply --server-side) before anything else changes. - Operator. The operator Deployment rolls: the new pod becomes Ready before the old one stops, so the dashboard and REST API stay up. The controllers restart on the new pod.
- Instances. The new operator re-renders every Instance. An Instance without
spec.image.taggets the new server image, and any change to the pod template rolls its pods. Withpersistence.enabledthe strategy isRecreate: the old pods stop before the new ones start, so each such Instance has a short outage. - Probe. The operator probes each server again. A probe that lands inside the rollout can record
ServerReachable=False/ProbeFailedwithDeadlineExceeded ... while waiting for connections to become ready,connection refusedori/o timeout, even though the new pods serve fine a moment later. A failed probe is retried after 30 seconds, so the condition turnsTrueshortly after the rollout. - Alerts. The WatcherBridge reconnects. If it reaches a server older than 1.211.0, which has no
StreamAlerts(typically an old pod that is still running), it logsServer has no StreamAlerts RPC; polling GetAlertsand polls every 30 seconds. When that old pod goes away, the next poll fails withUnavailable, the bridge logsServer unavailable while polling; retrying the alert stream on reconnectand opens the stream on the new server. No alert is lost meanwhile.
Verify
True. If ServerReachable still shows the False captured during the rollout, wait about 30 seconds for the retry.
Crossing these releases
Uninstall
Delete the Instances first, while the operator still runs: their finalizer (platform.chatcli.io/finalizer) is removed by the operator, and an Instance deleted after the operator is gone stays Terminating until you remove it by hand (kubectl patch instance <name> -n <ns> --type merge -p '{"metadata":{"finalizers":null}}'). Then run helm uninstall chatcli-operator -n chatcli-system (or make undeploy for the raw manifests, which also deletes the CRDs). helm uninstall leaves the CRDs (Helm never deletes crds/); deleting them deletes every custom resource.
Troubleshooting
Development
make manifests only prints the controller-gen command; the generated CRDs are committed in config/crd/bases/ and copied into both charts’ crds/.
Next steps
AIOps Platform
How the pipeline works inside
Incident Lifecycle
States, remediation actions and rollback
Server Mode
Flags, authentication and the gRPC RPCs
Production setup
Recipe: a production AIOps setup