Platform v2 Components
Notifications
SLO & SLA
Approvals
Multi-Cluster
Audit
Chaos Engineering
Decision Engine
Capacity & Cost
Web Dashboard
REST API & Dashboard
The operator also exposes a REST HTTP API on port8090 (chart value api.port, env CHATCLI_AIOPS_PORT), reached through the Service chatcli-operator in the operator namespace (not the Instance Service). Its 40+ endpoints cover incidents, AI insights, remediations, runbooks, approvals, SLOs, post-mortems, analytics (including LLM cost), clusters, federation, policies and audit. A Web Dashboard is embedded and served at / on the same port.
- Authentication: header
X-API-Key. Keys are read from the Secretchatcli-operator-secrets, keyapi-keys(fallback: ConfigMapchatcli-operator-config), in the operator namespace, as a YAML list of{key, role, description}. Roles areviewer<operator<admin; any other role string is denied. Changes are picked up within about 30 seconds; anapi-keysentry that is not valid YAML keeps the last valid key set in force, and a Secret without the entry falls back to the ConfigMap. Create it yourself, or let the operator chart render it (apiKeys.create: truewithapiKeys.entries). With no keys configured every/api/call returns401, unlessCHATCLI_OPERATOR_DEV_MODE=true(admin without a key, development only). - Rate limit: 30 requests per minute per client host without a valid API key, 600 per minute per valid key (
429withRetry-After). - CORS: deny-all unless
CHATCLI_CORS_ALLOWED_ORIGINS/CHATCLI_CORS_ORIGINare set. TLS: setCHATCLI_AIOPS_TLS_CERTandCHATCLI_AIOPS_TLS_KEYto serve HTTPS (TLS 1.3). - Local preview:
make dash-previewinoperator/serves the dashboard over synthetic data, no cluster needed, athttp://127.0.0.1:8085with the API keypreview.
replicaCount > 1 works behind the Service without pinning requests to the leader.deploy/grafana/. The dashboards-configmap.yaml there holds only the two ServiceMonitors (operator and server, the latter selecting app.kubernetes.io/name: chatcli), not the dashboards: create the ConfigMap by naming each JSON file (see Web Dashboard) or import the JSON files by hand.
Pipeline Overview
Internal Components
1. WatcherBridge (watcher_bridge.go)
The WatcherBridge is the pipeline entry point. It implements the controller-runtime manager.Runnable interface and runs as a manager-managed goroutine.
Responsibilities:
- No temporal component: A continuous problem (e.g., CrashLoopBackOff) generates only one Anomaly
- Resource UID: a deleted and recreated workload gets a new UID, so its alerts are not swallowed by the old entry
- TTL: 30 minutes by default (Instance
spec.aiops.dedupTTLMinutes, 5–1440) — expired hashes are pruned automatically - Invalidation: When an Issue is resolved, contained or escalated, dedup entries for the affected resource are invalidated, allowing immediate recurrence detection
- Result: Avoids duplicates during an active problem; detects recurrence after resolution
Lists Instance CRs in the whole cluster
Selects the first Instance with Status.Ready=true
Connects over gRPC with TLS 1.3
dns:///<name>.<namespace>.svc.cluster.local:<port>, with the credential and CA taken from the Instance spec. TLS is mandatory: the Instance needs spec.server.tls.enabled: true and a certificate valid for that name.Retry
2. AnomalyReconciler (anomaly_controller.go)
Watches Anomaly CRs and correlates them into Issues.
Flow:
Receives Anomaly CR
Status.Correlated = false.Attaches to an active Issue
Suppression checks
spec.aiops.resolutionCooldownMinutes, default 10; 0 turns the cooldown off) or when the noise reducer flags it (repetitive, flapping, seasonal).Groups anomalies and calculates risk score and severity
CorrelationEngine.FindRelatedAnomalies() for uncorrelated anomalies on the same resource within 10 minutes.Creates the Issue CR
<resource>-<signal>-<unix time>, with labels platform.chatcli.io/inc-id (INC-YYYYMMDD-NNN), platform.chatcli.io/resource and platform.chatcli.io/signal.Marks Anomaly as correlated
Correlated = true with reference to the Issue.3. CorrelationEngine (correlation.go)
Correlation engine that groups anomalies into incidents.
Correlation Algorithm:
pod_restart (25) + memory_high (15) = risk 40 -> Medium. Adding error_rate (30) = risk 70 -> High. An Issue opened by an oom_kill anomaly is Critical whatever the score.
Source Mapping: the Issue source mirrors the Anomaly source (watcher, prometheus, events, logs, webhook); an unknown source maps to prometheus.
4. IssueReconciler (issue_controller.go)
Manages the complete lifecycle of an Issue through a state machine.
States and Transitions:
Contained means the plan silenced the workload (for example ScaleDeployment to 0 with containment=true) without fixing it: the Issue carries status.requiresHumanAction: true and status.requiredAction, and it moves to Resolved only once the workload is restored. Escalated Issues are re-checked every 30 seconds and auto-resolve when the resource is healthy again, unless the Instance sets spec.aiops.enableAutoResolve: false. The health check understands Deployments, StatefulSets, DaemonSets, Jobs (healthy once Complete) and Nodes (healthy when Ready); an Issue on any other kind is not auto-resolved. Failed is terminal.
The AIOps settings (spec.aiops) come from the Instance the WatcherBridge uses: the WatcherBridge labels each Anomaly with platform.chatcli.io/instance and platform.chatcli.io/instance-namespace, the Issue inherits them, and without the labels the first Ready Instance applies.
handleDetected()
handleDetected()
- Sets
detectedAtandmaxRemediationAttempts(default: 5, configurable via Instanceaiops.maxRemediationAttempts; read from the Instance the WatcherBridge uses) - Creates AIInsight CR
<issue>-insightwith owner reference (Issue -> AIInsight), annotated with the candidate Runbooks - Transitions to
Analyzing - Runs the federation checks (cascade detection, cross-cluster correlation), best effort
- Requeues after 10 seconds
handleAnalyzing()
handleAnalyzing()
- Checks if AIInsight has
Analysispopulated - Searches for matching manual Runbook (
findMatchingRunbook— tiered matching) - If manual Runbook found ->
createRemediationPlan()(manual has precedence) - If no manual Runbook but AIInsight has
SuggestedActions->generateRunbookFromAI()->createRemediationPlan()using the auto-generated Runbook - If none ->
createAgenticRemediationPlan()(AgenticMode=true, no pre-defined actions — AI decides each step) - Transitions to
Remediating
findMatchingRunbook() -- Tiered Matching
findMatchingRunbook() -- Tiered Matching
- Tier 1: SignalType + Severity + ResourceKind (exact match, preferred)
- Tier 2: Severity + ResourceKind (fallback when signal doesn’t match)
SignalTyperesolved from:issue.Spec.SignalType-> fallbackissue.Labels["platform.chatcli.io/signal"]- Runbooks are searched in every namespace, not only the Issue’s: the Issue’s namespace first, then the others, each Runbook once
generateRunbookFromAI()
generateRunbookFromAI()
- Materializes
SuggestedActionsfrom AI as a reusable Runbook CR - Name:
auto-{signal}-{severity}-{kind}-{hash}(sanitized;hash= first 6 hex chars of the SHA256 of the analysis, so different root causes produce different Runbooks) - Labels:
platform.chatcli.io/auto-generated=true - Trigger: SignalType + Severity + ResourceKind (for future reuse)
- Uses
CreateOrUpdatefor idempotency
handleRemediating()
handleRemediating()
- Finds the most recent RemediationPlan (
findLatestRemediationPlan) - If
Completed-> IssueResolved(orContainedwhen the plan applied a containment action) + PostMortem CR (timeline, root cause, impact, lessons) + invalidates dedup for the resource- If agentic plan: also generates a reusable Runbook from successful steps
- If
Failedand remaining attempts -> re-analysis: collects failure evidence (collectFailureEvidence), clears AIInsight analysis, returns toAnalyzingstate with failure context - If
Failedand max attempts ->Escalated+ invalidates dedup for the resource
- Each retry triggers AI re-analysis with context from previous failures
- AI receives
previous_failure_contextwith evidence from failed attempts - The prompt instructs: “Do not repeat the same actions. Analyze why they failed and suggest a fundamentally different approach”
- Generates a new auto-generated Runbook when the new analysis differs (the name carries a hash of the analysis)
5. AIInsightReconciler (aiinsight_controller.go)
Watches AIInsight CRs and calls the AnalyzeIssue RPC to populate the analysis.
Flow:
Checks existing analysis
Status.Analysis is already populated (skip if yes).Checks connectivity
Fetches context
Collects K8s context
KubernetesContextBuilder (deployment, pods, events, revisions).Reads failure context
platform.chatcli.io/failure-context (if re-analysis).Builds request
AnalyzeIssueRequest with Issue data + K8s context + failure context.Calls AnalyzeIssue RPC
AnalyzeIssue RPC via ServerClient.Populates status
Status.Analysis, Confidence, Recommendations, SuggestedActions. Clears failure-context annotation after re-analysis completes.k8s_context.go):
Collects real cluster context for Deployments, StatefulSets, DaemonSets, Jobs, CronJobs, and HPAs (max 15000 chars):
- Resource Status: replicas, conditions, containers, images + resources (each type has a dedicated context builder)
- StatefulSet: replicas, update strategy, partition, PodManagementPolicy, VolumeClaimTemplates
- DaemonSet: desired/current/ready/available/unavailable, nodeSelector, tolerations
- Job/CronJob: active/succeeded/failed, completions, parallelism, schedule, lastSuccessful
- HPA: min/max replicas, current/desired, target utilization, current metrics, maxed-out detection
- Pod Details (up to 5 pods, unhealthy first): phase, restart count, container states
- Recent Events (last 15): type, reason, message, count
- Revision History: Last 5 revisions (ReplicaSets) with image diff between revisions
log_analyzer.go):
Advanced application log analysis (beyond the basic 50-line tail):
- Stack Trace Extraction: detects and extracts stack traces from Java (Exception/Caused by), Go (panic/goroutine), Python (Traceback), Node.js (Error at)
- Error Pattern Detection: 24+ critical patterns categorized (crash, connectivity, dns, auth, storage, tls, database, cache, messaging)
- Structured Log Parsing: extracts error/warn entries from JSON logs (fields level, msg, error, timestamp, logger)
- Init Container Logs: analyzes init container logs (reveals startup failures)
- Sidecar Logs: analyzes sidecar logs (istio-proxy, envoy, datadog-agent, etc.)
- Critical Lines: extracts FATAL/PANIC lines with 3 lines of context before/after
- Temporal Window: fetches logs by temporal window (10min before the incident), not just tail
metrics_collector.go):
Prometheus queries for quantitative data during analysis:
- CPU/Memory: usage trends 30min before → during → 15min after the incident
- Request/Error Rate: HTTP requests and 5xx per second
- Latency: P50, P95, P99 histogram percentiles
- HPA Metrics: current vs desired replicas, CPU target
- Network: receive/transmit bytes/s
- Trend Analysis: detects spikes, drops, sustained_high/low with % change calculation
- Enabled via:
PROMETHEUS_URLenv var on the operator
gitops_detector.go):
Detects and integrates with GitOps tools:
- Helm Releases: detects via Secrets type
helm.sh/release.v1, status (deployed/failed/pending-upgrade), chart version, previous revision for rollback - ArgoCD Applications: sync status (Synced/OutOfSync), health (Healthy/Degraded), conditions, last sync result
- Flux Kustomizations: ready status, source ref, conditions, last applied
source_controller.go):
Code-aware diagnostics when SourceRepository CRD is configured:
- Git Correlation: finds commits in the 30min before the incident
- Suspected Commit: identifies the most likely commit (score by temporal proximity + volume of changes)
- Code Extraction: extracts code snippets referenced in stack traces (file path + line number → source code)
- Config Analysis: reads Dockerfile, values.yaml, Chart.yaml for deploy context
- Credentials: Secret keys
token,username+password, orssh-key+known_hosts; HTTPS credentials go through aGIT_ASKPASShelper and are never written to.git/config; SSH verifies host keys againstknown_hosts(spec.sshHostKeyPolicy: acceptNewtrusts the first key when the Secret has none). See Source repositories
cascade_analyzer.go):
Cross-service cascade failure analysis:
- Dependency Graph: discovers dependencies via Services + EndpointSlices
- Temporal Correlation: finds active issues in the same namespace and cross-namespace within a 15-20min window
- Cascade Chain: orders services by detection time (first = root cause)
- Root Cause Service: identifies the service that originated the cascade
blast_radius.go):
Impact prediction before action execution:
- PDB Check: verifies if the action would violate PodDisruptionBudgets
- Quota Check: verifies ResourceQuotas (>90% used = warning)
- Node Capacity: counts pods on node for cordon/drain actions
- Affected Services: discovers which Services would be impacted
- Risk Level: classifies as low/medium/high/critical
6. RemediationReconciler (remediation_controller.go)
Executes the actions defined in a RemediationPlan.
Supported Actions (54 types, plus Custom, which is always rejected):
Deployment / Generic (19 actions + Custom):
7. ServerClient (grpc_client.go)
Shared gRPC client between the WatcherBridge, the AIInsightReconciler and the RemediationReconciler.
Server and Operator Interaction
StreamAlerts and GetAlerts RPCs
The server pushes K8s Watcher alerts overStreamAlerts (server streaming with heartbeats; see Server Mode) and still exposes them for one-shot reads via gRPC:
ObservabilityStore of each MultiWatcher target, filters by namespace if specified, and returns active alerts.
AnalyzeIssue RPC
The server receives the Issue context and calls the LLM for analysis:- Issue context (name, namespace, resource, severity, risk score, description)
- The action catalog (the 54 action types, with their parameters and resource-kind rules)
- Instructions to return structured JSON with
analysis,confidence,recommendations, andactionsfields
- Removes markdown codeblocks (
```json ... ```) - Parses JSON into
analysisResult - Clamps confidence between 0.0 and 1.0
- If parsing fails -> uses raw response as analysis with confidence 0.5
AgenticStep RPC
The server receives the Issue context, history of previous steps, and updated K8s context, and decides the next action:- Role + Issue details: incident context (type, severity, resource)
- Kubernetes context: real cluster state (refreshed at each step via KubernetesContextBuilder)
- Tool definitions: the action catalog + “Observe” (no action, wait for next context)
- Conversation history: each previous step formatted with reasoning -> action -> observation
- Instructions: respond JSON, budget (step N of M), safety rules
resolved=true, the response includes data for PostMortem generation (summary, root_cause, impact, lessons_learned, prevention_actions). When diverges_from_insight is true and divergence_reason is empty, the operator records the step as rejected and does not execute the proposed action.
PostMortem Generation
When any remediation plan completes (standard or agentic, including a containment that leaves the IssueContained), the IssueReconciler automatically generates:
PostMortem CR
Created viageneratePostMortem(), named pm-<issue> in the Issue’s namespace:
- Trending: detection of recurring incidents (count in the last 30 days, related PostMortems)
- Cascade Chain: cascade failure chain if there are correlated cross-service issues
- Git Correlation: suspected commit (SHA, author, changed files, confidence)
- GitOps Context: Helm/ArgoCD/Flux state at the time of the incident
Auto-generated Runbook (Agentic)
Created viagenerateAgenticRunbook():
- Name:
agentic-{signal}-{severity}-{kind}(sanitized) - Steps: only steps with successful actions
- Labels:
auto-generated=true,source=agentic - Uses
CreateOrUpdate(reused for future incidents of the same type)
Operator Prometheus Metrics
The operator exposes Prometheus metrics on its metrics port (8080, plain HTTP, path /metrics; chart serviceMonitor.enabled creates a ServiceMonitor). The pipeline metrics:
chatcli_operator_* metrics, listed on their pages; controller-runtime adds its default reconcile metrics.
Tests
The operator’s unit tests run over the controller-runtime fake client and cover every component on this page:Run Tests
operator/integration) proves what the fake client cannot: CRD schemas and required fields, status subresources, owner references and controllers reacting to each other’s writes. Its scenarios are an Instance provisioning its owned workload only when a credential is configured, one anomaly travelling Anomaly → Issue → AIInsight → RemediationPlan → scaled Deployment → completed plan → resolved Issue → PostMortem, an ApprovalPolicy parking a plan until a human approves, and an IncidentSLA recording a resolution violation. CI runs it with KUBEBUILDER_ASSETS exported and counts it toward coverage.
Ownership Diagram (Garbage Collection)
- Instance is the owner of the namespaced resources it creates (Deployment, Service, ConfigMaps, SA, PVC, watcher Role/RoleBinding); a cross-namespace watcher ClusterRoleBinding is removed by the Instance finalizer instead
- Issue is the owner of AIInsight, RemediationPlan, and PostMortem (cascade delete)
- Anomalies are independent (no owner) to preserve history
AIOps Deployment Checklist
Install Operator via Helm (CRDs + RBAC + Deployment + Dashboard)
Create the Secrets the Instance needs
- the LLM provider keys (for example
ANTHROPIC_API_KEY), referenced byspec.apiKeys.name - a server token (or JWT material): an in-cluster server listens on
0.0.0.0and refuses to run without a credential, and the operator does not create the Deployment without one (AuthenticationConfigured=False) - a TLS Secret (
tls.crt,tls.key, optionallyca.crt) valid for<instance>.<namespace>.svc.cluster.local: the operator always dials the server over TLS 1.3
Create Instance CR
chatcli-watcher (created by the chart).Create one Instance for AIOps per cluster: the pipeline connects to the first ready Instance it finds, cluster-wide.Verify server
kubectl get instances -A — READY must be true; check the conditions AuthenticationConfigured and ServerReachable with kubectl describe instance chatcli -n chatcli.Verify AIOps pipeline
kubectl get anomalies -A— anomalies being detectedkubectl get issues -A— issues being createdkubectl get aiinsights -A— AI analyzing
(Optional) Enable the dashboard and REST API
chatcli-operator-secrets (key api-keys) in chatcli-system, then kubectl -n chatcli-system port-forward svc/chatcli-operator 8090:8090 and open http://localhost:8090.(Optional) Create manual Runbooks
Monitor metrics
serviceMonitor.enabled=true in the chart).