Skip to main content
The ChatCLI AIOps Platform is an autonomous system that detects problems in Kubernetes, analyzes root causes with AI, and executes automatic remediations — all orchestrated by native Kubernetes CRDs. This page covers the internal architecture in depth. For configuration and usage examples, see K8s Operator.

Platform v2 Components

Notifications

NotificationPolicy & EscalationPolicy

SLO & SLA

ServiceLevelObjective & IncidentSLA

Approvals

ApprovalPolicy & ApprovalRequest

Multi-Cluster

ClusterRegistration & Federation

Audit

AuditEvent (immutable trail)

Chaos Engineering

ChaosExperiment with safety checks

Decision Engine

Opt-in confidence gate and circuit breaker

Capacity & Cost

Capacity forecast, noise reduction, LLM cost per incident

Web Dashboard

Embedded UI on the operator’s port 8090

REST API & Dashboard

The operator also exposes a REST HTTP API on port 8090 (chart value api.port, env CHATCLI_AIOPS_PORT), reached through the Service chatcli-operator in the operator namespace (not the Instance Service). Its 40+ endpoints cover incidents, AI insights, remediations, runbooks, approvals, SLOs, post-mortems, analytics (including LLM cost), clusters, federation, policies and audit. A Web Dashboard is embedded and served at / on the same port.
  • Authentication: header X-API-Key. Keys are read from the Secret chatcli-operator-secrets, key api-keys (fallback: ConfigMap chatcli-operator-config), in the operator namespace, as a YAML list of {key, role, description}. Roles are viewer < operator < admin; any other role string is denied. Changes are picked up within about 30 seconds; an api-keys entry that is not valid YAML keeps the last valid key set in force, and a Secret without the entry falls back to the ConfigMap. Create it yourself, or let the operator chart render it (apiKeys.create: true with apiKeys.entries). With no keys configured every /api/ call returns 401, unless CHATCLI_OPERATOR_DEV_MODE=true (admin without a key, development only).
  • Rate limit: 30 requests per minute per client host without a valid API key, 600 per minute per valid key (429 with Retry-After).
  • CORS: deny-all unless CHATCLI_CORS_ALLOWED_ORIGINS / CHATCLI_CORS_ORIGIN are set. TLS: set CHATCLI_AIOPS_TLS_CERT and CHATCLI_AIOPS_TLS_KEY to serve HTTPS (TLS 1.3).
  • Local preview: make dash-preview in operator/ serves the dashboard over synthetic data, no cluster needed, at http://127.0.0.1:8085 with the API key preview.
Every operator replica serves the REST API and the dashboard on 8090, so replicaCount > 1 works behind the Service without pinning requests to the leader.
For the complete reference, see the API Reference. Four Grafana dashboards (JSON) live in deploy/grafana/. The dashboards-configmap.yaml there holds only the two ServiceMonitors (operator and server, the latter selecting app.kubernetes.io/name: chatcli), not the dashboards: create the ConfigMap by naming each JSON file (see Web Dashboard) or import the JSON files by hand.

Pipeline Overview

Internal Components

1. WatcherBridge (watcher_bridge.go)

The WatcherBridge is the pipeline entry point. It implements the controller-runtime manager.Runnable interface and runs as a manager-managed goroutine. Responsibilities: SHA256 Dedup:
  • No temporal component: A continuous problem (e.g., CrashLoopBackOff) generates only one Anomaly
  • Resource UID: a deleted and recreated workload gets a new UID, so its alerts are not swallowed by the old entry
  • TTL: 30 minutes by default (Instance spec.aiops.dedupTTLMinutes, 5–1440) — expired hashes are pruned automatically
  • Invalidation: When an Issue is resolved, contained or escalated, dedup entries for the affected resource are invalidated, allowing immediate recurrence detection
  • Result: Avoids duplicates during an active problem; detects recurrence after resolution
Server Discovery:
1

Lists Instance CRs in the whole cluster

2

Selects the first Instance with Status.Ready=true

Only one Instance per cluster drives AIOps: the bridge (and the AIInsight and Remediation reconcilers, which share its client) talk to the first ready Instance it finds.
3

Connects over gRPC with TLS 1.3

Target dns:///<name>.<namespace>.svc.cluster.local:<port>, with the credential and CA taken from the Instance spec. TLS is mandatory: the Instance needs spec.server.tls.enabled: true and a certificate valid for that name.
4

Retry

If the connection fails, retries on the next poll cycle (30s). After 3 silent stream attempts the bridge drops the connection and rediscovers the Instance.

2. AnomalyReconciler (anomaly_controller.go)

Watches Anomaly CRs and correlates them into Issues. Flow:
1

Receives Anomaly CR

Newly created Anomaly with Status.Correlated = false.
2

Attaches to an active Issue

If a non-terminal Issue already exists for the same resource (kind, name, namespace), the Anomaly is attached to it; the Issue’s risk score is recalculated from every anomaly the Issue now holds in the correlation window, the one just attached included, and only ever goes up.
3

Suppression checks

Skips the Anomaly when the same resource was resolved within the resolution cooldown (Instance spec.aiops.resolutionCooldownMinutes, default 10; 0 turns the cooldown off) or when the noise reducer flags it (repetitive, flapping, seasonal).
4

Groups anomalies and calculates risk score and severity

Calls CorrelationEngine.FindRelatedAnomalies() for uncorrelated anomalies on the same resource within 10 minutes.
5

Creates the Issue CR

Name <resource>-<signal>-<unix time>, with labels platform.chatcli.io/inc-id (INC-YYYYMMDD-NNN), platform.chatcli.io/resource and platform.chatcli.io/signal.
6

Marks Anomaly as correlated

Sets Correlated = true with reference to the Issue.

3. CorrelationEngine (correlation.go)

Correlation engine that groups anomalies into incidents. Correlation Algorithm:
Risk Scoring (sum of weights, capped at 100): Severity Classification:
Example: a Deployment with pod_restart (25) + memory_high (15) = risk 40 -> Medium. Adding error_rate (30) = risk 70 -> High. An Issue opened by an oom_kill anomaly is Critical whatever the score. Source Mapping: the Issue source mirrors the Anomaly source (watcher, prometheus, events, logs, webhook); an unknown source maps to prometheus.

4. IssueReconciler (issue_controller.go)

Manages the complete lifecycle of an Issue through a state machine. States and Transitions: Contained means the plan silenced the workload (for example ScaleDeployment to 0 with containment=true) without fixing it: the Issue carries status.requiresHumanAction: true and status.requiredAction, and it moves to Resolved only once the workload is restored. Escalated Issues are re-checked every 30 seconds and auto-resolve when the resource is healthy again, unless the Instance sets spec.aiops.enableAutoResolve: false. The health check understands Deployments, StatefulSets, DaemonSets, Jobs (healthy once Complete) and Nodes (healthy when Ready); an Issue on any other kind is not auto-resolved. Failed is terminal. The AIOps settings (spec.aiops) come from the Instance the WatcherBridge uses: the WatcherBridge labels each Anomaly with platform.chatcli.io/instance and platform.chatcli.io/instance-namespace, the Issue inherits them, and without the labels the first Ready Instance applies.
  1. Sets detectedAt and maxRemediationAttempts (default: 5, configurable via Instance aiops.maxRemediationAttempts; read from the Instance the WatcherBridge uses)
  2. Creates AIInsight CR <issue>-insight with owner reference (Issue -> AIInsight), annotated with the candidate Runbooks
  3. Transitions to Analyzing
  4. Runs the federation checks (cascade detection, cross-cluster correlation), best effort
  5. Requeues after 10 seconds
  1. Checks if AIInsight has Analysis populated
  2. Searches for matching manual Runbook (findMatchingRunbook — tiered matching)
  3. If manual Runbook found -> createRemediationPlan() (manual has precedence)
  4. If no manual Runbook but AIInsight has SuggestedActions -> generateRunbookFromAI() -> createRemediationPlan() using the auto-generated Runbook
  5. If none -> createAgenticRemediationPlan() (AgenticMode=true, no pre-defined actions — AI decides each step)
  6. Transitions to Remediating
  • Tier 1: SignalType + Severity + ResourceKind (exact match, preferred)
  • Tier 2: Severity + ResourceKind (fallback when signal doesn’t match)
  • SignalType resolved from: issue.Spec.SignalType -> fallback issue.Labels["platform.chatcli.io/signal"]
  • Runbooks are searched in every namespace, not only the Issue’s: the Issue’s namespace first, then the others, each Runbook once
  • Materializes SuggestedActions from AI as a reusable Runbook CR
  • Name: auto-{signal}-{severity}-{kind}-{hash} (sanitized; hash = first 6 hex chars of the SHA256 of the analysis, so different root causes produce different Runbooks)
  • Labels: platform.chatcli.io/auto-generated=true
  • Trigger: SignalType + Severity + ResourceKind (for future reuse)
  • Uses CreateOrUpdate for idempotency
  1. Finds the most recent RemediationPlan (findLatestRemediationPlan)
  2. If Completed -> Issue Resolved (or Contained when the plan applied a containment action) + PostMortem CR (timeline, root cause, impact, lessons) + invalidates dedup for the resource
    • If agentic plan: also generates a reusable Runbook from successful steps
  3. If Failed and remaining attempts -> re-analysis: collects failure evidence (collectFailureEvidence), clears AIInsight analysis, returns to Analyzing state with failure context
  4. If Failed and max attempts -> Escalated + invalidates dedup for the resource
Retry with Strategy Escalation:
  • Each retry triggers AI re-analysis with context from previous failures
  • AI receives previous_failure_context with evidence from failed attempts
  • The prompt instructs: “Do not repeat the same actions. Analyze why they failed and suggest a fundamentally different approach”
  • Generates a new auto-generated Runbook when the new analysis differs (the name carries a hash of the analysis)
Remediation Priority:

5. AIInsightReconciler (aiinsight_controller.go)

Watches AIInsight CRs and calls the AnalyzeIssue RPC to populate the analysis. Flow:
1

Checks existing analysis

Checks if Status.Analysis is already populated (skip if yes).
2

Checks connectivity

Checks if server is connected (requeue 15s if not).
3

Fetches context

Fetches parent Issue for context.
4

Collects K8s context

Collects K8s context via KubernetesContextBuilder (deployment, pods, events, revisions).
5

Reads failure context

Reads failure context from annotation platform.chatcli.io/failure-context (if re-analysis).
6

Builds request

Builds AnalyzeIssueRequest with Issue data + K8s context + failure context.
7

Calls AnalyzeIssue RPC

Calls AnalyzeIssue RPC via ServerClient.
8

Populates status

Populates Status.Analysis, Confidence, Recommendations, SuggestedActions. Clears failure-context annotation after re-analysis completes.
KubernetesContextBuilder (k8s_context.go): Collects real cluster context for Deployments, StatefulSets, DaemonSets, Jobs, CronJobs, and HPAs (max 15000 chars):
  • Resource Status: replicas, conditions, containers, images + resources (each type has a dedicated context builder)
  • StatefulSet: replicas, update strategy, partition, PodManagementPolicy, VolumeClaimTemplates
  • DaemonSet: desired/current/ready/available/unavailable, nodeSelector, tolerations
  • Job/CronJob: active/succeeded/failed, completions, parallelism, schedule, lastSuccessful
  • HPA: min/max replicas, current/desired, target utilization, current metrics, maxed-out detection
  • Pod Details (up to 5 pods, unhealthy first): phase, restart count, container states
  • Recent Events (last 15): type, reason, message, count
  • Revision History: Last 5 revisions (ReplicaSets) with image diff between revisions
LogAnalyzer (log_analyzer.go): Advanced application log analysis (beyond the basic 50-line tail):
  • Stack Trace Extraction: detects and extracts stack traces from Java (Exception/Caused by), Go (panic/goroutine), Python (Traceback), Node.js (Error at)
  • Error Pattern Detection: 24+ critical patterns categorized (crash, connectivity, dns, auth, storage, tls, database, cache, messaging)
  • Structured Log Parsing: extracts error/warn entries from JSON logs (fields level, msg, error, timestamp, logger)
  • Init Container Logs: analyzes init container logs (reveals startup failures)
  • Sidecar Logs: analyzes sidecar logs (istio-proxy, envoy, datadog-agent, etc.)
  • Critical Lines: extracts FATAL/PANIC lines with 3 lines of context before/after
  • Temporal Window: fetches logs by temporal window (10min before the incident), not just tail
MetricsCollector (metrics_collector.go): Prometheus queries for quantitative data during analysis:
  • CPU/Memory: usage trends 30min before → during → 15min after the incident
  • Request/Error Rate: HTTP requests and 5xx per second
  • Latency: P50, P95, P99 histogram percentiles
  • HPA Metrics: current vs desired replicas, CPU target
  • Network: receive/transmit bytes/s
  • Trend Analysis: detects spikes, drops, sustained_high/low with % change calculation
  • Enabled via: PROMETHEUS_URL env var on the operator
GitOpsDetector (gitops_detector.go): Detects and integrates with GitOps tools:
  • Helm Releases: detects via Secrets type helm.sh/release.v1, status (deployed/failed/pending-upgrade), chart version, previous revision for rollback
  • ArgoCD Applications: sync status (Synced/OutOfSync), health (Healthy/Degraded), conditions, last sync result
  • Flux Kustomizations: ready status, source ref, conditions, last applied
SourceCodeAnalyzer (source_controller.go): Code-aware diagnostics when SourceRepository CRD is configured:
  • Git Correlation: finds commits in the 30min before the incident
  • Suspected Commit: identifies the most likely commit (score by temporal proximity + volume of changes)
  • Code Extraction: extracts code snippets referenced in stack traces (file path + line number → source code)
  • Config Analysis: reads Dockerfile, values.yaml, Chart.yaml for deploy context
  • Credentials: Secret keys token, username + password, or ssh-key + known_hosts; HTTPS credentials go through a GIT_ASKPASS helper and are never written to .git/config; SSH verifies host keys against known_hosts (spec.sshHostKeyPolicy: acceptNew trusts the first key when the Secret has none). See Source repositories
CascadeAnalyzer (cascade_analyzer.go): Cross-service cascade failure analysis:
  • Dependency Graph: discovers dependencies via Services + EndpointSlices
  • Temporal Correlation: finds active issues in the same namespace and cross-namespace within a 15-20min window
  • Cascade Chain: orders services by detection time (first = root cause)
  • Root Cause Service: identifies the service that originated the cascade
BlastRadiusPredictor (blast_radius.go): Impact prediction before action execution:
  • PDB Check: verifies if the action would violate PodDisruptionBudgets
  • Quota Check: verifies ResourceQuotas (>90% used = warning)
  • Node Capacity: counts pods on node for cordon/drain actions
  • Affected Services: discovers which Services would be impacted
  • Risk Level: classifies as low/medium/high/critical
AnalyzeIssueRequest:

6. RemediationReconciler (remediation_controller.go)

Executes the actions defined in a RemediationPlan. Supported Actions (54 types, plus Custom, which is always rejected): Deployment / Generic (19 actions + Custom): StatefulSet (9 actions): DaemonSet (7 actions): Job (9 actions): CronJob (10 actions):
Safety Checks (pre-execution): Scale to 0 replicas blocked (Deployment and StatefulSet) unless the action carries containment=true, which marks a deliberate stop-the-bleeding step and moves the Issue to Contained. AdjustResources limit cannot be less than request (all resource types). DeletePod/DeleteStatefulSetPod refuses if only 1 pod exists. ForceDeleteStatefulSetPod requires explicit pod name. RecreateStatefulSetPVC requires confirm=true. Custom actions are blocked. Blast radius prediction checks PDB violations, resource quotas, and affected services before execution — now generalized for all workload types via getPodTemplateLabels.Automatic Rollback (post-failure): Before any action, a structured ResourceSnapshot captures the complete resource state. For Deployments: replicas, images, CPU/memory, HPA. For StatefulSets: replicas, containers, updateStrategy, partition. For DaemonSets: containers, updateStrategy, maxUnavailable. For Jobs: suspend, parallelism, backoffLimit, activeDeadlineSeconds, containers. For CronJobs: suspend, schedule, concurrencyPolicy, history limits, containers. If an action fails or health verification expires (90s), the RollbackEngine automatically restores the resource to the pre-remediation state. Works for Deployments, StatefulSets, DaemonSets, Jobs, CronJobs, Nodes, and HPAs.
Execution Flow (Standard):
Execution Flow (Agentic):

7. ServerClient (grpc_client.go)

Shared gRPC client between the WatcherBridge, the AIInsightReconciler and the RemediationReconciler.

Server and Operator Interaction

StreamAlerts and GetAlerts RPCs

The server pushes K8s Watcher alerts over StreamAlerts (server streaming with heartbeats; see Server Mode) and still exposes them for one-shot reads via gRPC:
The server handler iterates over the ObservabilityStore of each MultiWatcher target, filters by namespace if specified, and returns active alerts.

AnalyzeIssue RPC

The server receives the Issue context and calls the LLM for analysis:
Structured Prompt: The server builds a prompt that includes:
  1. Issue context (name, namespace, resource, severity, risk score, description)
  2. The action catalog (the 54 action types, with their parameters and resource-kind rules)
  3. Instructions to return structured JSON with analysis, confidence, recommendations, and actions fields
Response Parsing:
  1. Removes markdown codeblocks (```json ... ```)
  2. Parses JSON into analysisResult
  3. Clamps confidence between 0.0 and 1.0
  4. If parsing fails -> uses raw response as analysis with confidence 0.5

AgenticStep RPC

The server receives the Issue context, history of previous steps, and updated K8s context, and decides the next action:
AgenticStep Prompt: The server builds a structured prompt with:
  1. Role + Issue details: incident context (type, severity, resource)
  2. Kubernetes context: real cluster state (refreshed at each step via KubernetesContextBuilder)
  3. Tool definitions: the action catalog + “Observe” (no action, wait for next context)
  4. Conversation history: each previous step formatted with reasoning -> action -> observation
  5. Instructions: respond JSON, budget (step N of M), safety rules
When resolved=true, the response includes data for PostMortem generation (summary, root_cause, impact, lessons_learned, prevention_actions). When diverges_from_insight is true and divergence_reason is empty, the operator records the step as rejected and does not execute the proposed action.

PostMortem Generation

When any remediation plan completes (standard or agentic, including a containment that leaves the Issue Contained), the IssueReconciler automatically generates:

PostMortem CR

Created via generatePostMortem(), named pm-<issue> in the Issue’s namespace: In addition to the fields above, the PostMortem is automatically enriched with:
  • Trending: detection of recurring incidents (count in the last 30 days, related PostMortems)
  • Cascade Chain: cascade failure chain if there are correlated cross-service issues
  • Git Correlation: suspected commit (SHA, author, changed files, confidence)
  • GitOps Context: Helm/ArgoCD/Flux state at the time of the incident
The PostMortem CR is owned by the Issue (cascade delete).

Auto-generated Runbook (Agentic)

Created via generateAgenticRunbook():
  • Name: agentic-{signal}-{severity}-{kind} (sanitized)
  • Steps: only steps with successful actions
  • Labels: auto-generated=true, source=agentic
  • Uses CreateOrUpdate (reused for future incidents of the same type)

Operator Prometheus Metrics

The operator exposes Prometheus metrics on its metrics port (8080, plain HTTP, path /metrics; chart serviceMonitor.enabled creates a ServiceMonitor). The pipeline metrics: Approvals, SLA/SLO, notifications, escalation, federation and chaos experiments have their own chatcli_operator_* metrics, listed on their pages; controller-runtime adds its default reconcile metrics.

Tests

The operator’s unit tests run over the controller-runtime fake client and cover every component on this page:

Run Tests

The integration suite (operator/integration) proves what the fake client cannot: CRD schemas and required fields, status subresources, owner references and controllers reacting to each other’s writes. Its scenarios are an Instance provisioning its owned workload only when a credential is configured, one anomaly travelling Anomaly → Issue → AIInsight → RemediationPlan → scaled Deployment → completed plan → resolved Issue → PostMortem, an ApprovalPolicy parking a plan until a human approves, and an IncidentSLA recording a resolution violation. CI runs it with KUBEBUILDER_ASSETS exported and counts it toward coverage.

Ownership Diagram (Garbage Collection)

  • Instance is the owner of the namespaced resources it creates (Deployment, Service, ConfigMaps, SA, PVC, watcher Role/RoleBinding); a cross-namespace watcher ClusterRoleBinding is removed by the Instance finalizer instead
  • Issue is the owner of AIInsight, RemediationPlan, and PostMortem (cascade delete)
  • Anomalies are independent (no owner) to preserve history

AIOps Deployment Checklist

1

Install Operator via Helm (CRDs + RBAC + Deployment + Dashboard)

2

Create the Secrets the Instance needs

  • the LLM provider keys (for example ANTHROPIC_API_KEY), referenced by spec.apiKeys.name
  • a server token (or JWT material): an in-cluster server listens on 0.0.0.0 and refuses to run without a credential, and the operator does not create the Deployment without one (AuthenticationConfigured=False)
  • a TLS Secret (tls.crt, tls.key, optionally ca.crt) valid for <instance>.<namespace>.svc.cluster.local: the operator always dials the server over TLS 1.3
3

Create Instance CR

A watcher target outside the Instance namespace needs the pre-provisioned ClusterRole chatcli-watcher (created by the chart).Create one Instance for AIOps per cluster: the pipeline connects to the first ready Instance it finds, cluster-wide.
4

Verify server

kubectl get instances -A — READY must be true; check the conditions AuthenticationConfigured and ServerReachable with kubectl describe instance chatcli -n chatcli.
5

Verify AIOps pipeline

  • kubectl get anomalies -A — anomalies being detected
  • kubectl get issues -A — issues being created
  • kubectl get aiinsights -A — AI analyzing
6

(Optional) Enable the dashboard and REST API

Create the Secret chatcli-operator-secrets (key api-keys) in chatcli-system, then kubectl -n chatcli-system port-forward svc/chatcli-operator 8090:8090 and open http://localhost:8090.
7

(Optional) Create manual Runbooks

Create manual Runbooks for specific scenarios.
8

Monitor metrics

Monitor operator metrics via Prometheus (serviceMonitor.enabled=true in the chart).

Next Steps

K8s Operator

Configuration and examples

K8s Watcher

Collection and budget details

Server Mode

StreamAlerts, GetAlerts, AnalyzeIssue, and AgenticStep RPCs

K8s Monitoring

Recipe: K8s Monitoring with AI