Overview
The ChatCLI AIOps platform manages incidents (Issue CRs, short name iss) through a state machine with 7 states, and remediation plans (RemediationPlan CRs, short name rp) through a state machine with 7 states. Understanding this lifecycle is essential for operators who need to intervene when automatic remediation fails.
Incident States
State Machine Flow
Detection Phase (Detected)
When the watcher bridge turns a server alert into an Anomaly CR, the Anomaly controller correlates it:
- Existing incident — if a non-terminal Issue (anything except
Resolved,Escalated,Failed) already exists for the same resource (kind + name + namespace), the anomaly is attached to it and its risk score is raised if the new score (which counts the anomaly just attached) is higher - Resolution cooldown — if an Issue for the same resource was
ResolvedwithinresolutionCooldownMinutes, the anomaly is suppressed - Noise reduction — repetitive, seasonal, flapping or high alert-fatigue (score > 80) anomalies are suppressed
- Signal scoring — each signal type has a weight (
oom_kill=40,error_rate=30,pod_restart=25,deploy_failing=25,latency=20,pod_not_ready=20,cpu_high=15,memory_high=15, any other signal=10); the risk score sums all uncorrelated anomalies for the resource in the last 10 minutes, capped at 100 - Severity determination —
oom_killis alwayscritical; otherwise Critical (risk ≥ 80), High (≥ 60), Medium (≥ 30), Low (below 30) - Incident ID — format
INC-YYYYMMDD-NNN, stored in the labelplatform.chatcli.io/inc-id. The Issue name is<resource>-<signal>-<unix-timestamp>(e.g.payment-api-oom-kill-1773900000); the REST API andkubectladdress Issues by that name, not by the INC ID - Max remediation attempts — set on the first reconcile from
aiops.maxRemediationAttempts(default 5). If a runbook is later selected, itsmaxAttemptsreplaces this value (see the tip below)
Escalated counts as closed for correlation: while an Issue is Escalated, new anomalies for the same resource open a new Issue instead of attaching to the escalated one.Analysis Phase (Analyzing)
The system creates an AIInsight CR (<issue>-insight) for AI-powered analysis. At detection, ALL matching runbooks are injected into the AIInsight (annotations platform.chatcli.io/candidate-runbooks and platform.chatcli.io/runbook-context) for validation.
- Runbook candidate discovery (tiered):
- Tier 1: runbooks matching SignalType + Severity + ResourceKind
- Tier 2: runbooks matching Severity + ResourceKind with a different signal
- Multiple runbooks can exist per trigger (different root causes produce different runbooks)
- AI validates candidates: the LLM receives all candidate runbooks and evaluates each against the current root cause analysis:
RUNBOOK_APPROVED: <name>→ uses that runbook (fast path); if the name does not match any candidate, the first candidate is usedRUNBOOK_REJECTED→ skips all candidates, uses AI suggestions or agentic mode- Neither → uses the first candidate as default (backward compatibility)
- If no runbook was selected and the AI suggested actions → generates a new runbook from those actions and uses it
- If no runbook and no AI actions → enters Agentic Mode (AI-driven step-by-step)
- Creates the RemediationPlan
<issue>-plan-<attempt>and transitions toRemediating
Remediation Phase (Remediating)
Before executing, the remediation controller checks the approval gates in order: a matching ApprovalPolicy, then the local cluster’s tier (when CHATCLI_OPERATOR_CLUSTER_NAME is set), then the decision engine (when CHATCLI_OPERATOR_DECISION_ENGINE=true). A gated plan waits in WaitingApproval (see Approval Workflow). The gates fail closed: if the Issue, the AIInsight or the policies cannot be read, or the decision engine errors, the plan stays Pending with a Warning Event ApprovalGateUnavailable and is retried; a plan whose Issue is gone fails.
The plan then executes using a ReAct loop (Reason-Act-Observe):
- Pre-flight snapshot of the target workload captured for rollback
- For each action in the plan:
- OBSERVE — from the second action on, checks if the resource is already healthy. If so, stops immediately without executing the remaining actions (early exit)
- ACT — executes the action and records a checkpoint
- If the action fails → automatic rollback to the pre-flight snapshot
- Final health verification (polls every 10 seconds, for up to 90 seconds; on timeout, rollback to the pre-flight snapshot)
- On success →
Resolved(orContained, if a containment action ran) + PostMortem generated - On failure → re-analysis with the failure context (next attempt) or escalation
AdjustResources followed by RollbackDeployment which would undo the fix) and reduces operational impact to the minimum necessary.
Retry Mechanism
When the latest plan endsFailed or RolledBack:
- Attempt < max attempts: the failure evidence of all failed plans is written to the AIInsight (annotation
platform.chatcli.io/failure-context), the analysis is cleared, and the Issue goes back toAnalyzing— potentially selecting a different runbook or strategy - All attempts exhausted: transitions to
Escalated
Every way a plan can end
Failed counts as a failed attempt — including a rejected or expired approval. Rejecting an approval therefore triggers re-analysis and a new plan (and, eventually, escalation); it does not stop the incident.Escalated State — What Operators Must Do
When an incident reachesEscalated, the system has exhausted all automatic options. Here’s what happens and what you need to do:
What the system does automatically:
- Starts the first matching enabled
EscalationPolicy(byseverities; a policy withdefaultPolicy: trueis the fallback) and notifies level 0 — skipped for chaos-induced Issues - Sends the notifications of any
NotificationPolicyrule that matches theEscalatedstate - Advances to the next level when the current level’s
timeoutMinutesexpires, up to the last level (see Notifications) - Records an
issue_escalatedaudit event - Keeps checking the resource every 30 seconds for auto-resolve (below)
-
Acknowledge the incident (records who is on it and stops the escalation):
Acknowledging writes the annotations
aiops.chatcli.io/acknowledged,-atand-by(the-byvalue is the caller’s API-key role, not a person; the request body is ignored). The escalation stops at its current level: no further level and no repeat, and the acknowledgement is recorded on the EscalationPolicy’sstatus.activeEscalationsentry. Snooze (/snooze, body{"duration": "1h"}, a positive duration) holds every notification exceptResolved, and the escalation, untilaiops.chatcli.io/snoozed-until; a held escalation page goes out when the snooze ends and the level’s timer restarts. See Notifications. - Investigate and fix the issue manually
-
Resolve the incident via one of three methods:
Method 1: REST API (recommended for automation/scripts;
operatorrole)Without?namespace=, the first Issue with that name in any namespace is used. The call returns409if the Issue is alreadyResolved. It sets the status and the annotationsaiops.chatcli.io/resolved-by(the role),resolved-atandmanual-resolution: "true". Method 2: Web Dashboard Navigate to the incident detail page and click the “Resolve” button. Enter the resolution note in the prompt (optional). Method 3: Kubernetes Direct (advanced) —statusis a subresource, so the patch must target it:
A manual resolution (REST, dashboard or
kubectl) does not generate a PostMortem and does not clear the watcher bridge dedup cache — identical alerts stay deduplicated until dedupTTLMinutes expires. PostMortems are generated only when a remediation plan completes.Auto-Resolve for Escalated Issues
When an incident reachesEscalated, the system keeps checking the resource every 30 seconds. If the resource recovers, the issue is automatically resolved with the message:
“Auto-resolved: resource recovered while awaiting human intervention”and the annotations
aiops.chatcli.io/resolved-by: auto-resolve and aiops.chatcli.io/auto-resolution: "true". “Recovered” depends on the kind:
Escalated Issues on any other kind (CronJob, Pod, …) never auto-resolve and must be resolved manually.
This handles cases where:
- An operator fixes the issue manually (
kubectl rollout undo, etc.) without using the API - The resource self-heals (e.g., a transient network issue resolves)
- A CI/CD pipeline deploys a fix while the incident is still open
Escalated and Contained) can be disabled via the Instance CRD: spec.aiops.enableAutoResolve: false. The Issue then stays in that state until resolved manually.
Configurable AIOps Parameters
All timing and retry parameters are configurable via the Instance CRDaiops section:
These settings are read from the Instance the watcher bridge uses. The bridge stamps its Instance on every Anomaly (labels
platform.chatcli.io/instance and platform.chatcli.io/instance-namespace) and the Issue carries them on, so the cooldown follows the Instance the alert came from; the other settings come from the first Ready Instance (the one the bridge connects to), or the first Instance when none is Ready.Remediation Plan States
Each incident may have multiple remediation plans (one per attempt, named<issue>-plan-<attempt>):
Agentic Remediation Mode
When no runbook matches and the AI suggested no actions, the system uses AI-driven agentic remediation:- AI proposes an action via the AgenticStep RPC
- Action is executed and the result is observed
- AI analyzes the observation and proposes the next action
- Loop continues until resolved or a guardrail stops it
- Max steps: 10 (configurable via
aiops.agenticMaxSteps) - Max time: 10 minutes per agentic plan (the convergence detector already stops it at 8 minutes)
- Convergence detection (counted in
chatcli_operator_agentic_convergence_stops_total):- Last 3 observations identical → force stop
- Alternating A→B→A→B action pattern → force stop
- Last 5 actions all failed → force stop
Decision Engine Confidence Thresholds
The decision engine is off by default (decisionEngine.enabled: true in the operator chart, i.e. CHATCLI_OPERATOR_DECISION_ENGINE=true). It only evaluates plans that no ApprovalPolicy or cluster tier already parked. When enabled, it determines whether a plan may run automatically, based on the adjusted confidence:
Adjustments to the AIInsight confidence: historical success rate of the first action type (+0.1 / +0.05 / −0.1, with at least 3 plans in 30 days), learned pattern match boost, −0.05 outside 09:00–18:00 UTC, −0.02 per active Issue above 3 (max −0.1), and severity (critical −0.1, high −0.05, low +0.05).
Circuit breaker: if 3+ remediations failed or rolled back in the same namespace in the last hour, the plan is blocked and waits for a human.
Plans that wait are parked on an ApprovalRequest under the synthetic policy
decision-engine: one approver, 30 minutes, then the plan fails as expired.
Rollback Engine
The rollback engine provides safety nets at two levels:- Pre-flight snapshot — captured before ANY action. Automatic rollbacks always restore this snapshot of the target workload.
- Per-action checkpoints — a snapshot recorded before EACH action (for node actions, a node snapshot). They are kept in
status.actionCheckpointsfor audit and the PostMortem timeline; they are not replayed automatically (there is no automatic partial rollback).
- Action execution fails
- Health verification times out (90 seconds)
- Deployment: replicas, container images, container resources (and HPA min/max)
- StatefulSet: replicas, images, resources, partition
- DaemonSet: images, resources, max unavailable
- Job/CronJob: suspend, deadline, backoff limit, parallelism
- Node: schedulable state — supported by the engine, but because automatic rollback restores the workload snapshot, a
CordonNode/DrainNodeis not undone automatically (useUncordonNode)
Remediation Action Types
The platform supports 54 typed remediation actions (plusCustom, which is treated as a no-op requiring manual intervention) across resource kinds:
Deployment and cluster-level (19 actions)
ScaleDeployment, RollbackDeployment, RestartDeployment, PatchConfig, AdjustResources, DeletePod, HelmRollback, ArgoSyncApp, AdjustHPA, RestartStatefulSetPod, CordonNode, UncordonNode, DrainNode, ResizePVC, RotateSecret, ExecDiagnostic, UpdateIngress, PatchNetworkPolicy, ApplyManifest
StatefulSet (9 actions)
ScaleStatefulSet, RestartStatefulSet, RollbackStatefulSet, AdjustStatefulSetResources, DeleteStatefulSetPod, ForceDeleteStatefulSetPod, UpdateStatefulSetStrategy, RecreateStatefulSetPVC, PartitionStatefulSetUpdate
DaemonSet (7 actions)
RestartDaemonSet, RollbackDaemonSet, AdjustDaemonSetResources, DeleteDaemonSetPod, UpdateDaemonSetStrategy, PauseDaemonSetRollout, CordonAndDeleteDaemonSetPod
Job (9 actions)
RetryJob, AdjustJobResources, DeleteFailedJob, SuspendJob, ResumeJob, AdjustJobParallelism, AdjustJobDeadline, AdjustJobBackoffLimit, ForceDeleteJobPods
CronJob (10 actions)
SuspendCronJob, ResumeCronJob, TriggerCronJob, AdjustCronJobResources, AdjustCronJobSchedule, AdjustCronJobDeadline, AdjustCronJobHistory, AdjustCronJobConcurrency, DeleteCronJobActiveJobs, ReplaceCronJobTemplate
Runbook Learning System
Node Failure — Remediation Flow
When a node has problems, the watcher detects the condition and the bridge emits an Anomaly whose resource kind isNode:
DrainNode cordons the node and then evicts its pods through the policy/v1 Eviction API with a 30-second grace period (DaemonSet and mirror pods are skipped), so PodDisruptionBudgets are enforced. An eviction refused with 429 (a PDB) or a server error is retried every 5 seconds until the optional timeout param (a Go duration, default 2m, at most 10m); the action then fails, naming the pods it could not evict. Require approval for node actions anyway. Node context (CPU, memory, pod count, conditions) is included in the AI analysis.
Health verification and auto-resolve understand a
Node target: it is healthy when its Ready condition is True (a cordon leaves it so). Rollback snapshots only cover workload kinds (Deployment, StatefulSet, DaemonSet, Job, CronJob), so a failed node plan cannot be rolled back automatically.How Runbooks Are Named
Runbooks generated from the AI’s suggested actions include a hash (first 6 hex characters of the SHA-256 of the analysis), so different causes produce different runbooks:agentic-{signal}-{severity}-{kind} (no hash — a later agentic success for the same trigger overwrites it) and keep only the steps that did not fail. Both carry the label platform.chatcli.io/auto-generated: "true".
Multi-Runbook Selection
When multiple runbooks match the same trigger (signal + severity + kind), the AI receives ALL candidates and selects the most appropriate one:RUNBOOK_REJECTED; if it suggests actions, a new runbook is created with a unique hash — expanding the library for future incidents.
Runbook Lifecycle
Because
auto-* runbooks are created before they are proven, review the library (kubectl get rb -A -l platform.chatcli.io/auto-generated=true) and delete runbooks that led to failed attempts. Over time, common failure modes are resolved via runbooks (seconds) instead of full AI analysis (minutes).
PostMortem Generation
When a remediation plan completes — the Issue becomesResolved or Contained — a PostMortem CR (pm-<issue>, short name pm, state Open) is generated containing:
- Timeline — chronological events from detection to resolution
- Root cause, summary and impact — from the AI analysis
- Actions executed — complete remediation history
- Lessons learned and prevention actions — AI recommendations
- Git correlation and GitOps context — recent changes that may have caused the issue
- Cascade chain — related incidents across services
- Trending — recurrence of similar incidents
PostMortems with requiresHumanAction
When the parent Issue is Contained, both the Issue and the PostMortem carry typed status fields:
- If a PostMortem with
requiresHumanAction: trueis set toClosedwithout the annotationaiops.chatcli.io/human-action-acknowledgedset to a truthy value (true,True,yes,ack,acknowledged), thePostMortemReconcilerreverts it toOpen(even after a forcedkubectl patch) - When auto-resolve fires (a human restored the replicas), the controller clears both fields on the Issue and sets the condition
RequiresHumanAction: False. The PostMortem keeps its flag until the human action is acknowledged - The REST API exposes the fields as top-level fields of the incident item (
requiresHumanAction,requiredAction, returned under bothspecandstatusof the response), so dashboards render them without fetching the PostMortem
Schema history (v1alpha1) — In 1.122.x these fields lived in
PostMortemSpec and were null at runtime. They now live in PostMortemStatus and IssueStatus. Helm installs re-apply the CRDs automatically (pre-install/pre-upgrade hook, crdUpgrade.enabled: true); with raw manifests, re-apply config/crd/bases/.operator role):
400 if the PostMortem does not require human action; on success it also clears status.requiresHumanAction immediately. In the web dashboard this appears as the “Ack Human Action” button on the PostMortem row when requiresHumanAction=true.
Chaos Engineering Correlation
An Issue created while aChaosExperiment targets the same resource (kind, name and namespace) — while the experiment is Running, or within 2 minutes after it became Completed or Aborted — automatically receives the labels:
platform.chatcli.io/source=chaos-experimentplatform.chatcli.io/chaos-experiment=<experiment-name>
See Chaos Engineering for details on the CR and the controller.
SLA Integration
Each incident severity can have anIncidentSLA:
- Response time — max time from detection until the Issue is seen in
AnalyzingorRemediating - Resolution time — max time from detection to resolution (an Issue that reaches
Escalatedis also checked against it) - Business hours — optionally count only time inside business hours
escalationPolicyRef/notificationPolicyRef— accepted by the CRD but not read by any controller: an SLA breach is recorded (status, metrics, audit) but does not trigger escalation or notifications by itself