Skip to main content

Overview

The ChatCLI AIOps platform manages incidents (Issue CRs, short name iss) through a state machine with 7 states, and remediation plans (RemediationPlan CRs, short name rp) through a state machine with 7 states. Understanding this lifecycle is essential for operators who need to intervene when automatic remediation fails.

Incident States

Contained state — A plan that completes after executing an action with containment: "true" (for example ScaleDeployment replicas=0 containment=true) does not mark the Issue Resolved. It transitions to Contained, which:
  • Is not terminal: it auto-resolves only when the workload is back with desired replicas > 0 and all replicas ready (checked every 60 seconds). For a DaemonSet, Job or Node the check is the one used for Escalated auto-resolve; other kinds stay Contained until resolved manually
  • Sets status.requiresHumanAction: true and status.requiredAction on the Issue, plus the Contained and RequiresHumanAction conditions
  • Generates a PostMortem with requiresHumanAction: true that cannot stay Closed until a human acknowledges it (see below)
  • Is counted in analytics/summary as containedIssues and as open
Configure a NotificationPolicy rule with states: [Contained] so humans are paged when this happens.

State Machine Flow

Detection Phase (Detected)

When the watcher bridge turns a server alert into an Anomaly CR, the Anomaly controller correlates it:
  1. Existing incident — if a non-terminal Issue (anything except Resolved, Escalated, Failed) already exists for the same resource (kind + name + namespace), the anomaly is attached to it and its risk score is raised if the new score (which counts the anomaly just attached) is higher
  2. Resolution cooldown — if an Issue for the same resource was Resolved within resolutionCooldownMinutes, the anomaly is suppressed
  3. Noise reduction — repetitive, seasonal, flapping or high alert-fatigue (score > 80) anomalies are suppressed
  4. Signal scoring — each signal type has a weight (oom_kill=40, error_rate=30, pod_restart=25, deploy_failing=25, latency=20, pod_not_ready=20, cpu_high=15, memory_high=15, any other signal=10); the risk score sums all uncorrelated anomalies for the resource in the last 10 minutes, capped at 100
  5. Severity determination — oom_kill is always critical; otherwise Critical (risk ≥ 80), High (≥ 60), Medium (≥ 30), Low (below 30)
  6. Incident ID — format INC-YYYYMMDD-NNN, stored in the label platform.chatcli.io/inc-id. The Issue name is <resource>-<signal>-<unix-timestamp> (e.g. payment-api-oom-kill-1773900000); the REST API and kubectl address Issues by that name, not by the INC ID
  7. Max remediation attempts — set on the first reconcile from aiops.maxRemediationAttempts (default 5). If a runbook is later selected, its maxAttempts replaces this value (see the tip below)
Escalated counts as closed for correlation: while an Issue is Escalated, new anomalies for the same resource open a new Issue instead of attaching to the escalated one.
To find an Issue by its incident ID:

Analysis Phase (Analyzing)

The system creates an AIInsight CR (<issue>-insight) for AI-powered analysis. At detection, ALL matching runbooks are injected into the AIInsight (annotations platform.chatcli.io/candidate-runbooks and platform.chatcli.io/runbook-context) for validation.
  1. Runbook candidate discovery (tiered):
    • Tier 1: runbooks matching SignalType + Severity + ResourceKind
    • Tier 2: runbooks matching Severity + ResourceKind with a different signal
    • Multiple runbooks can exist per trigger (different root causes produce different runbooks)
  2. AI validates candidates: the LLM receives all candidate runbooks and evaluates each against the current root cause analysis:
    • RUNBOOK_APPROVED: &lt;name&gt; → uses that runbook (fast path); if the name does not match any candidate, the first candidate is used
    • RUNBOOK_REJECTED → skips all candidates, uses AI suggestions or agentic mode
    • Neither → uses the first candidate as default (backward compatibility)
  3. If no runbook was selected and the AI suggested actions → generates a new runbook from those actions and uses it
  4. If no runbook and no AI actions → enters Agentic Mode (AI-driven step-by-step)
  5. Creates the RemediationPlan <issue>-plan-<attempt> and transitions to Remediating

Remediation Phase (Remediating)

Before executing, the remediation controller checks the approval gates in order: a matching ApprovalPolicy, then the local cluster’s tier (when CHATCLI_OPERATOR_CLUSTER_NAME is set), then the decision engine (when CHATCLI_OPERATOR_DECISION_ENGINE=true). A gated plan waits in WaitingApproval (see Approval Workflow). The gates fail closed: if the Issue, the AIInsight or the policies cannot be read, or the decision engine errors, the plan stays Pending with a Warning Event ApprovalGateUnavailable and is retried; a plan whose Issue is gone fails. The plan then executes using a ReAct loop (Reason-Act-Observe):
  1. Pre-flight snapshot of the target workload captured for rollback
  2. For each action in the plan:
    • OBSERVE — from the second action on, checks if the resource is already healthy. If so, stops immediately without executing the remaining actions (early exit)
    • ACT — executes the action and records a checkpoint
    • If the action fails → automatic rollback to the pre-flight snapshot
  3. Final health verification (polls every 10 seconds, for up to 90 seconds; on timeout, rollback to the pre-flight snapshot)
  4. On success → Resolved (or Contained, if a containment action ran) + PostMortem generated
  5. On failure → re-analysis with the failure context (next attempt) or escalation
This prevents contradictory actions from being executed (e.g., AdjustResources followed by RollbackDeployment which would undo the fix) and reduces operational impact to the minimum necessary.

Retry Mechanism

When the latest plan ends Failed or RolledBack:
  • Attempt < max attempts: the failure evidence of all failed plans is written to the AIInsight (annotation platform.chatcli.io/failure-context), the analysis is cleared, and the Issue goes back to Analyzing — potentially selecting a different runbook or strategy
  • All attempts exhausted: transitions to Escalated
Every way a plan can end Failed counts as a failed attempt — including a rejected or expired approval. Rejecting an approval therefore triggers re-analysis and a new plan (and, eventually, escalation); it does not stop the incident.

Escalated State — What Operators Must Do

When an incident reaches Escalated, the system has exhausted all automatic options. Here’s what happens and what you need to do: What the system does automatically:
  1. Starts the first matching enabled EscalationPolicy (by severities; a policy with defaultPolicy: true is the fallback) and notifies level 0 — skipped for chaos-induced Issues
  2. Sends the notifications of any NotificationPolicy rule that matches the Escalated state
  3. Advances to the next level when the current level’s timeoutMinutes expires, up to the last level (see Notifications)
  4. Records an issue_escalated audit event
  5. Keeps checking the resource every 30 seconds for auto-resolve (below)
What operators must do:
  1. Acknowledge the incident (records who is on it and stops the escalation):
    Acknowledging writes the annotations aiops.chatcli.io/acknowledged, -at and -by (the -by value is the caller’s API-key role, not a person; the request body is ignored). The escalation stops at its current level: no further level and no repeat, and the acknowledgement is recorded on the EscalationPolicy’s status.activeEscalations entry. Snooze (/snooze, body {"duration": "1h"}, a positive duration) holds every notification except Resolved, and the escalation, until aiops.chatcli.io/snoozed-until; a held escalation page goes out when the snooze ends and the level’s timer restarts. See Notifications.
  2. Investigate and fix the issue manually
  3. Resolve the incident via one of three methods: Method 1: REST API (recommended for automation/scripts; operator role)
    Without ?namespace=, the first Issue with that name in any namespace is used. The call returns 409 if the Issue is already Resolved. It sets the status and the annotations aiops.chatcli.io/resolved-by (the role), resolved-at and manual-resolution: "true". Method 2: Web Dashboard Navigate to the incident detail page and click the “Resolve” button. Enter the resolution note in the prompt (optional). Method 3: Kubernetes Direct (advanced) — status is a subresource, so the patch must target it:
A manual resolution (REST, dashboard or kubectl) does not generate a PostMortem and does not clear the watcher bridge dedup cache — identical alerts stay deduplicated until dedupTTLMinutes expires. PostMortems are generated only when a remediation plan completes.

Auto-Resolve for Escalated Issues

When an incident reaches Escalated, the system keeps checking the resource every 30 seconds. If the resource recovers, the issue is automatically resolved with the message:
“Auto-resolved: resource recovered while awaiting human intervention”
and the annotations aiops.chatcli.io/resolved-by: auto-resolve and aiops.chatcli.io/auto-resolution: "true". “Recovered” depends on the kind: Escalated Issues on any other kind (CronJob, Pod, …) never auto-resolve and must be resolved manually. This handles cases where:
  • An operator fixes the issue manually (kubectl rollout undo, etc.) without using the API
  • The resource self-heals (e.g., a transient network issue resolves)
  • A CI/CD pipeline deploys a fix while the incident is still open
Auto-resolve (for both Escalated and Contained) can be disabled via the Instance CRD: spec.aiops.enableAutoResolve: false. The Issue then stays in that state until resolved manually.

Configurable AIOps Parameters

All timing and retry parameters are configurable via the Instance CRD aiops section:
These settings are read from the Instance the watcher bridge uses. The bridge stamps its Instance on every Anomaly (labels platform.chatcli.io/instance and platform.chatcli.io/instance-namespace) and the Issue carries them on, so the cooldown follows the Instance the alert came from; the other settings come from the first Ready Instance (the one the bridge connects to), or the first Instance when none is Ready.
AI auto-generated runbooks (both standard and agentic) get maxAttempts = the Instance’s maxRemediationAttempts. Manually created runbooks via YAML or API use the CRD default (maxAttempts: 3) unless explicitly specified. When a candidate runbook is selected, its maxAttempts replaces the Issue’s max attempts — an incident matched to a manual runbook without maxAttempts escalates after 3 attempts, not 5.

Remediation Plan States

Each incident may have multiple remediation plans (one per attempt, named <issue>-plan-<attempt>):

Agentic Remediation Mode

When no runbook matches and the AI suggested no actions, the system uses AI-driven agentic remediation:
  1. AI proposes an action via the AgenticStep RPC
  2. Action is executed and the result is observed
  3. AI analyzes the observation and proposes the next action
  4. Loop continues until resolved or a guardrail stops it
Safety guardrails (any of them fails the plan):
  • Max steps: 10 (configurable via aiops.agenticMaxSteps)
  • Max time: 10 minutes per agentic plan (the convergence detector already stops it at 8 minutes)
  • Convergence detection (counted in chatcli_operator_agentic_convergence_stops_total):
    • Last 3 observations identical → force stop
    • Alternating A→B→A→B action pattern → force stop
    • Last 5 actions all failed → force stop

Decision Engine Confidence Thresholds

The decision engine is off by default (decisionEngine.enabled: true in the operator chart, i.e. CHATCLI_OPERATOR_DECISION_ENGINE=true). It only evaluates plans that no ApprovalPolicy or cluster tier already parked. When enabled, it determines whether a plan may run automatically, based on the adjusted confidence: Adjustments to the AIInsight confidence: historical success rate of the first action type (+0.1 / +0.05 / −0.1, with at least 3 plans in 30 days), learned pattern match boost, −0.05 outside 09:00–18:00 UTC, −0.02 per active Issue above 3 (max −0.1), and severity (critical −0.1, high −0.05, low +0.05). Circuit breaker: if 3+ remediations failed or rolled back in the same namespace in the last hour, the plan is blocked and waits for a human. Plans that wait are parked on an ApprovalRequest under the synthetic policy decision-engine: one approver, 30 minutes, then the plan fails as expired.

Rollback Engine

The rollback engine provides safety nets at two levels:
  1. Pre-flight snapshot — captured before ANY action. Automatic rollbacks always restore this snapshot of the target workload.
  2. Per-action checkpoints — a snapshot recorded before EACH action (for node actions, a node snapshot). They are kept in status.actionCheckpoints for audit and the PostMortem timeline; they are not replayed automatically (there is no automatic partial rollback).
Automatic rollback triggers:
  • Action execution fails
  • Health verification times out (90 seconds)
What a snapshot restores:
  • Deployment: replicas, container images, container resources (and HPA min/max)
  • StatefulSet: replicas, images, resources, partition
  • DaemonSet: images, resources, max unavailable
  • Job/CronJob: suspend, deadline, backoff limit, parallelism
  • Node: schedulable state — supported by the engine, but because automatic rollback restores the workload snapshot, a CordonNode/DrainNode is not undone automatically (use UncordonNode)

Remediation Action Types

The platform supports 54 typed remediation actions (plus Custom, which is treated as a no-op requiring manual intervention) across resource kinds:

Deployment and cluster-level (19 actions)

ScaleDeployment, RollbackDeployment, RestartDeployment, PatchConfig, AdjustResources, DeletePod, HelmRollback, ArgoSyncApp, AdjustHPA, RestartStatefulSetPod, CordonNode, UncordonNode, DrainNode, ResizePVC, RotateSecret, ExecDiagnostic, UpdateIngress, PatchNetworkPolicy, ApplyManifest

StatefulSet (9 actions)

ScaleStatefulSet, RestartStatefulSet, RollbackStatefulSet, AdjustStatefulSetResources, DeleteStatefulSetPod, ForceDeleteStatefulSetPod, UpdateStatefulSetStrategy, RecreateStatefulSetPVC, PartitionStatefulSetUpdate

DaemonSet (7 actions)

RestartDaemonSet, RollbackDaemonSet, AdjustDaemonSetResources, DeleteDaemonSetPod, UpdateDaemonSetStrategy, PauseDaemonSetRollout, CordonAndDeleteDaemonSetPod

Job (9 actions)

RetryJob, AdjustJobResources, DeleteFailedJob, SuspendJob, ResumeJob, AdjustJobParallelism, AdjustJobDeadline, AdjustJobBackoffLimit, ForceDeleteJobPods

CronJob (10 actions)

SuspendCronJob, ResumeCronJob, TriggerCronJob, AdjustCronJobResources, AdjustCronJobSchedule, AdjustCronJobDeadline, AdjustCronJobHistory, AdjustCronJobConcurrency, DeleteCronJobActiveJobs, ReplaceCronJobTemplate

Runbook Learning System

Node Failure — Remediation Flow

When a node has problems, the watcher detects the condition and the bridge emits an Anomaly whose resource kind is Node:
DrainNode cordons the node and then evicts its pods through the policy/v1 Eviction API with a 30-second grace period (DaemonSet and mirror pods are skipped), so PodDisruptionBudgets are enforced. An eviction refused with 429 (a PDB) or a server error is retried every 5 seconds until the optional timeout param (a Go duration, default 2m, at most 10m); the action then fails, naming the pods it could not evict. Require approval for node actions anyway. Node context (CPU, memory, pod count, conditions) is included in the AI analysis.
Health verification and auto-resolve understand a Node target: it is healthy when its Ready condition is True (a cordon leaves it so). Rollback snapshots only cover workload kinds (Deployment, StatefulSet, DaemonSet, Job, CronJob), so a failed node plan cannot be rolled back automatically.
The platform builds a library of learned strategies over time, reusable for future incidents with the same trigger.

How Runbooks Are Named

Runbooks generated from the AI’s suggested actions include a hash (first 6 hex characters of the SHA-256 of the analysis), so different causes produce different runbooks:
Runbooks learned from a successful agentic plan are named agentic-{signal}-{severity}-{kind} (no hash — a later agentic success for the same trigger overwrites it) and keep only the steps that did not fail. Both carry the label platform.chatcli.io/auto-generated: "true".

Multi-Runbook Selection

When multiple runbooks match the same trigger (signal + severity + kind), the AI receives ALL candidates and selects the most appropriate one:
If none of the candidates match the current root cause, the AI responds with RUNBOOK_REJECTED; if it suggests actions, a new runbook is created with a unique hash — expanding the library for future incidents.

Runbook Lifecycle

Because auto-* runbooks are created before they are proven, review the library (kubectl get rb -A -l platform.chatcli.io/auto-generated=true) and delete runbooks that led to failed attempts. Over time, common failure modes are resolved via runbooks (seconds) instead of full AI analysis (minutes).

PostMortem Generation

When a remediation plan completes — the Issue becomes Resolved or Contained — a PostMortem CR (pm-<issue>, short name pm, state Open) is generated containing:
  • Timeline — chronological events from detection to resolution
  • Root cause, summary and impact — from the AI analysis
  • Actions executed — complete remediation history
  • Lessons learned and prevention actions — AI recommendations
  • Git correlation and GitOps context — recent changes that may have caused the issue
  • Cascade chain — related incidents across services
  • Trending — recurrence of similar incidents
Issues resolved manually or by auto-resolve do not get a PostMortem. PostMortems can be reviewed and closed via the Review PostMortem and Close PostMortem API endpoints.

PostMortems with requiresHumanAction

When the parent Issue is Contained, both the Issue and the PostMortem carry typed status fields:
Guaranteed behaviors:
  • If a PostMortem with requiresHumanAction: true is set to Closed without the annotation aiops.chatcli.io/human-action-acknowledged set to a truthy value (true, True, yes, ack, acknowledged), the PostMortemReconciler reverts it to Open (even after a forced kubectl patch)
  • When auto-resolve fires (a human restored the replicas), the controller clears both fields on the Issue and sets the condition RequiresHumanAction: False. The PostMortem keeps its flag until the human action is acknowledged
  • The REST API exposes the fields as top-level fields of the incident item (requiresHumanAction, requiredAction, returned under both spec and status of the response), so dashboards render them without fetching the PostMortem
Schema history (v1alpha1) — In 1.122.x these fields lived in PostMortemSpec and were null at runtime. They now live in PostMortemStatus and IssueStatus. Helm installs re-apply the CRDs automatically (pre-install/pre-upgrade hook, crdUpgrade.enabled: true); with raw manifests, re-apply config/crd/bases/.
To acknowledge the action and unblock closing (operator role):
The REST call returns 400 if the PostMortem does not require human action; on success it also clears status.requiresHumanAction immediately. In the web dashboard this appears as the “Ack Human Action” button on the PostMortem row when requiresHumanAction=true.

Chaos Engineering Correlation

An Issue created while a ChaosExperiment targets the same resource (kind, name and namespace) — while the experiment is Running, or within 2 minutes after it became Completed or Aborted — automatically receives the labels:
  • platform.chatcli.io/source=chaos-experiment
  • platform.chatcli.io/chaos-experiment=<experiment-name>
These labels change platform behavior: See Chaos Engineering for details on the CR and the controller.

SLA Integration

Each incident severity can have an IncidentSLA:
  • Response time — max time from detection until the Issue is seen in Analyzing or Remediating
  • Resolution time — max time from detection to resolution (an Issue that reaches Escalated is also checked against it)
  • Business hours — optionally count only time inside business hours
  • escalationPolicyRef / notificationPolicyRef — accepted by the CRD but not read by any controller: an SLA breach is recorded (status, metrics, audit) but does not trigger escalation or notifications by itself
See SLOs & SLAs for details.