Skip to main content
The Decision Engine is the central component that determines when and how the AIOps platform should act autonomously. It combines calculated confidence, historical patterns, root cause enrichment, and convergence detection to make safe decisions in production.
The Decision Engine never acts blindly. Every decision goes through a pipeline of confidence adjustments, circuit breaker checks, and pattern validation before any action is executed.

Architecture Overview

Enabling the engine

The engine is off by default: ApprovalPolicies are then the only gate. Turn it on per operator:
or set CHATCLI_OPERATOR_DECISION_ENGINE=true on the operator Deployment. The RemediationReconciler then evaluates every plan that reaches Pending, after the ApprovalPolicies and the cluster tier had their say: a plan an ApprovalPolicy already parked (including an auto rule with autoApproveConditions) is not re-evaluated, a plan an auto rule without conditions let through still is. Every evaluation leaves its verdict on the plan, whether it ran or waited:
A plan that must wait is parked in WaitingApproval under the synthetic policy decision-engine: an ApprovalRequest with policyRef: decision-engine, one required approver and a 30-minute timeout, decided like any other request (dashboard, REST API, or the platform.chatcli.io/approve / platform.chatcli.io/reject annotation with the value <approver>:<reason>). No ApprovalPolicy object is needed.
Every channel records who decided and why in status.decisions: the annotation stores the name it carries, and the REST API and the dashboard store the typed name plus the identity of the API key (<name> (api-key: <identity>)). If the engine returns an error, the plan stays Pending with a Warning Event ApprovalGateUnavailable and is retried with backoff; it never runs ungated.
The engine reads the base confidence from the AIInsight named <issue>-insight. When that AIInsight is missing, the base confidence is 0 and the plan ends up manual; an error reading it (other than not found) keeps the plan Pending and retries. The AIInsight comes from the ChatCLI server the operator is connected to; only one Instance per cluster drives AIOps (the operator uses the first ready Instance it finds cluster-wide).

Base Confidence (AIInsight)

The entire process starts with the confidence field of the AIInsight CR, which is generated by the LLM provider during root cause analysis. This value represents the AI’s certainty about the diagnosis and suggested actions.

High Confidence

0.90 - 1.00 — The AI identified the problem with high precision. Well-known scenarios like OOMKilled, CrashLoopBackOff with invalid image.

Medium Confidence

0.70 - 0.89 — Probable diagnosis but with uncertainty. Performance issues, resource pressure, intermittent dependencies.

Low Confidence

0.50 - 0.69 — The AI does not have sufficient certainty. Complex problems with multiple possible causes.

Very Low Confidence

< 0.50 — Unknown scenario or insufficient data. Always requires human intervention.

Confidence Adjustment Factors

The base confidence from the AIInsight is adjusted by five factors, then clamped to 0.0–1.0. Every number below is what the operator applies.

1. Historical Success Rate

The engine counts RemediationPlans of the last 30 days in the Issue’s namespace whose actions include the plan’s first action type, and needs at least 3 finished ones (Completed, Failed or RolledBack) before it adjusts: Agentic plans carry no pre-planned actions, so this factor does not apply to them.

2. Pattern Match

When the Pattern Store of the Issue’s namespace holds a pattern with the same signal type, resource kind and severity that has been resolved successfully at least twice, its stored confidenceBoost is added. The boost is learned per pattern: successCount / (successCount + failureCount) × 0.15, so it is never more than +0.15.

3. Time of Day

The clock is UTC on purpose: the operator has no notion of the team’s local business hours.

4. Simultaneous Active Issues

Non-terminal Issues in the same namespace, the one being evaluated included:

5. Incident Severity

Practical Calculation Example

The same Issue at severity critical would end at 0.88 and be parked as manual: critical never runs unattended.

Decision Thresholds

The final confidence and the severity together decide how much autonomy the plan gets.
Requirements: confidence ≥ 0.95 and severity lowThe plan executes immediately. The verdict stays on the plan:
A plan that an ApprovalPolicy already parked is not evaluated again. A plan an auto rule without autoApproveConditions let through still is: the engine can only add caution, never remove a policy’s gate.

Circuit Breaker

The circuit breaker blocks every new plan in a namespace when remediations there keep failing, so the platform stops compounding damage.
1

Failure window

On each evaluation the engine counts RemediationPlans in the namespace that ended Failed or RolledBack within the last 1 hour. The time of failure is completedAt, or startedAt when a failure path did not stamp it, or the plan’s creation time.
2

Open

3 or more failures in the window open the breaker: the plan is parked as decision-mode: blocked with the reason Circuit breaker open: 3 remediations failed in last hour, and chatcli_operator_decision_engine_circuit_breaker_state{namespace} reads 1.
3

Close

The breaker closes on its own once the failures age out of the hour. There is no manual reset: approve the parked requests you want to run, or fix the cause and wait.
The state is derived from the plans on every evaluation, not kept in memory, so an operator restart does not lose it.

Pattern Store

The Pattern Store is the platform’s pattern learning system. It lets AIOps “remember” how past incidents ended and feed that memory back into the pattern match adjustment.

Fingerprint

Each pattern is keyed by a fingerprint of the Issue’s signal type, resource kind and severity, lowercased:
The result is a 24-character hex string. Two Issues share a pattern exactly when those three values match; the resource name and the namespace’s other Issues play no part.

ConfigMap Storage

Patterns live in a ConfigMap named chatcli-pattern-store in the Issue’s namespace (one per namespace that has had remediations), created by the operator on first use. Each data key is a fingerprint and its value is the pattern as JSON:

When patterns are recorded

The RemediationReconciler updates the pattern when a plan reaches a terminal state: Both paths recompute confidenceBoost and refresh lastSeenAt.

Confidence Boost Calculation

The boost is applied only once successCount ≥ 2.
A pattern match is visible only through its effect on the verdict: the final value in the plan’s platform.chatcli.io/confidence annotation. The operator does not write pattern details to the Issue, the AIInsight or the RemediationPlan.

Root Cause Analysis (RCA) Enrichment

Before asking the ChatCLI server for an analysis, the AIInsight controller gathers extra cluster context about the Issue and appends it to the analysis prompt as a text block (capped at 4,000 characters). It feeds the LLM’s diagnosis, and through it the AIInsight’s confidence; the decision engine does not read it directly, and it is not stored as a structured field on any CR. The enricher looks back 30 minutes from the Issue’s detectedAt (or its creation time): The block ends with a fixed list of “possible causes”, in this order and only when the matching signal exists: a recent deployment change, a recent ConfigMap change, each unhealthy Service, related active Issues, or, when none apply, No obvious external cause detected — may be resource exhaustion or application bug. The list is a heuristic hint for the LLM, not a scored ranking.

Convergence Detector

The Convergence Detector guards the agentic remediation loop. It inspects the plan’s spec.agenticHistory (each step’s action and observation) to stop loops that are stuck, oscillating, failing repeatedly or about to time out, instead of letting them burn the remaining steps.

IsConverged

True when the last 3 steps of the agentic history have the same observation (compared trimmed and lowercased, and not empty): the loop is no longer changing anything, for better or worse.

IsOscillating

True when the actions of the last 4 steps alternate between two different action types, A → B → A → B (for example ScaleDeployment, RestartDeployment, ScaleDeployment, RestartDeployment). Steps without an action break the pattern.
Oscillation is a strong signal that the remediation is undoing itself. The loop is stopped at that step and the plan fails; the Issue then follows its normal retry or escalation path (see below).

ShouldStop

The RemediationReconciler calls the detector before every agentic step, after the hard limits (max steps, 10-minute timeout):
A stop fails the plan with status.result set to Agentic loop stopped: <reason> (estimated progress NN%) and increments chatcli_operator_agentic_convergence_stops_total{reason} with converged, oscillating, timeout or failures. The Issue then retries with its next attempt or escalates, exactly as after any failed plan.

EstimateProgress

Estimates agentic loop progress from 0.0 to 1.0; the value only appears in the stop message above.
The result is capped at 1.0.

Complete Decision Flow

Decision Engine Metrics

Next Steps

Multi-Cluster Federation

See how the decision engine operates in multi-cluster environments with policies per tier.

Chaos Engineering

Validate engine decisions with controlled chaos experiments.

Audit and Compliance

Parked plans, approval decisions and remediation start, success and failure are recorded as AuditEvents.

AIOps Platform

Return to the complete AIOps platform overview.