The Decision Engine never acts blindly. Every decision goes through a pipeline of
confidence adjustments, circuit breaker checks, and pattern validation
before any action is executed.
Architecture Overview
Enabling the engine
The engine is off by default: ApprovalPolicies are then the only gate. Turn it on per operator:CHATCLI_OPERATOR_DECISION_ENGINE=true on the operator Deployment. The RemediationReconciler then evaluates every plan that reaches Pending, after the ApprovalPolicies and the cluster tier had their say: a plan an ApprovalPolicy already parked (including an auto rule with autoApproveConditions) is not re-evaluated, a plan an auto rule without conditions let through still is.
Every evaluation leaves its verdict on the plan, whether it ran or waited:
WaitingApproval under the synthetic policy decision-engine: an ApprovalRequest with policyRef: decision-engine, one required approver and a 30-minute timeout, decided like any other request (dashboard, REST API, or the platform.chatcli.io/approve / platform.chatcli.io/reject annotation with the value <approver>:<reason>). No ApprovalPolicy object is needed.
Every channel records who decided and why in
status.decisions: the annotation stores the name it carries, and the REST API and the dashboard store the typed name plus the identity of the API key (<name> (api-key: <identity>)). If the engine returns an error, the plan stays Pending with a Warning Event ApprovalGateUnavailable and is retried with backoff; it never runs ungated.<issue>-insight. When that AIInsight is missing, the base confidence is 0 and the plan ends up manual; an error reading it (other than not found) keeps the plan Pending and retries. The AIInsight comes from the ChatCLI server the operator is connected to; only one Instance per cluster drives AIOps (the operator uses the first ready Instance it finds cluster-wide).
Base Confidence (AIInsight)
The entire process starts with theconfidence field of the AIInsight CR, which is generated by the LLM provider during root cause analysis. This value represents the AI’s certainty about the diagnosis and suggested actions.
High Confidence
0.90 - 1.00 — The AI identified the problem with high precision. Well-known
scenarios like OOMKilled, CrashLoopBackOff with invalid image.
Medium Confidence
0.70 - 0.89 — Probable diagnosis but with uncertainty. Performance
issues, resource pressure, intermittent dependencies.
Low Confidence
0.50 - 0.69 — The AI does not have sufficient certainty. Complex problems
with multiple possible causes.
Very Low Confidence
< 0.50 — Unknown scenario or insufficient data. Always requires
human intervention.
Confidence Adjustment Factors
The base confidence from the AIInsight is adjusted by five factors, then clamped to0.0–1.0. Every number below is what the operator applies.
1. Historical Success Rate
The engine counts RemediationPlans of the last 30 days in the Issue’s namespace whose actions include the plan’s first action type, and needs at least 3 finished ones (Completed, Failed or RolledBack) before it adjusts:
Agentic plans carry no pre-planned actions, so this factor does not apply to them.
2. Pattern Match
When the Pattern Store of the Issue’s namespace holds a pattern with the same signal type, resource kind and severity that has been resolved successfully at least twice, its storedconfidenceBoost is added. The boost is learned per pattern: successCount / (successCount + failureCount) × 0.15, so it is never more than +0.15.
3. Time of Day
The clock is UTC on purpose: the operator has no notion of the team’s local business hours.
4. Simultaneous Active Issues
Non-terminal Issues in the same namespace, the one being evaluated included:5. Incident Severity
Practical Calculation Example
critical would end at 0.88 and be parked as manual: critical never runs unattended.
Decision Thresholds
The final confidence and the severity together decide how much autonomy the plan gets.- Auto-Remediation
- Auto with Notification
- Requires Approval
- Manual
Requirements: confidence ≥ 0.95 and severity
lowThe plan executes immediately. The verdict stays on the plan:A plan that an ApprovalPolicy already parked is not evaluated again. A plan an
auto rule without autoApproveConditions let through still is: the engine can only add caution, never remove a policy’s gate.Circuit Breaker
The circuit breaker blocks every new plan in a namespace when remediations there keep failing, so the platform stops compounding damage.1
Failure window
On each evaluation the engine counts RemediationPlans in the namespace that ended
Failed or RolledBack within the last 1 hour. The time of failure is completedAt, or startedAt when a failure path did not stamp it, or the plan’s creation time.2
Open
3 or more failures in the window open the breaker: the plan is parked as
decision-mode: blocked with the reason Circuit breaker open: 3 remediations failed in last hour, and chatcli_operator_decision_engine_circuit_breaker_state{namespace} reads 1.3
Close
The breaker closes on its own once the failures age out of the hour. There is no manual reset: approve the parked requests you want to run, or fix the cause and wait.
Pattern Store
The Pattern Store is the platform’s pattern learning system. It lets AIOps “remember” how past incidents ended and feed that memory back into the pattern match adjustment.Fingerprint
Each pattern is keyed by a fingerprint of the Issue’s signal type, resource kind and severity, lowercased:ConfigMap Storage
Patterns live in a ConfigMap namedchatcli-pattern-store in the Issue’s namespace (one per namespace that has had remediations), created by the operator on first use. Each data key is a fingerprint and its value is the pattern as JSON:
When patterns are recorded
The RemediationReconciler updates the pattern when a plan reaches a terminal state:
Both paths recompute
confidenceBoost and refresh lastSeenAt.
Confidence Boost Calculation
successCount ≥ 2.
A pattern match is visible only through its effect on the verdict: the final value in the plan’s
platform.chatcli.io/confidence annotation. The operator does not write pattern details to the Issue, the AIInsight or the RemediationPlan.Root Cause Analysis (RCA) Enrichment
Before asking the ChatCLI server for an analysis, the AIInsight controller gathers extra cluster context about the Issue and appends it to the analysis prompt as a text block (capped at 4,000 characters). It feeds the LLM’s diagnosis, and through it the AIInsight’sconfidence; the decision engine does not read it directly, and it is not stored as a structured field on any CR.
The enricher looks back 30 minutes from the Issue’s detectedAt (or its creation time):
The block ends with a fixed list of “possible causes”, in this order and only when the matching signal exists: a recent deployment change, a recent ConfigMap change, each unhealthy Service, related active Issues, or, when none apply,
No obvious external cause detected — may be resource exhaustion or application bug. The list is a heuristic hint for the LLM, not a scored ranking.
Convergence Detector
The Convergence Detector guards the agentic remediation loop. It inspects the plan’sspec.agenticHistory (each step’s action and observation) to stop loops that are stuck, oscillating, failing repeatedly or about to time out, instead of letting them burn the remaining steps.
IsConverged
True when the last 3 steps of the agentic history have the same observation (compared trimmed and lowercased, and not empty): the loop is no longer changing anything, for better or worse.IsOscillating
True when the actions of the last 4 steps alternate between two different action types, A → B → A → B (for exampleScaleDeployment, RestartDeployment, ScaleDeployment, RestartDeployment). Steps without an action break the pattern.
ShouldStop
The RemediationReconciler calls the detector before every agentic step, after the hard limits (max steps, 10-minute timeout):
A stop fails the plan with
status.result set to Agentic loop stopped: <reason> (estimated progress NN%) and increments chatcli_operator_agentic_convergence_stops_total{reason} with converged, oscillating, timeout or failures. The Issue then retries with its next attempt or escalates, exactly as after any failed plan.
EstimateProgress
Estimates agentic loop progress from 0.0 to 1.0; the value only appears in the stop message above.
The result is capped at
1.0.
Complete Decision Flow
Decision Engine Metrics
Next Steps
Multi-Cluster Federation
See how the decision engine operates in multi-cluster environments with policies
per tier.
Chaos Engineering
Validate engine decisions with controlled chaos experiments.
Audit and Compliance
Parked plans, approval decisions and remediation start, success and failure are recorded as AuditEvents.
AIOps Platform
Return to the complete AIOps platform overview.