Why Approval Workflows are Essential
Security
Prevents automatic remediation from causing greater impact than the original problem (e.g., accidental rollback in production)
Compliance
Every request, decision and expiry is kept on the
ApprovalRequest CR and recorded as an AuditEvent.Trust
Teams adopt AIOps more easily when they know that critical actions require human approval.
Flow Overview
The operator does not send a notification when an ApprovalRequest is created. NotificationPolicies are driven by Issue state changes, not by approval requests. Watch
kubectl get approvalrequests -A, the web dashboard, or alert on the metrics below to know a request is waiting.ApprovalPolicy CRD
TheApprovalPolicy (short name ap) defines rules that decide which remediation plans need approval and how that approval is obtained. It only applies to plans in its own namespace: the operator lists the enabled ApprovalPolicies in the namespace of the RemediationPlan (the namespace of the Issue), never across namespaces.
Spec Fields
ApprovalRule
Each rule defines a match + mode pair with specific configurations.ApprovalMatch
Defines which remediations are covered by this rule. The logic is AND between fields and OR within each field. An empty field matches everything, so a rule with an emptymatch matches every plan.
There is no “most restrictive rule wins” merging. Within a policy the first matching rule is used and the rest are ignored, so put the strict rules first. If several enabled ApprovalPolicies exist in the same namespace, the order in which they are evaluated is not guaranteed; keep one policy per namespace, or make their matches disjoint.
Three Approval Modes
- auto
- manual
- quorum
Auto: a matching The confidence comes from the Issue’s AIInsight; a missing AIInsight counts as confidence
auto rule without autoApproveConditions means “this plan does not need approval from the policy”. No ApprovalRequest is created and the plan proceeds (the cluster tier and the decision engine, when enabled, still evaluate it).Because the first match wins, an auto rule also shadows every rule below it for the plans it matches.With autoApproveConditions, the rule parks the plan in an ApprovalRequest, and the approval controller approves it automatically (status.autoApproved: true, decision by auto-policy) only when all conditions hold. Otherwise the request waits for a human decision like a manual rule, and still expires on timeoutMinutes:0, so the conditions are not met and a human decides.ChangeWindowSpec
A change window is set per rule (rules[].changeWindow), not at the policy level. It only affects requests raised by that rule.
There are no blackout dates and no override for critical severity.
The window controls when an approval takes effect, whatever channel the decision came from (annotation, REST API or dashboard):
- Decisions are recorded at any time, and a rejection takes effect immediately.
- The timeout clock only runs while the window is open, so a request raised at night keeps its full
timeoutMinutesfor when the approvers can act on it. - A request that already has the approvals it needs and waits only for the window does not expire. It carries the condition
ChangeWindow=False(reasonOutsideChangeWindow) and turnsApprovedwhen the window opens (re-checked every minute); approvals inside the window setChangeWindow=True(WithinChangeWindow).
ApprovalRequest CRD
TheApprovalRequest (short name ar) is created by the RemediationReconciler when a plan must wait. Its name is always approval-<plan-name>, it lives in the plan’s namespace, and it is owned by the RemediationPlan (deleting the plan deletes it).
Spec Fields
Root
BlastRadiusAssessment
ApprovalEvidence
Status
ApprovalRequest States
A single rejection is sufficient to block the plan, regardless of the number of approvals. A rejected or expired plan counts as a failed remediation attempt: the Issue goes back to
Analyzing for a new plan while attempts remain (which may raise a new ApprovalRequest), and becomes Escalated when the maximum is reached. Deleting the ApprovalRequest before a decision also fails the plan; it never lets it run.Blast Radius Calculator
Two calculations run, both informational: neither one blocks a plan on its own.How It Works
1
Assessment stored in the spec (at request creation)
For a
Deployment target, the calculator counts the pods matching the Deployment selector (falling back to spec.replicas when none are found) and the Services in the namespace whose selector matches the pod template labels. For other kinds it counts the pods in the namespace that have an owner reference with the target’s name, and reports 0 services.2
Risk level from the pod count
3
Adjustment by action type
A
RollbackDeployment action raises low to medium. A Custom action raises low or medium to high. Other action types do not change the level. Ingresses and estimated downtime are not computed.4
Prediction annotations (approval controller)
While the request is
Pending, the approval controller runs the blast radius predictor on the first action of the plan (PodDisruptionBudget, ResourceQuota, node capacity and affected Services checks) and stores the result in two annotations on the ApprovalRequest: platform.chatcli.io/blast-radius (text summary) and platform.chatcli.io/blast-risk-level. This happens once per request. The REST API returns the request’s annotations, so the dashboard’s approval card shows the risk level as a badge.Integration with RemediationReconciler
Complete Flow
When a RemediationPlan isPending, the RemediationReconciler runs three gates in order. The first one that parks the plan wins; the next ones are skipped.
The gate fails closed. If the Issue, the AIInsight or the ApprovalPolicies cannot be read, or the decision engine returns an error, the plan stays Pending, a Warning Event ApprovalGateUnavailable is recorded on it, and the reconcile is retried with backoff; the plan never runs because a gate errored. If creating the ApprovalRequest or annotating the plan fails, the plan also stays Pending and is retried. An ApprovalRequest that already exists with the same name is reused rather than skipped.
A plan whose parent Issue no longer exists fails (“Parent issue not found; approval policies cannot be evaluated without it”) instead of running ungated. A missing AIInsight is not an error: the evidence carries confidence 0, which only makes auto-approval conditions and the decision engine stricter. The cluster-tier gate follows the same rule: a failure to list the ClusterRegistrations keeps the plan Pending and retries, and a CHATCLI_OPERATOR_CLUSTER_NAME that matches no ClusterRegistration parks the plan for manual approval under the cluster-tier policy, with the reason Cluster name "<name>" (CHATCLI_OPERATOR_CLUSTER_NAME) is not registered: .... With the variable unset the tier gate does not apply.
Requests without an ApprovalPolicy
The cluster tier and the decision engine can park a plan even when no ApprovalPolicy exists. Their requests carry a syntheticpolicyRef (cluster-tier or decision-engine) and always follow the same built-in rule: manual mode, one approver, 30-minute timeout, no change window. They are approved or rejected exactly like any other request.
The decision engine also leaves its verdict on the plan as annotations:
platform.chatcli.io/decision-mode, platform.chatcli.io/confidence, platform.chatcli.io/risk and platform.chatcli.io/decision-reason. See the Decision Engine page for the thresholds.
Control Annotation
When a plan is parked, the reconciler setsplatform.chatcli.io/approval-pending on the RemediationPlan to the name of the ApprovalRequest:
WaitingApproval state plus the ApprovalRequest status, which the reconciler always reads directly, so deleting the annotation does not bypass approval. On approval or rejection the approval controller removes it; on rejection it also writes platform.chatcli.io/rejection-reason on the plan.
A request whose ApprovalPolicy (or rule) was deleted in the meantime is still evaluated, with its own requiredApprovers and timeoutMinutes, no change window and no automatic approval: it still needs a human and still expires.
How to Approve
Via kubectl
The recommended way to decide is an annotation on the ApprovalRequest:status.decisions, then removes the annotation (the decision is stored first, so it is never lost), and evaluates the rule: quorum, change window, timeout. A second decision from the same approver is ignored. For a quorum, each approver adds the annotation in turn after the previous one was consumed; if two people annotate before the controller runs, the second needs --overwrite and replaces the first.
<approver> is whatever text is written: the operator does not check it against the identity of the Kubernetes user. Restrict update/patch on approvalrequests with RBAC to the people allowed to approve.Via REST API
The operator REST API (port8090) exposes the requests. Authentication is the X-API-Key header; listing needs the viewer role, approve/reject need operator. Pass ?namespace= to target a namespace; without it the request is looked up by name across all namespaces.
Pending request, exactly like an annotation: an entry in status.decisions with the approver, the reason and a timestamp. It never sets the state itself; the ApprovalReconciler evaluates the decisions against the rule (requiredApprovers, change window) and moves the request to Approved or Rejected, so the response usually still shows Pending and the new entry in decisions.
approveris required in the body (400when empty). The approver recorded is<typed name> (api-key: <identity>), where the identity is thenameof the API key entry, else itsdescription, else akey-<hash>fingerprint (dev-modein dev mode). See Fail-Closed Authentication.- A quorum counts distinct API keys: two calls with the same key count once, whatever names are typed. Give each approver their own key.
409when the request is no longerPending, or when the same key already decided on it.404when the request does not exist.
resource is the Issue name, action the first requested action and reason the policy name. approvedBy, rejectedBy and decisionReason are derived from decisions, and decidedAt is set once the request is finished.
The web dashboard uses the same endpoints: it asks for your name (remembered in the browser) and shows the quorum progress on requests that need more than one approver.
Via Slack
There is no interactive Slack approval. The operator has no Slack callback endpoint and does not post approval requests to any channel.Complete YAML Examples
No Approval for Low-Risk Actions in Staging
Quorum of 2 Approvers for Production
Change Window Weekdays 9-18 UTC
Rollback Protection in a Critical Namespace
payments, auth, billing, …).
Auditing and Compliance
Annotation-based decisions are recorded in theApprovalRequest status:
kubectl get ar columns are Issue, Plan, State, Rule, Age.
The RemediationReconciler also writes AuditEvent CRs for the lifecycle: approval_requested when a plan is parked, and approval_approved, approval_rejected or approval_expired when it acts on the outcome (correlationId = Issue name). A human approval or rejection names the approvers from status.decisions as the event’s actor (actor.type: user, plus an approvers detail); an auto-approval or an expiry is attributed to the ApprovalReconciler. See Audit and Compliance.
ApprovalRequests are owned by their RemediationPlan and are deleted with it, so export them periodically if you need long-term records:
totalApproved, totalRejected, totalExpired and totalAutoApproved (not kept for synthetic requests). Each finished request is counted once (the request is marked platform.chatcli.io/policy-counted), including across operator restarts.
Prometheus Metrics
The approval workflow system exposes these metrics on the operator metrics endpoint (port8080):
There is no gauge of pending requests; count them with
kubectl get ar -A or the REST API (?state=Pending).
Recommended Prometheus alerts:
Next Steps
Notifications and Escalation
Multi-channel notification system and escalation policies
SLOs and SLAs
Service Level Objectives management with burn rate alerting
AIOps Platform
Deep-dive into the complete AIOps architecture
K8s Operator
Operator configuration and CRDs