Skip to main content
In production environments, not every automatic remediation should be executed without human oversight. The ChatCLI Approval Workflow lets you define policies that control which remediation plans must wait for a human, how many approvals they need, and during which change windows a decision can be applied.

Why Approval Workflows are Essential

Security

Prevents automatic remediation from causing greater impact than the original problem (e.g., accidental rollback in production)

Compliance

Every request, decision and expiry is kept on the ApprovalRequest CR and recorded as an AuditEvent.

Trust

Teams adopt AIOps more easily when they know that critical actions require human approval.
Without approval workflows, an AI that detects a false positive could execute an unnecessary rollback, affecting a healthy deployment. With approval policies, high-impact plans are parked until a human validates the analysis and blast radius.

Flow Overview

The operator does not send a notification when an ApprovalRequest is created. NotificationPolicies are driven by Issue state changes, not by approval requests. Watch kubectl get approvalrequests -A, the web dashboard, or alert on the metrics below to know a request is waiting.

ApprovalPolicy CRD

The ApprovalPolicy (short name ap) defines rules that decide which remediation plans need approval and how that approval is obtained. It only applies to plans in its own namespace: the operator lists the enabled ApprovalPolicies in the namespace of the RemediationPlan (the namespace of the Issue), never across namespaces.

Spec Fields

ApprovalRule

Each rule defines a match + mode pair with specific configurations.

ApprovalMatch

Defines which remediations are covered by this rule. The logic is AND between fields and OR within each field. An empty field matches everything, so a rule with an empty match matches every plan.
There is no “most restrictive rule wins” merging. Within a policy the first matching rule is used and the rest are ignored, so put the strict rules first. If several enabled ApprovalPolicies exist in the same namespace, the order in which they are evaluated is not guaranteed; keep one policy per namespace, or make their matches disjoint.

Three Approval Modes

Auto: a matching auto rule without autoApproveConditions means “this plan does not need approval from the policy”. No ApprovalRequest is created and the plan proceeds (the cluster tier and the decision engine, when enabled, still evaluate it).Because the first match wins, an auto rule also shadows every rule below it for the plans it matches.With autoApproveConditions, the rule parks the plan in an ApprovalRequest, and the approval controller approves it automatically (status.autoApproved: true, decision by auto-policy) only when all conditions hold. Otherwise the request waits for a human decision like a manual rule, and still expires on timeoutMinutes:
The confidence comes from the Issue’s AIInsight; a missing AIInsight counts as confidence 0, so the conditions are not met and a human decides.

ChangeWindowSpec

A change window is set per rule (rules[].changeWindow), not at the policy level. It only affects requests raised by that rule. There are no blackout dates and no override for critical severity. The window controls when an approval takes effect, whatever channel the decision came from (annotation, REST API or dashboard):
  • Decisions are recorded at any time, and a rejection takes effect immediately.
  • The timeout clock only runs while the window is open, so a request raised at night keeps its full timeoutMinutes for when the approvers can act on it.
  • A request that already has the approvals it needs and waits only for the window does not expire. It carries the condition ChangeWindow=False (reason OutsideChangeWindow) and turns Approved when the window opens (re-checked every minute); approvals inside the window set ChangeWindow=True (WithinChangeWindow).
A window that can never open (an unknown timezone, no valid day in allowedDays, or startHour equal to endHour) is logged, blocks approval, and the request expires on wall-clock time.

ApprovalRequest CRD

The ApprovalRequest (short name ar) is created by the RemediationReconciler when a plan must wait. Its name is always approval-<plan-name>, it lives in the plan’s namespace, and it is owned by the RemediationPlan (deleting the plan deletes it).

Spec Fields

Root

BlastRadiusAssessment

ApprovalEvidence

Status

ApprovalRequest States

A single rejection is sufficient to block the plan, regardless of the number of approvals. A rejected or expired plan counts as a failed remediation attempt: the Issue goes back to Analyzing for a new plan while attempts remain (which may raise a new ApprovalRequest), and becomes Escalated when the maximum is reached. Deleting the ApprovalRequest before a decision also fails the plan; it never lets it run.

Blast Radius Calculator

Two calculations run, both informational: neither one blocks a plan on its own.

How It Works

1

Assessment stored in the spec (at request creation)

For a Deployment target, the calculator counts the pods matching the Deployment selector (falling back to spec.replicas when none are found) and the Services in the namespace whose selector matches the pod template labels. For other kinds it counts the pods in the namespace that have an owner reference with the target’s name, and reports 0 services.
2

Risk level from the pod count

3

Adjustment by action type

A RollbackDeployment action raises low to medium. A Custom action raises low or medium to high. Other action types do not change the level. Ingresses and estimated downtime are not computed.
4

Prediction annotations (approval controller)

While the request is Pending, the approval controller runs the blast radius predictor on the first action of the plan (PodDisruptionBudget, ResourceQuota, node capacity and affected Services checks) and stores the result in two annotations on the ApprovalRequest: platform.chatcli.io/blast-radius (text summary) and platform.chatcli.io/blast-risk-level. This happens once per request. The REST API returns the request’s annotations, so the dashboard’s approval card shows the risk level as a badge.

Integration with RemediationReconciler

Complete Flow

When a RemediationPlan is Pending, the RemediationReconciler runs three gates in order. The first one that parks the plan wins; the next ones are skipped. The gate fails closed. If the Issue, the AIInsight or the ApprovalPolicies cannot be read, or the decision engine returns an error, the plan stays Pending, a Warning Event ApprovalGateUnavailable is recorded on it, and the reconcile is retried with backoff; the plan never runs because a gate errored. If creating the ApprovalRequest or annotating the plan fails, the plan also stays Pending and is retried. An ApprovalRequest that already exists with the same name is reused rather than skipped. A plan whose parent Issue no longer exists fails (“Parent issue not found; approval policies cannot be evaluated without it”) instead of running ungated. A missing AIInsight is not an error: the evidence carries confidence 0, which only makes auto-approval conditions and the decision engine stricter. The cluster-tier gate follows the same rule: a failure to list the ClusterRegistrations keeps the plan Pending and retries, and a CHATCLI_OPERATOR_CLUSTER_NAME that matches no ClusterRegistration parks the plan for manual approval under the cluster-tier policy, with the reason Cluster name "<name>" (CHATCLI_OPERATOR_CLUSTER_NAME) is not registered: .... With the variable unset the tier gate does not apply.

Requests without an ApprovalPolicy

The cluster tier and the decision engine can park a plan even when no ApprovalPolicy exists. Their requests carry a synthetic policyRef (cluster-tier or decision-engine) and always follow the same built-in rule: manual mode, one approver, 30-minute timeout, no change window. They are approved or rejected exactly like any other request. The decision engine also leaves its verdict on the plan as annotations: platform.chatcli.io/decision-mode, platform.chatcli.io/confidence, platform.chatcli.io/risk and platform.chatcli.io/decision-reason. See the Decision Engine page for the thresholds.

Control Annotation

When a plan is parked, the reconciler sets platform.chatcli.io/approval-pending on the RemediationPlan to the name of the ApprovalRequest:
The annotation is informational. The gate is the plan’s WaitingApproval state plus the ApprovalRequest status, which the reconciler always reads directly, so deleting the annotation does not bypass approval. On approval or rejection the approval controller removes it; on rejection it also writes platform.chatcli.io/rejection-reason on the plan. A request whose ApprovalPolicy (or rule) was deleted in the meantime is still evaluated, with its own requiredApprovers and timeoutMinutes, no change window and no automatic approval: it still needs a human and still expires.

How to Approve

Via kubectl

The recommended way to decide is an annotation on the ApprovalRequest:
Annotation format:
The approval controller (it polls pending requests every 15 seconds) records the decision in status.decisions, then removes the annotation (the decision is stored first, so it is never lost), and evaluates the rule: quorum, change window, timeout. A second decision from the same approver is ignored. For a quorum, each approver adds the annotation in turn after the previous one was consumed; if two people annotate before the controller runs, the second needs --overwrite and replaces the first.
<approver> is whatever text is written: the operator does not check it against the identity of the Kubernetes user. Restrict update/patch on approvalrequests with RBAC to the people allowed to approve.

Via REST API

The operator REST API (port 8090) exposes the requests. Authentication is the X-API-Key header; listing needs the viewer role, approve/reject need operator. Pass ?namespace= to target a namespace; without it the request is looked up by name across all namespaces.
The call records one decision on a Pending request, exactly like an annotation: an entry in status.decisions with the approver, the reason and a timestamp. It never sets the state itself; the ApprovalReconciler evaluates the decisions against the rule (requiredApprovers, change window) and moves the request to Approved or Rejected, so the response usually still shows Pending and the new entry in decisions.
  • approver is required in the body (400 when empty). The approver recorded is <typed name> (api-key: <identity>), where the identity is the name of the API key entry, else its description, else a key-<hash> fingerprint (dev-mode in dev mode). See Fail-Closed Authentication.
  • A quorum counts distinct API keys: two calls with the same key count once, whatever names are typed. Give each approver their own key.
  • 409 when the request is no longer Pending, or when the same key already decided on it. 404 when the request does not exist.
API response (example):
In this representation resource is the Issue name, action the first requested action and reason the policy name. approvedBy, rejectedBy and decisionReason are derived from decisions, and decidedAt is set once the request is finished. The web dashboard uses the same endpoints: it asks for your name (remembered in the browser) and shows the quorum progress on requests that need more than one approver.

Via Slack

There is no interactive Slack approval. The operator has no Slack callback endpoint and does not post approval requests to any channel.

Complete YAML Examples

No Approval for Low-Risk Actions in Staging

A plan that contains a rollback and a restart matches the first rule and waits for approval. Plans matching no rule are not gated by this policy.

Quorum of 2 Approvers for Production

Change Window Weekdays 9-18 UTC

There is no override for critical incidents: a critical plan raised at 3am waits until 9am like any other. If critical incidents must be handled at night, put a rule without changeWindow for severities: [critical] above the windowed rule.

Rollback Protection in a Critical Namespace

Because a policy only covers its own namespace, create one per namespace you want to protect (payments, auth, billing, …).

Auditing and Compliance

Annotation-based decisions are recorded in the ApprovalRequest status:
The default kubectl get ar columns are Issue, Plan, State, Rule, Age. The RemediationReconciler also writes AuditEvent CRs for the lifecycle: approval_requested when a plan is parked, and approval_approved, approval_rejected or approval_expired when it acts on the outcome (correlationId = Issue name). A human approval or rejection names the approvers from status.decisions as the event’s actor (actor.type: user, plus an approvers detail); an auto-approval or an expiry is attributed to the ApprovalReconciler. See Audit and Compliance. ApprovalRequests are owned by their RemediationPlan and are deleted with it, so export them periodically if you need long-term records:
The ApprovalPolicy status keeps totalApproved, totalRejected, totalExpired and totalAutoApproved (not kept for synthetic requests). Each finished request is counted once (the request is marked platform.chatcli.io/policy-counted), including across operator restarts.

Prometheus Metrics

The approval workflow system exposes these metrics on the operator metrics endpoint (port 8080): There is no gauge of pending requests; count them with kubectl get ar -A or the REST API (?state=Pending). Recommended Prometheus alerts:

Next Steps

Notifications and Escalation

Multi-channel notification system and escalation policies

SLOs and SLAs

Service Level Objectives management with burn rate alerting

AIOps Platform

Deep-dive into the complete AIOps architecture

K8s Operator

Operator configuration and CRDs