NotificationReconciler), which watches Issue resources. NotificationPolicy and EscalationPolicy have no controller of their own: the notification controller reads them whenever an Issue changes.
Overview
What triggers a notification
The only trigger is a change ofstatus.state on an Issue. The controller stores the last state it handled in the Issue annotation platform.chatcli.io/last-notified-state, and it evaluates the policies once for each new state: Detected, Analyzing, Remediating, Contained, Resolved, Escalated, Failed.
Other events reach a channel only if they turn into an Issue state change:
Issues created by a
ChaosExperiment (label platform.chatcli.io/source: chaos-experiment) still produce normal state-change notifications. They never start an escalation, so a chaos drill never pages anyone through an EscalationPolicy.NotificationPolicy CRD
ANotificationPolicy (short name np) declares a set of named channels, a list of rules that pick channels by name, throttling, and optional message templates.
How policies are applied
- Every enabled policy in every namespace is evaluated for every Issue in the cluster. The namespace of the policy does not limit its scope. Use the rule’s
namespacesfilter for that. - Inside a policy, every matching rule sends to its channels. The same channel can receive the same Issue twice if two rules match, but the dedup window (below) usually suppresses the second send.
- A rule that names a channel missing from
spec.channelsis skipped with the log linechannel not found in policy.
Spec Fields
NotificationPolicySpec
NotificationChannel
NotificationRule
The filters sit directly on the rule (there is nomatch block). An omitted or empty filter matches everything. The logic is AND between filters and OR within a filter.
ThrottleConfig
Templates
templates replaces the message body (the title and email subject stay the same). The key is chosen from the new state:
The key
remediation_completed is accepted but never used. Templates are not Go text/template: they are plain placeholder replacement, and only these exact tokens are replaced: {{.Name}} (Issue name), {{.Namespace}}, {{.Severity}}, {{.State}}, {{.Resource}} (Kind/name), {{.Description}}, {{.Source}}, {{.SignalType}}, {{.RiskScore}}. Anything else (conditionals, functions, other fields) is sent literally.
Without a template, the body is the Issue’s spec.description, or Issue <name> on <Kind>/<name> transitioned to <State>. when the description is empty.
Message content
Every channel receives the same message:- Title:
<severity emoji> [<SEVERITY>] <namespace>/<resource name> — <State>, for example🔴 [CRITICAL] production/api-gateway — Detected. - Body: the template or description (see above).
- Fields:
Source,SignalType,RiskScore, plusCorrelationID,RemediationAttempts(n/max) andResolutionwhen they are set. - Color by severity: critical
#FF0000, high#FF8C00, medium#FFD700, low#00CC00.
Status
Escalation notifications are not counted in this status. Every delivery attempt, including escalation sends, is also recorded as an
AuditEvent with eventType: notification_sent (see Audit and Compliance).
Notification Channels
1. Slack
Posts to a Slack Incoming Webhook using Block Kit.Full Slack configuration
Full Slack configuration
No other keys are read (there is no mention or icon option). To mention a group, put the mention in a template, for example
issue_created: "<!subteam^S0123ABC> {{.Description}}".Payload sent:SignalType, RiskScore, …) come from a map, so their order can change between messages.2. PagerDuty
Sends events to the PagerDuty Events API v2 (https://events.pagerduty.com/v2/enqueue; the endpoint is fixed).
Full PagerDuty configuration
Full PagerDuty configuration
No other keys are read. The severity mapping and the dedup key are fixed (the
severity_map key mentioned in the CRD field description is not implemented):Deduplication:
dedup_key is always chatcli-<resource namespace>-<issue name>, so every notified state of one Issue updates the same PagerDuty alert.Payload sent:Resolved state, the event uses event_action: resolve with the same dedup_key. This happens whenever a rule sends Resolved to this channel: resolves to PagerDuty are never throttled.3. OpsGenie
Creates alerts through the OpsGenie Alert API (https://api.opsgenie.com/v2/alerts; the endpoint is fixed, so accounts on the EU instance are not supported).
Full OpsGenie configuration
Full OpsGenie configuration
Priority mapping (fixed):
critical → P1, high → P2, medium → P3, low → P4.Responders:alias is chatcli-<resource namespace>-<issue name>, source is ChatCLI AIOps, entity is the resource, and details has resource, namespace, severity, state and issue.Automatic close: for the Resolved state, the channel closes the alert by alias instead of creating one (same conditions as PagerDuty).4. Email
Sends an HTML email over SMTP.Full Email configuration
Full Email configuration
There are no
cc, bcc, subject or HTML template options. The subject is always [<SEVERITY>] <title> and the body is a fixed HTML layout with the fields table.TLS behavior:- Implicit TLS (SMTPS) is used on port
465, or on any port withsmtp_tls: implicit: the connection is encrypted from the first byte. - Otherwise the connection starts in plain text and is upgraded with STARTTLS when the server advertises it.
smtp_tls: starttlsforces this mode even on port465. - If the connection is not encrypted and
smtp_useris set, authentication fails (Go refuses to sendPLAINcredentials over an unencrypted connection, except tolocalhost). - The whole conversation, dial included, is bounded by
smtp_timeout(default30s), so an unreachable or silent server fails the send instead of holding the reconcile.
5. Webhook
Sends the message as JSON to any HTTP endpoint, optionally signed with HMAC-SHA256.Full Webhook configuration
Full Webhook configuration
Every request has
Content-Type: application/json and User-Agent: ChatCLI-AIOps/1.0. The timeout is 30 seconds and there are no retries.HMAC-SHA256 signing:When secret is set, the request carries the header X-Signature-256 with the HMAC-SHA256 of the raw body:6. Microsoft Teams
Posts an Adaptive Card (version 1.4) to a Teams incoming webhook URL.Full Microsoft Teams configuration
Full Microsoft Teams configuration
No other keys are read.Generated card:
- A large, bold title (the message title)
- The body text
- A FactSet with
Severity,Resource,Namespace,State,Issueand the extra fields - A footer with the generation time
themeColor with the severity color without # (FF0000, FF8C00, FFD700, 00CC00).EscalationPolicy CRD
AnEscalationPolicy (short name ep) is an ordered chain of levels. It is used when an Issue enters the Escalated state, which the Issue controller sets when automatic remediation gives up (for example, all remediation attempts failed).
Spec Fields
EscalationLevel
EscalationTarget
How Escalation Works
- Start. When an Issue changes to
Escalated(and it is not chaos-induced), the controller picks a policy: the first enabled policy, from any namespace, whoseseveritiescontains the Issue’s severity, or that has noseverities. A policy withdefaultPolicy: trueis used only when no policy matched. If nothing is found, no escalation happens (log lineno escalation policy found for issue). - Level 1 is notified right away. The Issue gets the annotations below and the policy status gets an
activeEscalationsentry. - Advance. When
timeoutMinutesof the current level has passed, the next level is notified with the title<emoji> ESCALATION [<SEVERITY>] <issue> — Level <n>: <level name>. The new level and its start time are saved on the Issue, so the chain moves forward level by level (L1 → L2 → L3) and each level is notified once, plus its repeats (repeatIntervalMinutes). - Last level. The chain stops there. The last level is re-sent only if it has a
repeatIntervalMinutes. - Stop. An acknowledgement freezes the chain at its current level: no further level and no repeat (see below). The escalation ends when the Issue reaches
Resolved: the escalation annotations are removed and the Issue’s entry leavesstatus.activeEscalations.
notifyChannels (Secrets are read from that policy’s namespace). They do not go through the rules, the throttle or the templates.
Issue annotations used for tracking:
status.activeEscalations entry carries issueName, currentLevel (0-based), escalatedAt and, once the Issue is acknowledged, acknowledgedAt and acknowledgedBy. The entry is removed when the Issue resolves; status.totalEscalations counts every escalation started.
Acknowledgement and stopping an escalation
Both actions go through the REST API (operator role,X-API-Key header) or the dashboard:
- Acknowledge.
POST /api/v1/incidents/{name}/acknowledgeadds the annotationsaiops.chatcli.io/acknowledged,aiops.chatcli.io/acknowledged-atandaiops.chatcli.io/acknowledged-by(the caller’s role). The escalation stops at its current level: no further level is reached and no repeat is sent. The acknowledgement is stamped on the policy’sstatus.activeEscalationsentry (acknowledgedAt,acknowledgedBy). An Issue acknowledged before it reachesEscalatednever starts an escalation. State-change notifications keep flowing. - Snooze.
POST /api/v1/incidents/{name}/snoozewith a body such as{"duration": "30m"}recordsaiops.chatcli.io/snoozed-untilandaiops.chatcli.io/snoozed-by. A duration of zero or less is rejected with400. Until the snooze ends, the Issue sends no notification exceptResolved(a state change that happens during the snooze is not sent later), and its escalation holds its level: no advance and no repeat. A level message that falls inside the snooze is held (escalation-pending-notify) and sent when the snooze ends, and the level’s timer restarts at that moment. - Acknowledging in PagerDuty or OpsGenie does nothing on the cluster side: there is no inbound webhook.
POST /api/v1/incidents/{name}/resolve (operator role) or from the dashboard. That also sends the Resolved notifications that resolve the PagerDuty event and close the OpsGenie alert.
Complete Examples
Notification Policy: Slack + PagerDuty
30s window lets the Slack channels see each transition of a fast recovery. The Resolved event that resolves the PagerDuty incident is never throttled, whatever the window.
Escalation Policy with two levels
email-leadership must be a channel defined in some NotificationPolicy. After the second level is reached, nothing more is sent until the Issue is resolved; add repeatIntervalMinutes to the second level to keep reminding until someone acknowledges.
SLO violation alerts
SLO burn-rate and budget-exhaustion alerts arrive as Issues withsignalType: slo_violation and source: watcher:
Troubleshooting
Notifications are not being sent
Notifications are not being sent
Diagnostic checklist:
- Check that the policy exists and is enabled (policies in any namespace apply):
- Check the policy status for delivery errors:
- Check the operator logs (
rule matched,notification sent,failed to send notification,channel not found in policy,failed to resolve channel config):
- Compare the Issue with the rule filters. Remember that
namespacesis compared withspec.resource.namespace:
-
Check whether the throttle dropped it (log line
notification throttled). Another state of the same Issue may have been sent to that channel withindeduplicationWindow. -
If
platform.chatcli.io/last-notified-statealready equals the current state, that state was already handled and is not evaluated again.
The policy is rejected by kubectl apply
The policy is rejected by kubectl apply
- Every
configvalue must be a string: quote numbers (smtp_port: "587") and booleans (tls_skip_verify: "true"), and write lists as comma-separated strings. - Every channel needs
name,typeandconfig(useconfig: {}when the values come fromsecretRef). - Rule filters sit directly on the rule; a
match:block is not part of the schema.
Slack returns 404 or invalid_payload error
Slack returns 404 or invalid_payload error
- Confirm that the
webhook_urlis correct and the Slack app is still installed in the workspace - Test the webhook manually:
PagerDuty does not create or resolve incidents
PagerDuty does not create or resolve incidents
- Confirm that the
routing_keyis an Events API v2 Integration Key (not a REST API key) - Confirm that the service in PagerDuty is active, and check the payload in the PagerDuty Event Debugger
- If incidents are never resolved, make sure a rule sends
Resolvedto the channel (resolves to PagerDuty are never throttled)
Emails are not arriving
Emails are not arriving
- Test SMTP connectivity from inside the cluster:
- On port
465the channel uses implicit TLS; on587or25it upgrades with STARTTLS. Setsmtp_tlsif your server uses a non-standard port. - A send that fails with a timeout hit
smtp_timeout(default30s) - Confirm that the Secret has the keys
smtp_userandsmtp_password(notusername/password) - Check the recipients’ spam folder
Escalation does not start or does not advance
Escalation does not start or does not advance
- Escalation starts only when the Issue enters
Escalated. Checkkubectl get issue <name> -o jsonpath='{.status.state}'. - Chaos-induced Issues never escalate.
- Check the annotations
platform.chatcli.io/escalation-level,escalation-timeandescalation-policy. - Make sure each level has
notifyChannelswhose names exist in an enabled NotificationPolicy. - Look for
escalation initiated,escalation advancedandno escalation policy found for issuein the operator logs. - An acknowledged Issue does not advance, and a snoozed one holds its level until
aiops.chatcli.io/snoozed-until.
Webhook returns signature error
Webhook returns signature error
- Read the signature from the
X-Signature-256header - Confirm that the
secretin the policy’s Secret is the same one the receiver uses - Compute the HMAC over the raw body, before parsing the JSON
- Use
hmac.compare_digest(or equivalent) to avoid timing attacks
Prometheus Metrics
The operator exposes these metrics on its metrics endpoint (port8080, path /metrics):
There is no metric for throttled notifications. Throttling is visible only in the logs (
notification throttled).
Recommended Prometheus alerts:
Next Steps
SLOs and SLAs
Service Level Objectives management with burn rate alerting
Approval Workflow
Change control with approval policies and blast radius
AIOps Platform
Deep-dive into the AIOps architecture
K8s Operator
Operator configuration and CRDs