Skip to main content
The AIOps platform notification system sends Issue state changes to the right teams on the right channels (Slack, PagerDuty, OpsGenie, email, generic webhooks and Microsoft Teams). Escalation policies add a timed chain of levels for Issues that automatic remediation could not fix. Everything described on this page is done by one controller in the operator, the notification controller (NotificationReconciler), which watches Issue resources. NotificationPolicy and EscalationPolicy have no controller of their own: the notification controller reads them whenever an Issue changes.

Overview

What triggers a notification

The only trigger is a change of status.state on an Issue. The controller stores the last state it handled in the Issue annotation platform.chatcli.io/last-notified-state, and it evaluates the policies once for each new state: Detected, Analyzing, Remediating, Contained, Resolved, Escalated, Failed. Other events reach a channel only if they turn into an Issue state change:
IncidentSLA.spec.notificationPolicyRef, IncidentSLA.spec.escalationPolicyRef and ServiceLevelObjective.spec.alertPolicy.notificationPolicyRef are reserved: the CRDs accept them, but no controller reads them yet. Routing is decided only by the rules of your NotificationPolicies. To route SLO alerts, match signalTypes: [slo_violation].
Issues created by a ChaosExperiment (label platform.chatcli.io/source: chaos-experiment) still produce normal state-change notifications. They never start an escalation, so a chaos drill never pages anyone through an EscalationPolicy.

NotificationPolicy CRD

A NotificationPolicy (short name np) declares a set of named channels, a list of rules that pick channels by name, throttling, and optional message templates.
The Secrets referenced above live in the same namespace as the policy:

How policies are applied

  • Every enabled policy in every namespace is evaluated for every Issue in the cluster. The namespace of the policy does not limit its scope. Use the rule’s namespaces filter for that.
  • Inside a policy, every matching rule sends to its channels. The same channel can receive the same Issue twice if two rules match, but the dedup window (below) usually suppresses the second send.
  • A rule that names a channel missing from spec.channels is skipped with the log line channel not found in policy.

Spec Fields

NotificationPolicySpec

NotificationChannel

config is a map of strings. Numbers, booleans, lists and maps must be written as strings: smtp_port: "587", tls_skip_verify: "true", to: "a@example.com,b@example.com", headers: '{"X-Env":"prod"}'. An unquoted 587 or a YAML list is rejected by the API server.
GET /api/v1/policies/notification on the operator REST API returns the full spec, including config, to any key with the viewer role. Keep webhook URLs, routing keys, API keys and passwords in the secretRef Secret, not in config.

NotificationRule

The filters sit directly on the rule (there is no match block). An omitted or empty filter matches everything. The logic is AND between filters and OR within a filter.
Always list states. Without it, every transition of the Issue (Detected → Analyzing → Remediating → …) is a candidate. Also include Resolved for PagerDuty and OpsGenie channels: that is what resolves or closes the alert on their side.

ThrottleConfig

Throttled notifications are dropped, not queued, and the state is still marked as handled, so it is never sent later. With the default 5m window, an Issue that goes Detected → Analyzing → Remediating → Resolved within five minutes sends only the first matching state to Slack, email, webhook and Teams channels. For channels that must see every transition, use a short window such as 30s.One exception: a Resolved notification to a PagerDuty or OpsGenie channel is never throttled, by either deduplicationWindow or maxPerHour, because it is what resolves the event or closes the alert on their side.The throttle state is kept in the operator’s memory: it resets when the operator restarts. The window is measured per channel name, so two policies that use the same channel name share it.

Templates

templates replaces the message body (the title and email subject stay the same). The key is chosen from the new state: The key remediation_completed is accepted but never used. Templates are not Go text/template: they are plain placeholder replacement, and only these exact tokens are replaced: {{.Name}} (Issue name), {{.Namespace}}, {{.Severity}}, {{.State}}, {{.Resource}} (Kind/name), {{.Description}}, {{.Source}}, {{.SignalType}}, {{.RiskScore}}. Anything else (conditionals, functions, other fields) is sent literally. Without a template, the body is the Issue’s spec.description, or Issue <name> on <Kind>/<name> transitioned to <State>. when the description is empty.

Message content

Every channel receives the same message:
  • Title: <severity emoji> [<SEVERITY>] <namespace>/<resource name> — <State>, for example 🔴 [CRITICAL] production/api-gateway — Detected.
  • Body: the template or description (see above).
  • Fields: Source, SignalType, RiskScore, plus CorrelationID, RemediationAttempts (n/max) and Resolution when they are set.
  • Color by severity: critical #FF0000, high #FF8C00, medium #FFD700, low #00CC00.
The message does not include the AI analysis, the remediation plan or a dashboard link. Every HTTP channel uses a 30-second timeout and makes a single attempt, with no retries.

Status

Escalation notifications are not counted in this status. Every delivery attempt, including escalation sends, is also recorded as an AuditEvent with eventType: notification_sent (see Audit and Compliance).

Notification Channels

1. Slack

Posts to a Slack Incoming Webhook using Block Kit.
No other keys are read (there is no mention or icon option). To mention a group, put the mention in a template, for example issue_created: "<!subteam^S0123ABC> {{.Description}}".Payload sent:
The extra fields (SignalType, RiskScore, …) come from a map, so their order can change between messages.
Minimal example:

2. PagerDuty

Sends events to the PagerDuty Events API v2 (https://events.pagerduty.com/v2/enqueue; the endpoint is fixed).
No other keys are read. The severity mapping and the dedup key are fixed (the severity_map key mentioned in the CRD field description is not implemented):Deduplication: dedup_key is always chatcli-<resource namespace>-<issue name>, so every notified state of one Issue updates the same PagerDuty alert.Payload sent:
Automatic resolution: when a notification is sent for the Resolved state, the event uses event_action: resolve with the same dedup_key. This happens whenever a rule sends Resolved to this channel: resolves to PagerDuty are never throttled.

3. OpsGenie

Creates alerts through the OpsGenie Alert API (https://api.opsgenie.com/v2/alerts; the endpoint is fixed, so accounts on the EU instance are not supported).
Priority mapping (fixed): critical → P1, high → P2, medium → P3, low → P4.Responders:
The alert alias is chatcli-<resource namespace>-<issue name>, source is ChatCLI AIOps, entity is the resource, and details has resource, namespace, severity, state and issue.Automatic close: for the Resolved state, the channel closes the alert by alias instead of creating one (same conditions as PagerDuty).

4. Email

Sends an HTML email over SMTP.
There are no cc, bcc, subject or HTML template options. The subject is always [<SEVERITY>] <title> and the body is a fixed HTML layout with the fields table.TLS behavior:
  • Implicit TLS (SMTPS) is used on port 465, or on any port with smtp_tls: implicit: the connection is encrypted from the first byte.
  • Otherwise the connection starts in plain text and is upgraded with STARTTLS when the server advertises it. smtp_tls: starttls forces this mode even on port 465.
  • If the connection is not encrypted and smtp_user is set, authentication fails (Go refuses to send PLAIN credentials over an unencrypted connection, except to localhost).
  • The whole conversation, dial included, is bounded by smtp_timeout (default 30s), so an unreachable or silent server fails the send instead of holding the reconcile.
Example:
Never put SMTP credentials directly in the NotificationPolicy YAML. Put smtp_user and smtp_password in a Secret and reference it with secretRef.

5. Webhook

Sends the message as JSON to any HTTP endpoint, optionally signed with HMAC-SHA256.
Every request has Content-Type: application/json and User-Agent: ChatCLI-AIOps/1.0. The timeout is 30 seconds and there are no retries.HMAC-SHA256 signing:When secret is set, the request carries the header X-Signature-256 with the HMAC-SHA256 of the raw body:
Validation on the receiver:
JSON payload sent:

6. Microsoft Teams

Posts an Adaptive Card (version 1.4) to a Teams incoming webhook URL.
No other keys are read.Generated card:
  • A large, bold title (the message title)
  • The body text
  • A FactSet with Severity, Resource, Namespace, State, Issue and the extra fields
  • A footer with the generation time
The message also carries themeColor with the severity color without # (FF0000, FF8C00, FFD700, 00CC00).

EscalationPolicy CRD

An EscalationPolicy (short name ep) is an ordered chain of levels. It is used when an Issue enters the Escalated state, which the Issue controller sets when automatic remediation gives up (for example, all remediation attempts failed).

Spec Fields

EscalationLevel

EscalationTarget

Targets are informational only. The operator does not contact users, teams or on-call schedules. It writes the list (for example oncall:sre-primary, user:sre-lead@example.com) into the escalation message and sends that message to the level’s notifyChannels. A level without notifyChannels notifies nobody.

How Escalation Works

  1. Start. When an Issue changes to Escalated (and it is not chaos-induced), the controller picks a policy: the first enabled policy, from any namespace, whose severities contains the Issue’s severity, or that has no severities. A policy with defaultPolicy: true is used only when no policy matched. If nothing is found, no escalation happens (log line no escalation policy found for issue).
  2. Level 1 is notified right away. The Issue gets the annotations below and the policy status gets an activeEscalations entry.
  3. Advance. When timeoutMinutes of the current level has passed, the next level is notified with the title <emoji> ESCALATION [<SEVERITY>] <issue> — Level <n>: <level name>. The new level and its start time are saved on the Issue, so the chain moves forward level by level (L1 → L2 → L3) and each level is notified once, plus its repeats (repeatIntervalMinutes).
  4. Last level. The chain stops there. The last level is re-sent only if it has a repeatIntervalMinutes.
  5. Stop. An acknowledgement freezes the chain at its current level: no further level and no repeat (see below). The escalation ends when the Issue reaches Resolved: the escalation annotations are removed and the Issue’s entry leaves status.activeEscalations.
Escalation messages are sent to every channel in any enabled NotificationPolicy whose name is listed in notifyChannels (Secrets are read from that policy’s namespace). They do not go through the rules, the throttle or the templates. Issue annotations used for tracking:
Each status.activeEscalations entry carries issueName, currentLevel (0-based), escalatedAt and, once the Issue is acknowledged, acknowledgedAt and acknowledgedBy. The entry is removed when the Issue resolves; status.totalEscalations counts every escalation started.

Acknowledgement and stopping an escalation

Both actions go through the REST API (operator role, X-API-Key header) or the dashboard:
  • Acknowledge. POST /api/v1/incidents/{name}/acknowledge adds the annotations aiops.chatcli.io/acknowledged, aiops.chatcli.io/acknowledged-at and aiops.chatcli.io/acknowledged-by (the caller’s role). The escalation stops at its current level: no further level is reached and no repeat is sent. The acknowledgement is stamped on the policy’s status.activeEscalations entry (acknowledgedAt, acknowledgedBy). An Issue acknowledged before it reaches Escalated never starts an escalation. State-change notifications keep flowing.
  • Snooze. POST /api/v1/incidents/{name}/snooze with a body such as {"duration": "30m"} records aiops.chatcli.io/snoozed-until and aiops.chatcli.io/snoozed-by. A duration of zero or less is rejected with 400. Until the snooze ends, the Issue sends no notification except Resolved (a state change that happens during the snooze is not sent later), and its escalation holds its level: no advance and no repeat. A level message that falls inside the snooze is held (escalation-pending-notify) and sent when the snooze ends, and the level’s timer restarts at that moment.
  • Acknowledging in PagerDuty or OpsGenie does nothing on the cluster side: there is no inbound webhook.
To end an escalation, resolve the Issue, for example with POST /api/v1/incidents/{name}/resolve (operator role) or from the dashboard. That also sends the Resolved notifications that resolve the PagerDuty event and close the OpsGenie alert.

Complete Examples

Notification Policy: Slack + PagerDuty

The short 30s window lets the Slack channels see each transition of a fast recovery. The Resolved event that resolves the PagerDuty incident is never throttled, whatever the window.

Escalation Policy with two levels

email-leadership must be a channel defined in some NotificationPolicy. After the second level is reached, nothing more is sent until the Issue is resolved; add repeatIntervalMinutes to the second level to keep reminding until someone acknowledges.

SLO violation alerts

SLO burn-rate and budget-exhaustion alerts arrive as Issues with signalType: slo_violation and source: watcher:

Troubleshooting

Diagnostic checklist:
  1. Check that the policy exists and is enabled (policies in any namespace apply):
  1. Check the policy status for delivery errors:
  1. Check the operator logs (rule matched, notification sent, failed to send notification, channel not found in policy, failed to resolve channel config):
  1. Compare the Issue with the rule filters. Remember that namespaces is compared with spec.resource.namespace:
  1. Check whether the throttle dropped it (log line notification throttled). Another state of the same Issue may have been sent to that channel within deduplicationWindow.
  2. If platform.chatcli.io/last-notified-state already equals the current state, that state was already handled and is not evaluated again.
  • Every config value must be a string: quote numbers (smtp_port: "587") and booleans (tls_skip_verify: "true"), and write lists as comma-separated strings.
  • Every channel needs name, type and config (use config: {} when the values come from secretRef).
  • Rule filters sit directly on the rule; a match: block is not part of the schema.
  • Confirm that the webhook_url is correct and the Slack app is still installed in the workspace
  • Test the webhook manually:
  • Confirm that the routing_key is an Events API v2 Integration Key (not a REST API key)
  • Confirm that the service in PagerDuty is active, and check the payload in the PagerDuty Event Debugger
  • If incidents are never resolved, make sure a rule sends Resolved to the channel (resolves to PagerDuty are never throttled)
  • Test SMTP connectivity from inside the cluster:
  • On port 465 the channel uses implicit TLS; on 587 or 25 it upgrades with STARTTLS. Set smtp_tls if your server uses a non-standard port.
  • A send that fails with a timeout hit smtp_timeout (default 30s)
  • Confirm that the Secret has the keys smtp_user and smtp_password (not username/password)
  • Check the recipients’ spam folder
  • Escalation starts only when the Issue enters Escalated. Check kubectl get issue <name> -o jsonpath='{.status.state}'.
  • Chaos-induced Issues never escalate.
  • Check the annotations platform.chatcli.io/escalation-level, escalation-time and escalation-policy.
  • Make sure each level has notifyChannels whose names exist in an enabled NotificationPolicy.
  • Look for escalation initiated, escalation advanced and no escalation policy found for issue in the operator logs.
  • An acknowledged Issue does not advance, and a snoozed one holds its level until aiops.chatcli.io/snoozed-until.
  • Read the signature from the X-Signature-256 header
  • Confirm that the secret in the policy’s Secret is the same one the receiver uses
  • Compute the HMAC over the raw body, before parsing the JSON
  • Use hmac.compare_digest (or equivalent) to avoid timing attacks

Prometheus Metrics

The operator exposes these metrics on its metrics endpoint (port 8080, path /metrics): There is no metric for throttled notifications. Throttling is visible only in the logs (notification throttled). Recommended Prometheus alerts:

Next Steps

SLOs and SLAs

Service Level Objectives management with burn rate alerting

Approval Workflow

Change control with approval policies and blast radius

AIOps Platform

Deep-dive into the AIOps architecture

K8s Operator

Operator configuration and CRDs