Skip to main content
The ChatCLI AIOps platform manages Service Level Objectives (SLOs) and incident Service Level Agreements (SLAs) through two Kubernetes CRDs. The SLO controller computes an SLI, an error budget and multi-window burn rates (following the Google SRE model) and opens an Issue when a burn-rate window fires. The SLA controller times every Issue against response and resolution targets per severity, with optional business hours.
Read this before you design SLOs. The SLO controller does not query Prometheus. Every SLI is computed from the operator’s own Issue and Anomaly objects (see How each indicator is measured). metricSource: prometheus, prometheusQuery, latencyPercentile and latencyThresholdMs are accepted by the CRD but currently have no effect; an SLO with metricSource: prometheus says so in its status with the condition MetricSourceSupported=False (reason PrometheusNotEvaluated). Use request-based PromQL SLOs in your own Prometheus/Alertmanager setup if you need them; use this CRD for incident-based availability and error-budget tracking of the workloads the operator watches.

SLO vs SLA: Understanding the Difference

Best practice is to define SLOs that are more stringent than SLAs. If your SLA guarantees 99.9%, set the SLO at 99.95%. This creates a safety margin (internal error budget) that allows detecting degradations before the SLA is violated.

ServiceLevelObjective CRD

The ServiceLevelObjective defines a reliability target for a service, with error budget tracking and burn-rate alerts.
The controller writes the status (never set it yourself):

Spec Fields

Root

SLOIndicator

How each indicator is measured

All lookups are restricted to the SLO’s own namespace. Issues and Anomalies are created in the namespace of the affected workload, so the SLO must live there too. What this means in practice:
  • availability is the only time-based indicator, and the one whose error budget maps to “allowed downtime”. Minutes from overlapping Issues are added up, so two concurrent Issues on the same service count double.
  • The anomaly-based indicators measure the share of anomalies of a given kind, not a share of requests. With no anomalies in the window the SLI is 1.0. If every anomaly in the window is of the counted kind, the SLI is 0.
  • The Issues the SLO opens itself carry the label platform.chatcli.io/service=<serviceName>, but availability leaves every slo_violation Issue out of the incident minutes: a page is the consequence of downtime, not more downtime, so an open SLO violation Issue does not deepen the breach.
  • A window that cannot be parsed (see below) becomes zero, which makes the SLI 1.0 and every burn rate 0. The mistake is silent.

SLOTarget

SLOAlertPolicy

BurnRateWindow

How the Calculation Works (Google SRE Model)

The controller reconciles every SLO every 60 seconds and also whenever the SLO or one of the Issues it owns changes.

Error Budget

The error budget is the maximum amount of “error” allowed within the SLO window.
For an availability SLO over a 30-day window, this means:
The status fields derive from it:
A 100 target has a zero budget: the budget is then either fully remaining (SLI = 1) or exhausted.

Burn Rate

The burn rate indicates how fast the error budget is being consumed.
The controller always computes and publishes four fixed windows: burnRate1h, burnRate6h, burnRate24h and burnRate72h. The windows you list in burnRateWindows are computed separately each reconcile, only to decide whether to fire.
1

Compute the error rate in each window

2

Compute the burn rate

Divide the error rate by the error budget.
3

Check both windows

An alert fires only when the burn rate is >= threshold in the short AND the long window.
4

Open an Issue

A new alert records an entry in status.activeAlerts, increments chatcli_operator_slo_violations_total, writes an slo_violation AuditEvent and opens an Issue (see What fires). The Issue then goes through the normal incident flow: notifications, escalation, SLA timing.
There are no built-in windows: nothing fires until you list burnRateWindows. For a 30-day SLO, these are the commonly used values:
To derive a threshold: burn_rate_threshold = window_days / days_until_exhaustion. For a 30-day SLO where you want to alert when the budget would be exhausted in about 2 days: 30 / 2.08 = 14.4x.
These thresholds were designed for request-based SLIs. They fit the time-based availability indicator. With the anomaly-ratio indicators, burn rates jump in large steps: one error_rate anomaly among 2 in the window gives an error rate of 0.5, i.e. a burn rate of 500x for a 99.9% target. Choose targets and thresholds for those indicators accordingly.

Complete Numerical Example

Consider an availability SLO of 99.9% over 30 days for api-gateway, and one incident: an Issue on api-gateway was detected 20 hours ago and resolved 20 minutes later.

Error Budget Tracking

Status Fields

The CRD has no state field; the REST API derives Healthy/AtRisk/Breached from the status (see Checking SLOs). There are no budget warning thresholds (50/25/10%). Build those as Prometheus alerts on chatcli_operator_slo_error_budget_remaining (see below).

What fires when

Issues opened by an SLO have spec.source: watcher, spec.signalType: slo_violation, spec.resource = indicator.resource (or the SLO itself when it is not set), a riskScore based on the budget consumed (50/75/90/100 steps), the labels platform.chatcli.io/signal=slo_violation and platform.chatcli.io/service=<serviceName>, and the annotations platform.chatcli.io/slo-name, slo-window and burn-rate. The SLO is their owner, so deleting the SLO deletes them.
A budget that recovers and runs out again pages again: each exhaustion opens exactly one Issue.

Checking SLOs

CURRENT is a fraction, not a percentage. BUDGETREMAINING% is status.errorBudgetRemainingPercentage, the share of the budget left, in percent.
The operator REST API (header X-API-Key, role viewer or higher) exposes SLOs read-only: Each SLO in the REST responses also carries burnRate1h, burnRate6h, burnRate24h, burnRate72h, activeAlerts (how many alerts are firing), errorBudgetUsed (the consumed part, in the same unit as errorBudgetTotal) and a derived state: In /budget, burnRate is the 1-hour burn rate (1.0 = consuming the budget exactly at the allowed pace). GET /api/v1/analytics/summary counts AtRisk and Breached SLOs in slosAtRisk.

IncidentSLA CRD

An IncidentSLA defines the response and resolution targets for one severity. Create one object per severity.
Status written by the controller:

Spec Fields

responseTime and resolutionTime accept whole days in front of a Go duration (1d, 2d12h); weeks and fractional days (1.5d) are not accepted. The CRD validates the pattern, so the API server rejects a malformed value. A value that passes the pattern but still cannot be used (for example 0m) sets the condition Ready=False with reason InvalidDuration, and the SLA is not enforced until it is fixed; a valid SLA reports Ready=True.

Which SLA applies to an Issue

The controller evaluates every Issue. It uses the first IncidentSLA with the Issue’s severity in the Issue’s namespace. If there is none, it uses the first one with that severity found in any namespace. You can therefore keep a cluster-wide default set in one namespace and override it per namespace. Keep a single IncidentSLA per severity per namespace: with several, which one wins is not defined.

How response and resolution are measured

The start of the clock is the Issue’s status.detectedAt (or its creation time). The timing has limits you should know about:
  • Breaches are detected on state transitions, not by a running timer. An Issue that stays in Detected past responseTime is not flagged until it moves to Analyzing/Remediating. An open Issue past resolutionTime is not flagged until it becomes Resolved, Escalated or Failed.
  • An Issue that never passes through Analyzing or Remediating never gets a response check.
  • Contained is not a terminal state: the resolution clock keeps running until Resolved/Escalated/Failed.
  • An Issue is timed only once. If an Escalated Issue is resolved later, it is not re-evaluated.
  • The controller marks the Issue with the annotations platform.chatcli.io/sla-response-checked, platform.chatcli.io/sla-resolution-checked and, on a breach, platform.chatcli.io/sla-violated (response, resolution or both).

What a breach does

Each breach:
  • appends a record to status.recentViolations (only the last 50 are kept),
  • increments totalViolations, activeViolations and chatcli_operator_sla_violations_total{severity,type}, and records the SLA on the Issue (annotation platform.chatcli.io/sla-name),
  • sets lastViolationAt and the condition SLAViolation=True (reason responseViolation or resolutionViolation),
  • writes an sla_breach AuditEvent.
It does not notify anyone or escalate by itself. To be alerted, alert on chatcli_operator_sla_violations_total in Prometheus.
activeViolations counts violations whose Issue is not resolved yet. When the Issue reaches Resolved, its violations are subtracted from the SLA named in platform.chatcli.io/sla-name (even if the Issue’s severity changed since), and when it drops to 0 the condition becomes SLAViolation=False (reason NoActiveViolation). An Issue that ends Escalated or Failed keeps its violations active until it is resolved.

BusinessHoursSpec

There is no holiday calendar: the clock runs on every listed weekday.

How the Business Hours Clock Works

With businessHoursOnly: true and businessHours set, only time inside the window counts. Outside it the clock is paused.
1

Incident detected

Issue detected at 17:45 (Friday).
2

Clock counts 15 minutes (Friday)

From 17:45 to 18:00 = 15 minutes of SLA clock. Clock pauses at 18:00 (end of business hours).
3

Weekend: clock paused

Saturday and Sunday are not in workDays. Accumulated SLA time: 15 minutes.
4

Monday: clock resumes

Clock resumes at 09:00 on Monday. If the incident is resolved at 10:30 on Monday:
  • Friday: 15 minutes
  • Monday: 1h30 = 90 minutes
  • Total SLA: 105 minutes (1h45)
5

Compliance evaluation

With a critical SLA of resolutionTime: 1h:
  • SLA time spent: 105 minutes
  • Limit: 60 minutes
  • VIOLATION
With a high SLA of resolutionTime: 4h:
  • SLA time spent: 105 minutes
  • Limit: 240 minutes
  • WITHIN SLA
For critical incidents, consider leaving businessHoursOnly off and using a 24/7 clock. Critical production issues should not wait for the next business day. Business hours are set per IncidentSLA, so each severity can have its own choice.

CompliancePercentage Calculation

  • totalIssuesTracked counts Issues when they close (Resolved, Escalated or Failed). totalViolations counts breaches: an Issue that breaches both response and resolution counts twice. Compliance is clamped at 0 and is 100 while nothing has been tracked.
  • Each IncidentSLA has its own compliance, which is also its severity’s compliance (chatcli_operator_sla_compliance_percentage{severity}). The counters are cumulative since the object was created: there is no rolling period.
  • averageResponseTime and averageResolutionTime are recomputed when an Issue is resolved. They cover all Issues of that severity in the cluster, not just the SLA’s namespace. The response time is taken from the Issue’s Analyzing condition.
The compliance report at GET /api/v1/analytics/compliance covers the absolute period from-to (default: the last 7 days). With IncidentSLA objects in scope, its SLA section counts the response and resolution violations their controller recorded on each Issue (platform.chatcli.io/sla-violated), and IncidentSLAs[] lists every SLA with its own counters (CompliancePercentage, ActiveViolations, TotalViolations, TotalIssuesTracked). Without any IncidentSLA, it falls back to counting every Escalated Issue as a resolution violation. The report keeps its PascalCase JSON keys, and durations are in nanoseconds. GET /api/v1/policies/sla lists the IncidentSLA objects (read-only, role viewer).

Complete YAML Examples

99.9% Availability SLO with Burn Rate Alerting

Route the Issues it opens with a NotificationPolicy rule (see Notifications):

Incident SLAs: Critical 5min/1h (24/7), Others in Business Hours

One IncidentSLA per severity:
medium-sla and low-sla rely on the businessHours defaults (09:00-18:00, Monday to Friday). The object must still be present: businessHoursOnly: true without businessHours falls back to a 24/7 clock.

Latency SLO for One Deployment (Anomaly-Based)

The SLI counts Anomalies with signalType: latency (raised by the watcher from HighLatency/Latency alerts) against all Anomalies of payment-service in the window.

Grafana Dashboards

The repository ships 4 Grafana dashboards in deploy/grafana/, built on the operator metrics:

SLO Burn Rate & SLA Compliance

slo-burn-rate.json: error budget remaining, current SLI, burn rate per fixed window (1h/6h/24h/72h) against the 14.4/6/3/1x reference lines, time until budget exhaustion, SLA compliance, response/resolution time distribution and violations by type.

Remediation Stats & Operator Health

remediation-stats.json: remediation statistics, plus an SLA section with compliance and p95 response/resolution time by severity.

AIOps Overview

aiops-overview.json: platform-wide view of issues, anomalies and remediations.

Incident Timeline & Workflows

incident-timeline.json: incident flow from detection through analysis, remediation and resolution.
Importing the dashboards (from a checkout of the chatcli repository):

Prometheus Metrics

The operator exposes these metrics on its metrics port (8080, path /metrics).

SLO Metrics

SLA Metrics

The gauges are only updated while the SLO or SLA is reconciled. Deleting an SLO does not remove its last values from /metrics until the operator restarts.
Recommended Prometheus alerts:

Next Steps

Notifications and Escalation

Multi-channel notification system and automatic escalation

Approval Workflow

Change control with approval policies and blast radius

AIOps Platform

Deep-dive into the AIOps architecture

K8s Operator

Operator configuration and CRDs