Issue when a burn-rate window fires. The SLA controller times every Issue against response and resolution targets per severity, with optional business hours.
SLO vs SLA: Understanding the Difference
Best practice is to define SLOs that are more stringent than SLAs. If your SLA guarantees 99.9%, set the SLO at 99.95%. This creates a safety margin (internal error budget) that allows detecting degradations before the SLA is violated.
ServiceLevelObjective CRD
TheServiceLevelObjective defines a reliability target for a service, with error budget tracking and burn-rate alerts.
Spec Fields
Root
SLOIndicator
How each indicator is measured
All lookups are restricted to the SLO’s own namespace. Issues and Anomalies are created in the namespace of the affected workload, so the SLO must live there too.
What this means in practice:
availabilityis the only time-based indicator, and the one whose error budget maps to “allowed downtime”. Minutes from overlapping Issues are added up, so two concurrent Issues on the same service count double.- The anomaly-based indicators measure the share of anomalies of a given kind, not a share of requests. With no anomalies in the window the SLI is
1.0. If every anomaly in the window is of the counted kind, the SLI is0. - The Issues the SLO opens itself carry the label
platform.chatcli.io/service=<serviceName>, butavailabilityleaves everyslo_violationIssue out of the incident minutes: a page is the consequence of downtime, not more downtime, so an open SLO violation Issue does not deepen the breach. - A window that cannot be parsed (see below) becomes zero, which makes the SLI
1.0and every burn rate0. The mistake is silent.
SLOTarget
SLOAlertPolicy
BurnRateWindow
- Availability
- Error Rate
- Latency
- Throughput
How the Calculation Works (Google SRE Model)
The controller reconciles every SLO every 60 seconds and also whenever the SLO or one of the Issues it owns changes.Error Budget
The error budget is the maximum amount of “error” allowed within the SLO window.availability SLO over a 30-day window, this means:
The status fields derive from it:
100 target has a zero budget: the budget is then either fully remaining (SLI = 1) or exhausted.
Burn Rate
The burn rate indicates how fast the error budget is being consumed.burnRate1h, burnRate6h, burnRate24h and burnRate72h. The windows you list in burnRateWindows are computed separately each reconcile, only to decide whether to fire.
1
Compute the error rate in each window
2
Compute the burn rate
Divide the error rate by the error budget.
3
Check both windows
An alert fires only when the burn rate is >= threshold in the short AND the long window.
4
Open an Issue
A new alert records an entry in
status.activeAlerts, increments chatcli_operator_slo_violations_total, writes an slo_violation AuditEvent and opens an Issue (see What fires). The Issue then goes through the normal incident flow: notifications, escalation, SLA timing.Multi-Window Alerting: Recommended Thresholds
There are no built-in windows: nothing fires until you listburnRateWindows. For a 30-day SLO, these are the commonly used values:
These thresholds were designed for request-based SLIs. They fit the time-based
availability indicator. With the anomaly-ratio indicators, burn rates jump in large steps: one error_rate anomaly among 2 in the window gives an error rate of 0.5, i.e. a burn rate of 500x for a 99.9% target. Choose targets and thresholds for those indicators accordingly.Complete Numerical Example
Consider anavailability SLO of 99.9% over 30 days for api-gateway, and one incident: an Issue on api-gateway was detected 20 hours ago and resolved 20 minutes later.
Error Budget Tracking
Status Fields
The CRD has no state field; the REST API derives
Healthy/AtRisk/Breached from the status (see Checking SLOs). There are no budget warning thresholds (50/25/10%). Build those as Prometheus alerts on chatcli_operator_slo_error_budget_remaining (see below).
What fires when
Issues opened by an SLO have
spec.source: watcher, spec.signalType: slo_violation, spec.resource = indicator.resource (or the SLO itself when it is not set), a riskScore based on the budget consumed (50/75/90/100 steps), the labels platform.chatcli.io/signal=slo_violation and platform.chatcli.io/service=<serviceName>, and the annotations platform.chatcli.io/slo-name, slo-window and burn-rate. The SLO is their owner, so deleting the SLO deletes them.
A budget that recovers and runs out again pages again: each exhaustion opens exactly one Issue.
Checking SLOs
CURRENT is a fraction, not a percentage. BUDGETREMAINING% is status.errorBudgetRemainingPercentage, the share of the budget left, in percent.X-API-Key, role viewer or higher) exposes SLOs read-only:
Each SLO in the REST responses also carries
burnRate1h, burnRate6h, burnRate24h, burnRate72h, activeAlerts (how many alerts are firing), errorBudgetUsed (the consumed part, in the same unit as errorBudgetTotal) and a derived state:
In
/budget, burnRate is the 1-hour burn rate (1.0 = consuming the budget exactly at the allowed pace). GET /api/v1/analytics/summary counts AtRisk and Breached SLOs in slosAtRisk.
IncidentSLA CRD
AnIncidentSLA defines the response and resolution targets for one severity. Create one object per severity.
Spec Fields
responseTime and resolutionTime accept whole days in front of a Go duration (1d, 2d12h); weeks and fractional days (1.5d) are not accepted. The CRD validates the pattern, so the API server rejects a malformed value. A value that passes the pattern but still cannot be used (for example 0m) sets the condition Ready=False with reason InvalidDuration, and the SLA is not enforced until it is fixed; a valid SLA reports Ready=True.Which SLA applies to an Issue
The controller evaluates everyIssue. It uses the first IncidentSLA with the Issue’s severity in the Issue’s namespace. If there is none, it uses the first one with that severity found in any namespace. You can therefore keep a cluster-wide default set in one namespace and override it per namespace. Keep a single IncidentSLA per severity per namespace: with several, which one wins is not defined.
How response and resolution are measured
The start of the clock is the Issue’sstatus.detectedAt (or its creation time).
The timing has limits you should know about:
- Breaches are detected on state transitions, not by a running timer. An Issue that stays in
DetectedpastresponseTimeis not flagged until it moves toAnalyzing/Remediating. An open Issue pastresolutionTimeis not flagged until it becomesResolved,EscalatedorFailed. - An Issue that never passes through
AnalyzingorRemediatingnever gets a response check. Containedis not a terminal state: the resolution clock keeps running untilResolved/Escalated/Failed.- An Issue is timed only once. If an
EscalatedIssue is resolved later, it is not re-evaluated. - The controller marks the Issue with the annotations
platform.chatcli.io/sla-response-checked,platform.chatcli.io/sla-resolution-checkedand, on a breach,platform.chatcli.io/sla-violated(response,resolutionor both).
What a breach does
Each breach:- appends a record to
status.recentViolations(only the last 50 are kept), - increments
totalViolations,activeViolationsandchatcli_operator_sla_violations_total{severity,type}, and records the SLA on the Issue (annotationplatform.chatcli.io/sla-name), - sets
lastViolationAtand the conditionSLAViolation=True(reasonresponseViolationorresolutionViolation), - writes an
sla_breachAuditEvent.
chatcli_operator_sla_violations_total in Prometheus.
activeViolations counts violations whose Issue is not resolved yet. When the Issue reaches Resolved, its violations are subtracted from the SLA named in platform.chatcli.io/sla-name (even if the Issue’s severity changed since), and when it drops to 0 the condition becomes SLAViolation=False (reason NoActiveViolation). An Issue that ends Escalated or Failed keeps its violations active until it is resolved.BusinessHoursSpec
There is no holiday calendar: the clock runs on every listed weekday.
How the Business Hours Clock Works
WithbusinessHoursOnly: true and businessHours set, only time inside the window counts. Outside it the clock is paused.
1
Incident detected
Issue detected at 17:45 (Friday).
2
Clock counts 15 minutes (Friday)
From 17:45 to 18:00 = 15 minutes of SLA clock.
Clock pauses at 18:00 (end of business hours).
3
Weekend: clock paused
Saturday and Sunday are not in
workDays.
Accumulated SLA time: 15 minutes.4
Monday: clock resumes
Clock resumes at 09:00 on Monday.
If the incident is resolved at 10:30 on Monday:
- Friday: 15 minutes
- Monday: 1h30 = 90 minutes
- Total SLA: 105 minutes (1h45)
5
Compliance evaluation
With a
critical SLA of resolutionTime: 1h:- SLA time spent: 105 minutes
- Limit: 60 minutes
- VIOLATION
high SLA of resolutionTime: 4h:- SLA time spent: 105 minutes
- Limit: 240 minutes
- WITHIN SLA
CompliancePercentage Calculation
totalIssuesTrackedcounts Issues when they close (Resolved,EscalatedorFailed).totalViolationscounts breaches: an Issue that breaches both response and resolution counts twice. Compliance is clamped at 0 and is 100 while nothing has been tracked.- Each
IncidentSLAhas its own compliance, which is also its severity’s compliance (chatcli_operator_sla_compliance_percentage{severity}). The counters are cumulative since the object was created: there is no rolling period. averageResponseTimeandaverageResolutionTimeare recomputed when an Issue is resolved. They cover all Issues of that severity in the cluster, not just the SLA’s namespace. The response time is taken from the Issue’sAnalyzingcondition.
The compliance report at
GET /api/v1/analytics/compliance covers the absolute period from-to (default: the last 7 days). With IncidentSLA objects in scope, its SLA section counts the response and resolution violations their controller recorded on each Issue (platform.chatcli.io/sla-violated), and IncidentSLAs[] lists every SLA with its own counters (CompliancePercentage, ActiveViolations, TotalViolations, TotalIssuesTracked). Without any IncidentSLA, it falls back to counting every Escalated Issue as a resolution violation. The report keeps its PascalCase JSON keys, and durations are in nanoseconds. GET /api/v1/policies/sla lists the IncidentSLA objects (read-only, role viewer).Complete YAML Examples
99.9% Availability SLO with Burn Rate Alerting
NotificationPolicy rule (see Notifications):
Incident SLAs: Critical 5min/1h (24/7), Others in Business Hours
OneIncidentSLA per severity:
medium-sla and low-sla rely on the businessHours defaults (09:00-18:00, Monday to Friday). The object must still be present: businessHoursOnly: true without businessHours falls back to a 24/7 clock.Latency SLO for One Deployment (Anomaly-Based)
signalType: latency (raised by the watcher from HighLatency/Latency alerts) against all Anomalies of payment-service in the window.
Grafana Dashboards
The repository ships 4 Grafana dashboards indeploy/grafana/, built on the operator metrics:
SLO Burn Rate & SLA Compliance
slo-burn-rate.json: error budget remaining, current SLI, burn rate per fixed window (1h/6h/24h/72h) against the 14.4/6/3/1x reference lines, time until budget exhaustion, SLA compliance, response/resolution time distribution and violations by type.Remediation Stats & Operator Health
remediation-stats.json: remediation statistics, plus an SLA section with compliance and p95 response/resolution time by severity.AIOps Overview
aiops-overview.json: platform-wide view of issues, anomalies and remediations.Incident Timeline & Workflows
incident-timeline.json: incident flow from detection through analysis, remediation and resolution.Prometheus Metrics
The operator exposes these metrics on its metrics port (8080, path /metrics).
SLO Metrics
SLA Metrics
The gauges are only updated while the SLO or SLA is reconciled. Deleting an SLO does not remove its last values from
/metrics until the operator restarts.Next Steps
Notifications and Escalation
Multi-channel notification system and automatic escalation
Approval Workflow
Change control with approval policies and blast radius
AIOps Platform
Deep-dive into the AIOps architecture
K8s Operator
Operator configuration and CRDs