Skip to main content
The ChatCLI AIOps platform includes three subsystems that sit next to the incident pipeline: Capacity Planner (an on-demand forecast for the resources behind your incidents), Noise Reducer (suppression of repetitive, seasonal, flapping and high-fatigue anomalies), and Cost Tracker (LLM spend per incident, served over the REST API).

Capacity Planner

The Capacity Planner produces a CPU and memory report for each Deployment that appears as the resource of an Issue. It is not a background job: it runs only when GET /api/v1/analytics/capacity is called, and it keeps no history of its own.
The planner does not read real usage and collects no usage history. It reads no metrics-server and no Prometheus data (PROMETHEUS_URL is not used here). “Current usage” is the requests of the Deployment’s first container (UsageSource: "requests"), and the percentage is requests divided by limits. Without history there is no trend to fit, so the response says Trend.Direction: "insufficient_history" and HistoryAvailable: false, and it never projects an exhaustion date. Treat the output as a requests-versus-limits check plus an incident count, not as a usage forecast.

What It Computes

For every distinct resource (kind/namespace/name) among the Issues listed (all Issues, or those in ?namespace=), the planner:
  1. Reads the Deployment with that name and namespace. If there is none (the resource is a StatefulSet, a DaemonSet, a Pod or has been deleted), the resource is skipped with no error.
  2. Takes the first container’s resources.limits and resources.requests for CPU and memory, and computes usage% = requests / limits × 100 (0 when there is no limit).
  3. Correlates the resource with the incidents of its namespace inside the window (below).
  4. Sets an urgency and one recommendation from what it knows: the requests/limits percentage and the incident correlation.
The window defaults to 7 days. When both from and to are passed (RFC 3339), the window becomes their difference, still ending now.

Trend and Urgency

Urgency has only these two values: without history nothing is ever urgent. The dashboard’s capacity banner lists the resources with Urgency: "plan".

Response Fields

The endpoint returns kind: CapacityForecastList with one CapacityForecast per resource. Field names are the Go names (no JSON renaming), except the nested resource reference:

Correlation with Incidents

A bottleneck raises Urgency to plan and, when requests are within limits, sets the recommendation.

Recommendation Generation

The first matching rule wins:

How to Use

The endpoint needs the viewer role (GET only). Each call lists every Issue in scope and reads one Deployment plus the namespace’s Issues per resource, so keep the calls occasional on large clusters.

Noise Reducer

The Noise Reducer runs inside the Anomaly controller. It is consulted only for an anomaly that would open a new Issue: an anomaly whose resource already has an active Issue is correlated to it first, and one that arrives within the resolution cooldown (spec.aiops.resolutionCooldownMinutes of the Instance the anomaly came from, default 10, 0 turns it off) after a resolution is suppressed before the Noise Reducer is reached. The four checks run in order and the first one that fires wins. A suppressed anomaly is marked status.correlated: true with the reason recorded as suppressed:<reason>, no Issue is created, and chatcli_operator_anomalies_processed_total{result="suppressed_noise"} is incremented. If a check fails (for example, a List error), the anomaly is not suppressed. Resources are matched by name within the anomaly’s namespace.

Strategy 1: Repetitive Suppression

Severity is not part of the match.

Strategy 2: Seasonal Patterns

Every anomaly that is not suppressed is recorded as a seasonal occurrence. Patterns are stored per namespace in the ConfigMap chatcli-seasonal-patterns, key patterns, as a JSON array. SeasonalPattern (JSON): Algorithm:
Hours and weekdays come from the operator pod’s clock (UTC unless you set TZ on the operator; the image ships tzdata). The hour distance does not wrap around midnight (23:00 and 00:00 are 23 hours apart).
occurrences counts anomalies, not distinct weeks. Three anomalies for the same signal and resource in the same hour slot of one day are enough to suppress later ones in that slot, on that day and on the same weekday every week after. There is no confidence score and no pruning: patterns stay in the ConfigMap until you edit or delete it.

Strategy 3: Flap Detection

There is no separate flapping flag, hold-off timer or consolidated alert: each new anomaly is checked the same way, and suppression stops once fewer than three resolved Issues remain inside the 24-hour window.

Strategy 4: Alert Fatigue Scoring

In practice, as few as 7 anomalies for a resource in 24 hours, all already correlated, with a recent reference anomaly, cross the threshold (35 + 30 + 20 = 85), and further anomalies for it are suppressed until the count drops.
The recency reference is the last item of the namespace’s unfiltered Anomaly list, not the newest anomaly of the resource, so recency_score can reflect another resource. The threshold > 80 is fixed and not configurable.

Cost Tracker

The Cost Tracker books the LLM spend of every incident from the usage the server reports, and serves the aggregate over the REST API.

LLM Costs per Provider

Every analysis (AnalyzeIssue, AIInsight controller) and every agentic remediation step (AgenticStep, Remediation controller) is booked from the token usage the server reports on the reply, attributed to the provider and model that actually answered (the server’s fallback chain may route elsewhere). When the reply carries no usage (a server older than the usage fields), analyses fall back to characters divided by four for both input and output, and agentic steps book zero input tokens and the reasoning length divided by four as output. The price per token is resolved in this order:
  1. an entry for the provider in the chatcli-cost-config ConfigMap of the Issue’s namespace (the cluster operator’s word);
  2. the shared pricing engine (llm/pricing), the same per-model tables, overrides and subscription rules the CLI’s /cost uses (including a CHATCLI_MODEL_PRICING override set on the operator process, for example through the chart’s extraEnv), so a ledger here and a session there agree on the price of the same call;
  3. a per-provider default for a model the engine does not know (for example CLAUDEAI 3/3/15, OPENAI 10/10/30, OPENROUTER 2/2/8, DEVIN 0,anythingelse0, anything else 1/$3 per million input/output tokens).
Only input and output rates are applied: cache-read and cache-write discounts are not modeled on the ledger. The server also reports cost_usd / cost_known on the reply itself (see server mode); the ledger does not use those fields.

Cost Configuration

Prices are overridable per provider via the chatcli-cost-config ConfigMap, key pricing, one entry per provider name exactly as the server reports it (CLAUDEAI, OPENAI, …). The tracker looks it up in the namespace of the Issue being booked, the same namespace as its ledger, so create one in every namespace whose incidents you want to price differently:
An entry applies to every model of that provider; there is no per-model key. Without the ConfigMap, for a provider it does not list, or when the pricing value is not valid JSON, the pricing engine’s per-model rates apply. The ConfigMap is read on every booking, so an update applies from the next call on; calls already booked keep the price they were booked at.

IncidentCost

One ledger entry per incident, written by the AIInsight and Remediation controllers on every LLM call:
Each call is priced once, at the rates of the provider and model that answered it, and added to the entry’s total. A later call on another model (for example after a fallback) never reprices the earlier ones; provider and model only name the latest call.
Calculation example (with the CLAUDEAI override above):
The usage the server reports on each response (usage on AnalyzeIssue and AgenticStep) is what gets booked, so the ledger reflects real token counts, not estimates.
Two calls booked at the same time in one namespace do not overwrite each other: the ledger update carries the version it read, and a write conflict is retried. A booking that still fails (the ConfigMap size limit, missing RBAC, a ledger that cannot be read or an entry that cannot be decoded) is logged by the controller (Failed to book the analysis cost on the ledger or Failed to book the agentic step cost on the ledger); the analysis or step itself is not affected, and an unreadable entry is never overwritten with a fresh one.

CostSummary

Aggregation over a period, served by the REST API (viewer role, GET only), as kind: CostSummary with the summary under spec: from and to (RFC 3339) set an absolute period: with both, it is exactly [from, to]; with only from, it runs to now; with only to, it is the 30 days ending at to; with neither, it is the last 30 days. An entry is counted in full when its recordedAt (last booking) falls in the period; entries without recordedAt are always counted. Without namespace, the summary aggregates every chatcli-cost-ledger in the cluster. A namespace without a ledger returns zeros; any other failure to read a ledger returns 500 (failed to read the cost ledger: ...) instead of zeros.
Downtime cost and return-on-investment figures are not computed: the platform has no trustworthy source for revenue per minute or engineer hourly rates, and numbers built on assumed constants would be misleading. Combine totalLLMCost with your own incident economics if you need them.

Storage Architecture (ConfigMaps)

These subsystems persist their data in ConfigMaps in the namespace of the Issue or Anomaly they concern (not the operator namespace). The Capacity Planner stores nothing.
ConfigMaps have a 1MB limit in Kubernetes, roughly 3,000 ledger entries. The operator does not compact the ledger or the seasonal patterns: on a namespace with that volume, prune old keys periodically, or new bookings fail (each failure is logged, see above).

Storage Format

One ConfigMap per namespace, chatcli-cost-ledger, one key per incident holding the IncidentCost JSON. The tracker creates it with the app.kubernetes.io/managed-by: chatcli-operator label, which is how the all-namespace summary finds every ledger:
Entries are never compacted or expired by the operator: delete keys, or the ConfigMap, to reset a namespace’s ledger.

Integrations

REST API

/api/v1/analytics/cost serves the ledger summary; /api/v1/analytics/capacity serves the capacity forecasts; /api/v1/analytics/summary exposes incident data.

Web Dashboard

The Overview view has a capacity-warnings banner fed by /analytics/capacity: it lists the resources with Urgency: "plan" and their recommendation. The dashboard does not show the cost ledger.

Grafana

The remediation-stats.json dashboard covers remediation success by action type. No bundled dashboard covers capacity or LLM cost.

AIOps Platform

Complete AIOps pipeline architecture and how these subsystems integrate.