Capacity Planner
The Capacity Planner produces a CPU and memory report for each Deployment that appears as the resource of an Issue. It is not a background job: it runs only whenGET /api/v1/analytics/capacity is called, and it keeps no history of its own.
What It Computes
For every distinct resource (kind/namespace/name) among the Issues listed (all Issues, or those in ?namespace=), the planner:
- Reads the Deployment with that name and namespace. If there is none (the resource is a StatefulSet, a DaemonSet, a Pod or has been deleted), the resource is skipped with no error.
- Takes the first container’s
resources.limitsandresources.requestsfor CPU and memory, and computesusage% = requests / limits × 100(0 when there is no limit). - Correlates the resource with the incidents of its namespace inside the window (below).
- Sets an urgency and one recommendation from what it knows: the requests/limits percentage and the incident correlation.
from and to are passed (RFC 3339), the window becomes their difference, still ending now.
Trend and Urgency
Urgency has only these two values: without history nothing is ever urgent. The dashboard’s capacity banner lists the resources with Urgency: "plan".
Response Fields
The endpoint returnskind: CapacityForecastList with one CapacityForecast per resource. Field names are the Go names (no JSON renaming), except the nested resource reference:
CapacityForecast
CapacityForecast
ResourceTrend
ResourceTrend
ForecastResult
ForecastResult
IncidentResourceCorrelation
IncidentResourceCorrelation
Correlation with Incidents
Urgency to plan and, when requests are within limits, sets the recommendation.
Recommendation Generation
The first matching rule wins:How to Use
viewer role (GET only). Each call lists every Issue in scope and reads one Deployment plus the namespace’s Issues per resource, so keep the calls occasional on large clusters.
Noise Reducer
The Noise Reducer runs inside the Anomaly controller. It is consulted only for an anomaly that would open a new Issue: an anomaly whose resource already has an active Issue is correlated to it first, and one that arrives within the resolution cooldown (spec.aiops.resolutionCooldownMinutes of the Instance the anomaly came from, default 10, 0 turns it off) after a resolution is suppressed before the Noise Reducer is reached.
The four checks run in order and the first one that fires wins. A suppressed anomaly is marked status.correlated: true with the reason recorded as suppressed:<reason>, no Issue is created, and chatcli_operator_anomalies_processed_total{result="suppressed_noise"} is incremented. If a check fails (for example, a List error), the anomaly is not suppressed. Resources are matched by name within the anomaly’s namespace.
Strategy 1: Repetitive Suppression
Strategy 2: Seasonal Patterns
Every anomaly that is not suppressed is recorded as a seasonal occurrence. Patterns are stored per namespace in the ConfigMapchatcli-seasonal-patterns, key patterns, as a JSON array.
SeasonalPattern (JSON):
Algorithm:
TZ on the operator; the image ships tzdata). The hour distance does not wrap around midnight (23:00 and 00:00 are 23 hours apart).
Strategy 3: Flap Detection
Strategy 4: Alert Fatigue Scoring
The recency reference is the last item of the namespace’s unfiltered Anomaly list, not the newest anomaly of the resource, so
recency_score can reflect another resource. The threshold > 80 is fixed and not configurable.Cost Tracker
The Cost Tracker books the LLM spend of every incident from the usage the server reports, and serves the aggregate over the REST API.LLM Costs per Provider
Every analysis (AnalyzeIssue, AIInsight controller) and every agentic remediation step (AgenticStep, Remediation controller) is booked from the token usage the server reports on the reply, attributed to the provider and model that actually answered (the server’s fallback chain may route elsewhere). When the reply carries no usage (a server older than the usage fields), analyses fall back to characters divided by four for both input and output, and agentic steps book zero input tokens and the reasoning length divided by four as output.
The price per token is resolved in this order:
- an entry for the provider in the
chatcli-cost-configConfigMap of the Issue’s namespace (the cluster operator’s word); - the shared pricing engine (
llm/pricing), the same per-model tables, overrides and subscription rules the CLI’s/costuses (including aCHATCLI_MODEL_PRICINGoverride set on the operator process, for example through the chart’sextraEnv), so a ledger here and a session there agree on the price of the same call; - a per-provider default for a model the engine does not know (for example
CLAUDEAI15,OPENAI30,OPENROUTER8,DEVIN1/$3 per million input/output tokens).
cost_usd / cost_known on the reply itself (see server mode); the ledger does not use those fields.
Cost Configuration
Prices are overridable per provider via thechatcli-cost-config ConfigMap, key pricing, one entry per provider name exactly as the server reports it (CLAUDEAI, OPENAI, …). The tracker looks it up in the namespace of the Issue being booked, the same namespace as its ledger, so create one in every namespace whose incidents you want to price differently:
An entry applies to every model of that provider; there is no per-model key. Without the ConfigMap, for a provider it does not list, or when the
pricing value is not valid JSON, the pricing engine’s per-model rates apply. The ConfigMap is read on every booking, so an update applies from the next call on; calls already booked keep the price they were booked at.IncidentCost
One ledger entry per incident, written by the AIInsight and Remediation controllers on every LLM call:Each call is priced once, at the rates of the provider and model that answered it, and added to the entry’s total. A later call on another model (for example after a fallback) never reprices the earlier ones;
provider and model only name the latest call.CLAUDEAI override above):
usage on AnalyzeIssue and AgenticStep) is what gets booked, so the ledger reflects real token counts, not estimates.
Two calls booked at the same time in one namespace do not overwrite each other: the ledger update carries the version it read, and a write conflict is retried. A booking that still fails (the ConfigMap size limit, missing RBAC, a ledger that cannot be read or an entry that cannot be decoded) is logged by the controller (
Failed to book the analysis cost on the ledger or Failed to book the agentic step cost on the ledger); the analysis or step itself is not affected, and an unreadable entry is never overwritten with a fresh one.CostSummary
Aggregation over a period, served by the REST API (viewer role, GET only), as kind: CostSummary with the summary under spec:
from and to (RFC 3339) set an absolute period: with both, it is exactly [from, to]; with only from, it runs to now; with only to, it is the 30 days ending at to; with neither, it is the last 30 days. An entry is counted in full when its recordedAt (last booking) falls in the period; entries without recordedAt are always counted. Without namespace, the summary aggregates every chatcli-cost-ledger in the cluster. A namespace without a ledger returns zeros; any other failure to read a ledger returns 500 (failed to read the cost ledger: ...) instead of zeros.
totalLLMCost with your own incident economics if you need them.
Storage Architecture (ConfigMaps)
These subsystems persist their data in ConfigMaps in the namespace of the Issue or Anomaly they concern (not the operator namespace). The Capacity Planner stores nothing.Storage Format
One ConfigMap per namespace,chatcli-cost-ledger, one key per incident holding the IncidentCost JSON. The tracker creates it with the app.kubernetes.io/managed-by: chatcli-operator label, which is how the all-namespace summary finds every ledger:
Integrations
REST API
/api/v1/analytics/cost serves the ledger summary; /api/v1/analytics/capacity serves the capacity forecasts; /api/v1/analytics/summary exposes incident data.Web Dashboard
The Overview view has a capacity-warnings banner fed by
/analytics/capacity: it lists the resources with Urgency: "plan" and their recommendation. The dashboard does not show the cost ledger.Grafana
The
remediation-stats.json dashboard covers remediation success by action type. No bundled dashboard covers capacity or LLM cost.AIOps Platform
Complete AIOps pipeline architecture and how these subsystems integrate.