> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Capacity Planning and Cost Attribution

> On-demand capacity forecasts per incident resource, anomaly noise reduction, and LLM cost tracking per incident.

The ChatCLI AIOps platform includes three subsystems that sit next to the incident pipeline: **Capacity Planner** (an on-demand forecast for the resources behind your incidents), **Noise Reducer** (suppression of repetitive, seasonal, flapping and high-fatigue anomalies), and **Cost Tracker** (LLM spend per incident, served over the REST API).

## Capacity Planner

The Capacity Planner produces a CPU and memory report for each Deployment that appears as the resource of an Issue. It is not a background job: it runs only when `GET /api/v1/analytics/capacity` is called, and it keeps no history of its own.

<Warning>
  The planner does not read real usage and collects no usage history. It reads no metrics-server and no Prometheus data (`PROMETHEUS_URL` is not used here). "Current usage" is the **requests** of the Deployment's first container (`UsageSource: "requests"`), and the percentage is requests divided by limits. Without history there is no trend to fit, so the response says `Trend.Direction: "insufficient_history"` and `HistoryAvailable: false`, and it never projects an exhaustion date. Treat the output as a requests-versus-limits check plus an incident count, not as a usage forecast.
</Warning>

### What It Computes

For every distinct resource (`kind/namespace/name`) among the Issues listed (all Issues, or those in `?namespace=`), the planner:

1. Reads the **Deployment** with that name and namespace. If there is none (the resource is a StatefulSet, a DaemonSet, a Pod or has been deleted), the resource is skipped with no error.
2. Takes the first container's `resources.limits` and `resources.requests` for CPU and memory, and computes `usage% = requests / limits × 100` (0 when there is no limit).
3. Correlates the resource with the incidents of its namespace inside the window (below).
4. Sets an **urgency** and one recommendation from what it knows: the requests/limits percentage and the incident correlation.

The window defaults to **7 days**. When both `from` and `to` are passed (RFC 3339), the window becomes their difference, still ending now.

### Trend and Urgency

```text theme={"system"}
UsageSource      = "requests"
HistoryAvailable = false
Trend.Direction  = "insufficient_history"   (CPUTrendPerDay = MemoryTrendPerDay = 0)
Forecast.*ExhaustionDate = null, Forecast.DaysUntil*Exhaustion = 0

Urgency = "plan" if CPU usage% >= 80 or memory usage% >= 80 or ResourceIsBottleneck
          "none" otherwise
```

`Urgency` has only these two values: without history nothing is ever `urgent`. The dashboard's capacity banner lists the resources with `Urgency: "plan"`.

### Response Fields

The endpoint returns `kind: CapacityForecastList` with one `CapacityForecast` per resource. Field names are the Go names (no JSON renaming), except the nested resource reference:

<AccordionGroup>
  <Accordion title="CapacityForecast">
    | Field | Type | Description |
    | - | - | - |
    | `Resource` | `{kind, name, namespace}` | The Issue's resource |
    | `CurrentUsage.CPUMillicores` / `.MemoryBytes` | `int64` | First container's requests |
    | `Limits.CPUMillicores` / `.MemoryBytes` | `int64` | First container's limits |
    | `UsagePercentage.CPU` / `.Memory` | `float64` | Requests divided by limits × 100 |
    | `Trend` | `ResourceTrend` | See below |
    | `Forecast` | `ForecastResult` | See below |
    | `IncidentCorrelation` | `IncidentResourceCorrelation` | See below |
    | `UsageSource` | `string` | Where `CurrentUsage` comes from: always `requests` |
    | `HistoryAvailable` | `bool` | Always `false`: no usage samples are collected |
    | `Urgency` | `string` | `plan` or `none` (see above) |
  </Accordion>

  <Accordion title="ResourceTrend">
    | Field | Type | Description |
    | - | - | - |
    | `CPUTrendPerDay` | `float64` | Always `0` (no history to fit) |
    | `MemoryTrendPerDay` | `float64` | Always `0` (no history to fit) |
    | `Direction` | `string` | `insufficient_history` |
  </Accordion>

  <Accordion title="ForecastResult">
    | Field | Type | Description |
    | - | - | - |
    | `CPUExhaustionDate` | `time` or `null` | Always `null`: no exhaustion date is projected |
    | `MemoryExhaustionDate` | `time` or `null` | Always `null` |
    | `DaysUntilCPUExhaustion` | `int` | Always `0` |
    | `DaysUntilMemoryExhaustion` | `int` | Always `0` |
    | `Recommendation` | `string` | One recommendation (table below) |
  </Accordion>

  <Accordion title="IncidentResourceCorrelation">
    | Field | Type | Description |
    | - | - | - |
    | `IncidentsInWindow` | `int` | Issues created in the namespace inside the window |
    | `ResourceRelatedIncidents` | `int` | Of those, Issues whose resource name matches |
    | `ResourceIsBottleneck` | `bool` | `true` when `ResourceRelatedIncidents > 2` |
  </Accordion>
</AccordionGroup>

### Correlation with Incidents

```text theme={"system"}
For each Issue in the resource's namespace created inside the window (any state):
  IncidentsInWindow++
  If issue.spec.resource.name == resource name: ResourceRelatedIncidents++

ResourceIsBottleneck = ResourceRelatedIncidents > 2
```

A bottleneck raises `Urgency` to `plan` and, when requests are within limits, sets the recommendation.

### Recommendation Generation

The first matching rule wins:

| Condition | Recommendation |
| - | - |
| CPU or memory requests above 80% of limits | "Requests are above 80% of limits. Consider increasing limits proactively." |
| `ResourceIsBottleneck` | "Repeated incidents on this resource. Review its capacity." |
| Otherwise | "No usage history is collected, so no trend or exhaustion date can be projected; requests are within limits." |

### How to Use

```bash theme={"system"}
# all namespaces, last 7 days (default)
curl -H "X-API-Key: $KEY" "http://localhost:8090/api/v1/analytics/capacity"

# one namespace, 30-day window (only the difference between from and to is used)
curl -H "X-API-Key: $KEY" \
  "http://localhost:8090/api/v1/analytics/capacity?namespace=production&from=2026-08-28T00:00:00Z&to=2026-09-27T00:00:00Z"
```

```json theme={"system"}
{
  "apiVersion": "v1",
  "kind": "CapacityForecastList",
  "metadata": {"totalCount": 1, "page": 1, "pageSize": 1},
  "items": [
    {
      "Resource": {"kind": "Deployment", "name": "api-gateway", "namespace": "production"},
      "CurrentUsage": {"CPUMillicores": 900, "MemoryBytes": 536870912},
      "Limits": {"CPUMillicores": 1000, "MemoryBytes": 1073741824},
      "UsagePercentage": {"CPU": 90, "Memory": 50},
      "Trend": {"CPUTrendPerDay": 0, "MemoryTrendPerDay": 0, "Direction": "insufficient_history"},
      "Forecast": {
        "CPUExhaustionDate": null,
        "MemoryExhaustionDate": null,
        "DaysUntilCPUExhaustion": 0,
        "DaysUntilMemoryExhaustion": 0,
        "Recommendation": "Requests are above 80% of limits. Consider increasing limits proactively."
      },
      "IncidentCorrelation": {"IncidentsInWindow": 4, "ResourceRelatedIncidents": 3, "ResourceIsBottleneck": true},
      "UsageSource": "requests",
      "HistoryAvailable": false,
      "Urgency": "plan"
    }
  ]
}
```

The endpoint needs the `viewer` role (GET only). Each call lists every Issue in scope and reads one Deployment plus the namespace's Issues per resource, so keep the calls occasional on large clusters.

## Noise Reducer

The Noise Reducer runs inside the Anomaly controller. It is consulted only for an anomaly that would open a **new** Issue: an anomaly whose resource already has an active Issue is correlated to it first, and one that arrives within the resolution cooldown (`spec.aiops.resolutionCooldownMinutes` of the Instance the anomaly came from, default 10, `0` turns it off) after a resolution is suppressed before the Noise Reducer is reached.

The four checks run in order and the first one that fires wins. A suppressed anomaly is marked `status.correlated: true` with the reason recorded as `suppressed:<reason>`, no Issue is created, and `chatcli_operator_anomalies_processed_total{result="suppressed_noise"}` is incremented. If a check fails (for example, a List error), the anomaly is not suppressed. Resources are matched by **name** within the anomaly's namespace.

### Strategy 1: Repetitive Suppression

```text theme={"system"}
Suppression condition:
  - 5 or more Anomalies in the namespace created in the last 1 hour
    with the same signalType, resource kind and resource name
    (the incoming anomaly counts)
  - and no Issue for that resource name is in Analyzing or Remediating

Result:
  - reason "Repetitive: N identical anomalies in last hour with no state change"
```

Severity is not part of the match.

### Strategy 2: Seasonal Patterns

Every anomaly that is not suppressed is recorded as a seasonal occurrence. Patterns are stored per namespace in the ConfigMap `chatcli-seasonal-patterns`, key `patterns`, as a JSON array.

**SeasonalPattern (JSON):**

| Field | Type | Description |
| - | - | - |
| `signalType` | `string` | Anomaly signal type (for example `pod_restart`) |
| `resource` | `{kind, name, namespace}` | Resource of the first anomaly recorded for the pattern |
| `hourOfDay` | `int` | Hour of the day (0–23) when the pattern was first recorded |
| `dayOfWeek` | `string` | Weekday name (`Monday`, ...) |
| `occurrences` | `int` | Anomalies recorded in this slot |
| `firstSeen` / `lastSeen` | `time` | First and last recording |

**Algorithm:**

```text theme={"system"}
Recording (anomaly not suppressed):
  Find a pattern with the same signalType and resource name,
  the same weekday and |current hour - hourOfDay| <= 1
  If found: occurrences++, lastSeen = now
  Else:     append a new pattern with occurrences = 1

Suppression check:
  Same match (signalType, resource name, weekday, hour within ±1)
  AND occurrences >= 3  →  suppress
  reason "Seasonal: pattern detected on <weekday> ~<hour>:00 (<n> occurrences)"
```

Hours and weekdays come from the operator pod's clock (UTC unless you set `TZ` on the operator; the image ships tzdata). The hour distance does not wrap around midnight (23:00 and 00:00 are 23 hours apart).

<Warning>
  `occurrences` counts anomalies, not distinct weeks. Three anomalies for the same signal and resource in the same hour slot of one day are enough to suppress later ones in that slot, on that day and on the same weekday every week after. There is no confidence score and no pruning: patterns stay in the ConfigMap until you edit or delete it.
</Warning>

### Strategy 3: Flap Detection

```text theme={"system"}
Flap condition:
  - 3 or more Issues for the same resource name in the namespace,
    created in the last 24 hours and currently in state Resolved

Result:
  - the new anomaly is suppressed
  - reason "Flapping: N resolved→detected cycles in 24h for this resource"
```

There is no separate flapping flag, hold-off timer or consolidated alert: each new anomaly is checked the same way, and suppression stops once fewer than three resolved Issues remain inside the 24-hour window.

### Strategy 4: Alert Fatigue Scoring

```text theme={"system"}
Over the namespace's Anomalies of the same resource name in the last 24 hours:
  total      = number of anomalies
  correlated = anomalies with status.correlated = true
               (grouped into an Issue, or suppressed earlier)

  volume_score  = min(50, total * 5)
  resolve_score = int(correlated / total * 30)
  recency_score = 20, or 10 if the reference anomaly is older than 1h,
                  or 5 if older than 6h

  fatigue_score = min(100, volume_score + resolve_score + recency_score)

Suppress when fatigue_score > 80.
```

In practice, as few as 7 anomalies for a resource in 24 hours, all already correlated, with a recent reference anomaly, cross the threshold (35 + 30 + 20 = 85), and further anomalies for it are suppressed until the count drops.

<Note>
  The recency reference is the last item of the namespace's unfiltered Anomaly list, not the newest anomaly of the resource, so `recency_score` can reflect another resource. The threshold `> 80` is fixed and not configurable.
</Note>

## Cost Tracker

The Cost Tracker books the LLM spend of every incident from the usage the server reports, and serves the aggregate over the REST API.

### LLM Costs per Provider

Every analysis (`AnalyzeIssue`, AIInsight controller) and every agentic remediation step (`AgenticStep`, Remediation controller) is booked from the **token usage the server reports** on the reply, attributed to the provider and model that actually answered (the server's fallback chain may route elsewhere). When the reply carries no usage (a server older than the usage fields), analyses fall back to characters divided by four for both input and output, and agentic steps book zero input tokens and the reasoning length divided by four as output.

The price per token is resolved in this order:

1. an entry for the provider in the `chatcli-cost-config` ConfigMap of the **Issue's namespace** (the cluster operator's word);
2. the shared pricing engine (`llm/pricing`), the same per-model tables, overrides and subscription rules the CLI's `/cost` uses (including a `CHATCLI_MODEL_PRICING` override set on the operator process, for example through the chart's `extraEnv`), so a ledger here and a session there agree on the price of the same call;
3. a per-provider default for a model the engine does not know (for example `CLAUDEAI` $3/$15, `OPENAI` $10/$30, `OPENROUTER` $2/$8, `DEVIN` $0, anything else $1/\$3 per million input/output tokens).

Only input and output rates are applied: cache-read and cache-write discounts are not modeled on the ledger. The server also reports `cost_usd` / `cost_known` on the reply itself (see [server mode](/server/server-mode#reply-attribution-and-token-usage)); the ledger does not use those fields.

### Cost Configuration

Prices are overridable per provider via the `chatcli-cost-config` ConfigMap, key `pricing`, one entry per provider name exactly as the server reports it (`CLAUDEAI`, `OPENAI`, ...). The tracker looks it up **in the namespace of the Issue being booked**, the same namespace as its ledger, so create one in every namespace whose incidents you want to price differently:

```yaml theme={"system"}
apiVersion: v1
kind: ConfigMap
metadata:
  name: chatcli-cost-config
  namespace: production        # the namespace where the Issues live
data:
  pricing: |
    {
      "CLAUDEAI": {"InputPerMillion": 3.00, "OutputPerMillion": 15.00},
      "OPENAI":   {"InputPerMillion": 1.25, "OutputPerMillion": 10.00}
    }
```

<Note>
  An entry applies to every model of that provider; there is no per-model key. Without the ConfigMap, for a provider it does not list, or when the `pricing` value is not valid JSON, the pricing engine's per-model rates apply. The ConfigMap is read on every booking, so an update applies from the next call on; calls already booked keep the price they were booked at.
</Note>

### IncidentCost

One ledger entry per incident, written by the AIInsight and Remediation controllers on every LLM call:

| Field | Type | Description |
| - | - | - |
| `issueName` | `string` | Incident name |
| `llmCosts.totalInputTokens` | `int64` | Input tokens across all calls |
| `llmCosts.totalOutputTokens` | `int64` | Output tokens across all calls |
| `llmCosts.analysisCalls` | `int32` | `AnalyzeIssue` calls |
| `llmCosts.agenticSteps` | `int32` | `AgenticStep` calls |
| `llmCosts.estimatedCostUSD` | `float64` | Sum of every call's cost, each priced at its own provider and model rate (pricing engine or ConfigMap override) when it was booked |
| `llmCosts.provider` / `llmCosts.model` | `string` | Provider and model of the **latest** call |
| `engineerTimeSavedMinutes` | `float64` | Reserved for future estimation, `0` today |
| `totalCostUSD` | `float64` | Equals `llmCosts.estimatedCostUSD` |
| `recordedAt` | `time` | When the entry was last booked |

<Note>
  Each call is priced once, at the rates of the provider and model that answered it, and added to the entry's total. A later call on another model (for example after a fallback) never reprices the earlier ones; `provider` and `model` only name the latest call.
</Note>

**Calculation example** (with the `CLAUDEAI` override above):

```text theme={"system"}
Incident: CrashLoopBackOff on api-server
  - Provider: CLAUDEAI
  - LLM calls: 2 (AnalyzeIssue + 1 AgenticStep)
  - Input tokens: 8,500 | Output tokens: 2,200
  - Input cost:  8,500 / 1,000,000 * $3.00  = $0.0255
  - Output cost: 2,200 / 1,000,000 * $15.00 = $0.0330
  - estimatedCostUSD: $0.0585
```

The usage the server reports on each response (`usage` on `AnalyzeIssue` and `AgenticStep`) is what gets booked, so the ledger reflects real token counts, not estimates.

<Note>
  Two calls booked at the same time in one namespace do not overwrite each other: the ledger update carries the version it read, and a write conflict is retried. A booking that still fails (the ConfigMap size limit, missing RBAC, a ledger that cannot be read or an entry that cannot be decoded) is logged by the controller (`Failed to book the analysis cost on the ledger` or `Failed to book the agentic step cost on the ledger`); the analysis or step itself is not affected, and an unreadable entry is never overwritten with a fresh one.
</Note>

### CostSummary

Aggregation over a period, served by the REST API (`viewer` role, GET only), as `kind: CostSummary` with the summary under `spec`:

| Field | Type | Description |
| - | - | - |
| `periodStart` / `periodEnd` | `time` | Period boundaries (see below) |
| `totalLLMCost` | `float64` | Sum of the `estimatedCostUSD` of every entry in the window |
| `incidentCount` | `int` | Ledger entries in the window |
| `costPerIncident` | `float64` | `totalLLMCost / incidentCount` (`0` when there are none) |

`from` and `to` (RFC 3339) set an **absolute period**: with both, it is exactly `[from, to]`; with only `from`, it runs to now; with only `to`, it is the 30 days ending at `to`; with neither, it is the **last 30 days**. An entry is counted in full when its `recordedAt` (last booking) falls in the period; entries without `recordedAt` are always counted. Without `namespace`, the summary aggregates every `chatcli-cost-ledger` in the cluster. A namespace without a ledger returns zeros; any other failure to read a ledger returns `500` (`failed to read the cost ledger: ...`) instead of zeros.

```bash theme={"system"}
# last 30 days (default), all namespaces
curl -H "X-API-Key: $KEY" "http://localhost:8090/api/v1/analytics/cost"

# one namespace, 1 to 25 September
curl -H "X-API-Key: $KEY" "http://localhost:8090/api/v1/analytics/cost?namespace=production&from=2026-09-01T00:00:00Z&to=2026-09-25T00:00:00Z"
```

```json theme={"system"}
{
  "apiVersion": "v1",
  "kind": "CostSummary",
  "spec": {
    "periodStart": "2026-09-01T00:00:00Z",
    "periodEnd": "2026-09-25T00:00:00Z",
    "totalLLMCost": 1.84,
    "incidentCount": 23,
    "costPerIncident": 0.08
  }
}
```

Downtime cost and return-on-investment figures are not computed: the platform has no trustworthy source for revenue per minute or engineer hourly rates, and numbers built on assumed constants would be misleading. Combine `totalLLMCost` with your own incident economics if you need them.

## Storage Architecture (ConfigMaps)

These subsystems persist their data in ConfigMaps in the **namespace of the Issue or Anomaly** they concern (not the operator namespace). The Capacity Planner stores nothing.

| ConfigMap | Written by | Data | Retention |
| - | - | - | - |
| `chatcli-seasonal-patterns` | Noise Reducer | Key `patterns`: JSON array of seasonal patterns | Indefinite (never pruned) |
| `chatcli-pattern-store` | Remediation controller (resolution learning, not the Noise Reducer) | One key per fingerprint of signal type, resource kind and severity: successful actions, success/failure counts, average resolution time, confidence boost | Indefinite (never pruned) |
| `chatcli-cost-ledger` | Cost Tracker | One key per incident with the `IncidentCost` JSON | Indefinite (never pruned) |
| `chatcli-cost-config` | You | Per-provider pricing override (`pricing` key) | User-managed |

<Warning>
  ConfigMaps have a 1MB limit in Kubernetes, roughly 3,000 ledger entries. The operator does not compact the ledger or the seasonal patterns: on a namespace with that volume, prune old keys periodically, or new bookings fail (each failure is logged, see above).
</Warning>

### Storage Format

One ConfigMap per namespace, `chatcli-cost-ledger`, one key per incident holding the `IncidentCost` JSON. The tracker creates it with the `app.kubernetes.io/managed-by: chatcli-operator` label, which is how the all-namespace summary finds every ledger:

```yaml theme={"system"}
apiVersion: v1
kind: ConfigMap
metadata:
  name: chatcli-cost-ledger
  namespace: production
  labels:
    app.kubernetes.io/managed-by: chatcli-operator
data:
  api-gateway-oom-7f3a: |
    {
      "issueName": "api-gateway-oom-7f3a",
      "llmCosts": {
        "totalInputTokens": 8500,
        "totalOutputTokens": 2200,
        "analysisCalls": 1,
        "agenticSteps": 1,
        "estimatedCostUSD": 0.0585,
        "provider": "CLAUDEAI",
        "model": "claude-sonnet-5"
      },
      "engineerTimeSavedMinutes": 0,
      "totalCostUSD": 0.0585,
      "recordedAt": "2026-09-25T14:13:00Z"
    }
```

Entries are never compacted or expired by the operator: delete keys, or the ConfigMap, to reset a namespace's ledger.

## Integrations

<CardGroup cols={2}>
  <Card title="REST API" icon="code" href="/reference/api/overview">
    `/api/v1/analytics/cost` serves the ledger summary; `/api/v1/analytics/capacity` serves the capacity forecasts; `/api/v1/analytics/summary` exposes incident data.
  </Card>

  <Card title="Web Dashboard" icon="chart-mixed" href="/kubernetes/aiops/web-dashboard">
    The Overview view has a capacity-warnings banner fed by `/analytics/capacity`: it lists the resources with `Urgency: "plan"` and their recommendation. The dashboard does not show the cost ledger.
  </Card>

  <Card title="Grafana" icon="chart-area" href="/kubernetes/aiops/web-dashboard#grafana-dashboards">
    The `remediation-stats.json` dashboard covers remediation success by action type. No bundled dashboard covers capacity or LLM cost.
  </Card>

  <Card title="AIOps Platform" icon="brain" href="/kubernetes/aiops-platform">
    Complete AIOps pipeline architecture and how these subsystems integrate.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.