> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Web Dashboard and Grafana

> The operator's built-in web dashboard (API keys, publishing it behind an Ingress) and pre-configured Grafana dashboards for observability.

The ChatCLI AIOps platform includes a **built-in web dashboard** (self-contained SPA served by the operator) and **4 Grafana dashboards** (JSON in the repository) for observability of the autonomous operations pipeline.

## Web Dashboard

### Overview

The Web Dashboard is a **Single Page Application** embedded directly in the operator binary via Go `embed.FS` -- it does not require Node.js, npm, or any separate frontend build.

| Characteristic | Detail |
| - | - |
| **Technology** | HTML/CSS/JS vanilla in a single `index.html` (zero external dependencies) |
| **Packaging** | Go `embed.FS` -- compiled into the binary |
| **Port** | `8090` on the operator (chart `api.port`, env `CHATCLI_AIOPS_PORT`), Service `chatcli-operator` |
| **Theme** | Same visual language as the ChatCLI live dashboard: dark terminal palette, a light palette, monospace type, KPI tiles, segmented tab bar. A header switcher picks **System / Dark / Light**; System follows the OS preference |
| **Language** | Header switcher with **English** and **Português (BR)**; every label, table header, empty state, toast and confirmation prompt is translated in place, without a reload. Dates and numbers follow the selected language |
| **Responsive** | Adapts to desktop, tablet, and mobile |
| **Auto-refresh** | Configurable interval: Off, 10s, 30s (default), 1m, 2m, 5m — with manual refresh button |
| **Draft guard** | A form being filled in (post-mortem feedback, approval reason) is never wiped by a refresh: the countdown holds while a field is focused or has unsaved text, the header shows *paused while editing*, and the values are restored after any other re-render. Expanded rows are keyed by object name, so an open post-mortem stays open when the list reorders or the time range changes |
| **Sortable tables** | Click any column header to sort (▲ ascending / ▼ descending) |
| **Stable ordering** | Lists maintain consistent order between refreshes (incidents, approvals and audit newest first, SLOs by name) |
| **Authentication** | Same API key as the REST API (header `X-API-Key`); the role of the key decides which actions work |
| **URL** | `http://&lt;operator-host&gt;:8090/` |

<Note>
  The dashboard consumes the same REST API documented in the [API Reference](/reference/api/overview). Every operation available in the dashboard (acknowledge, snooze, resolve, approve, reject, post-mortem review, close, feedback and human-action acknowledgement) is an authenticated REST call and needs at least the `operator` role; a `viewer` key sees everything but its actions return `403`.
</Note>

### Theme and language

Both switchers live in the header, next to the auto-refresh control, and are keyboard accessible. The choice is stored in the browser and applied before the first paint, so a reload never flashes the wrong theme.

| Preference | Values | Storage key | Default |
| - | - | - | - |
| **Theme** | `system`, `dark`, `light` | `localStorage` `chatcli_dash_theme` | `system` — follows `prefers-color-scheme` |
| **Language** | `en`, `pt-BR` | `localStorage` `chatcli_dash_lang` | From the browser: `pt*` → `pt-BR`, anything else → `en` |

The query parameters `?theme=` and `?lang=` set the preference once and persist it, so a bookmark such as `http://localhost:8090/?lang=pt-BR&theme=dark` opens the dashboard in Portuguese and dark for every visit that follows. The `<html lang>` attribute tracks the language, and values that come from the API (severity, incident and plan states, approval decisions) keep their raw value in the markup while the visible label is translated.

<Tip>
  Product names stay in English in both languages: Issue, SLO, Runbook, PostMortem, AIInsight and RemediationPlan are the names of the CRDs the operator manages.
</Tip>

### Architecture

```text theme={"system"}
+-------------------------------------------------------------+
|                    Operator Binary                           |
|                                                              |
|  +--------------------+  +-------------------------------+   |
|  |  embed.FS          |  |  HTTP Server (:8090)          |   |
|  |  +-- index.html    |  |  +-- /          -> SPA        |   |
|  |   (HTML, CSS, JS   |  |  +-- /api/v1/   -> REST API   |   |
|  |    and i18n in one |  |  +-- /healthz   -> Health     |   |
|  |    file)           |  |  +-- /readyz    -> Ready      |   |
|  +--------------------+  +-------------------------------+   |
|                                                              |
|  +------------------------------------------------------+    |
|  |  controller-runtime client (cached reads)             |    |
|  |  Issues, AIInsights, Plans, PostMortems, SLOs, ...    |    |
|  +------------------------------------------------------+    |
+-------------------------------------------------------------+
```

The page itself (`/`) is served without authentication; every data call it makes goes to `/api/v1/` with the API key. On first load the dashboard asks for the key, checks it against `/api/v1/incidents`, and keeps it in the browser's `localStorage` (`chatcli_api_key`); **Change API key** in the header clears it.

<Note>
  The dashboard is served by the **operator**, on the `api` port (8090) of the Service `chatcli-operator` in the operator namespace, not by the ChatCLI Instance's Service. Every operator replica serves the dashboard and the API, so a Service with several replicas can route to any of them.
</Note>

### Dashboard Views

The dashboard has **11 views** accessible via tab navigation. A time-range filter in the header (**All time**, or **From**/**To**) is passed as `from`/`to` to the lists that support it.

<AccordionGroup>
  <Accordion title="1. Overview">
    Platform overview with aggregated metrics, built from `/analytics/summary`, `/analytics/compliance`, `/analytics/capacity`, `/analytics/remediation-stats`, `/analytics/mttd`, `/analytics/mttr` and the incident list.

    **Components:**

    | Component | Description |
    | - | - |
    | **Stats Cards** | 8 KPI tiles: Active Issues, **Contained (needs human)**, Resolved, Remediations (Issues resolved by remediation / Issues remediated), Success Rate, PostMortems (with the count pending human action), Pending Approvals, **Chaos drills (excluded from MTTD/MTTR)**. The Contained tile gets a violet outline when issues are in that state. |
    | **Capacity Warnings** | Banner listing the capacity forecasts with `Urgency: plan` (requests at 80% or more of limits, or repeated incidents on the resource), with recommendations |
    | **Issues by Severity** | Pie chart of incident distribution by severity (Critical/High/Medium/Low) |
    | **Platform Resources** | Totals: Issues, Remediations, PostMortems, Runbooks, SLOs, Pending Approvals, average risk score |
    | **Remediation Activity** | The last incidents with remediation attempts, as a timeline. A toggle flips between newest first (default) and oldest first; the choice is remembered in `localStorage` (`chatcli_dash_remediationSort`) |
    | **Remediation by Strategy** | Horizontal bars with the success rate per strategy (e.g., RestartDeployment 93%, RollbackDeployment 80%) |
    | **Compliance & SLA** | SLA compliance percentage, response/resolution SLA violations, approvals (auto/manual) |
    | **MTTD & MTTR Trend** | MTTD and MTTR averages and the MTTR trend of the last 10 days |
    | **Recent Incidents** | The most recent incidents with state, severity, and age, plus a filter bar mirroring the Incidents tab: severity, state, resource kind, chaos drills and namespace, with a Reset button. Filtering is client-side over the incidents already loaded, re-applied on every auto-refresh, and the selection is remembered in `localStorage` (`chatcli_dash_recentFilters`). Every column header sorts the panel (name, severity, state, resource, namespace, age); severity sorts by rank, not by label, and the sort is kept per panel across refreshes |

    The panels can be rearranged by dragging their headers; the order is remembered in `localStorage` (`panelOrder_overviewPanels`).
  </Accordion>

  <Accordion title="2. Incidents">
    Interactive table of the incidents with filters and actions.

    **Features:**

    | Feature | Description |
    | - | - |
    | **Filters** | Severity, state, resource kind, chaos drills (hidden by default), namespace, plus the header time range. Click **Apply** |
    | **Table** | Columns: Name (with `CHAOS` and `NEEDS HUMAN` badges), Severity, State, Resource, Namespace, Risk, Age, actions. Newest first; shows the 100 most recent matches (no pagination) |
    | **Sorting** | Click column header to sort (asc/desc) |
    | **Expansion** | Click row to expand details: description, resource, signal type, source, risk score, remediation attempts, resolution, detection/resolution time |
    | **AI Insight Preview** | Expanded rows load the AI analysis inline with confidence badge, provider/model, analysis summary (first 500 chars), and top 3 recommendations |
    | **Ack** | Acknowledges the incident (not shown for Resolved or Contained) |
    | **Snooze** | Asks for a duration such as `30m`, `1h` or `2h` (default `1h`) (not shown for Resolved or Contained) |
    | **Resolve** | Shown for Escalated and Contained incidents; asks for an optional resolution note and marks the Issue `Resolved` |

    Ack, Snooze and Resolve need the `operator` role. **Ack** stops the incident's escalation at its current level (no further level, no repeat) and stamps who acknowledged on the EscalationPolicy status. **Snooze** holds every notification except `Resolved`, and the escalation, until the snooze ends; a held escalation page goes out then and the level's timer restarts. See [acknowledgement and snooze](/kubernetes/aiops/notifications#acknowledgement-and-stopping-an-escalation).

    **Severity badges (dark theme):**

    | Severity | Color |
    | - | - |
    | Critical | Red (#E06C75) |
    | High | Orange (#FF8C69) |
    | Medium | Amber (#FFB454) |
    | Low | Green (#7FD962) |

    **State badges (dark theme):**

    | State | Color |
    | - | - |
    | Detected | Red (#E06C75) |
    | Analyzing | Amber (#FFB454) |
    | Remediating | Amber (#FFB454) |
    | **Contained** | **Violet (#C79BFF) with a pulse animation** — a human action is owed |
    | Resolved | Green (#7FD962) |
    | Escalated | Orange (#FF8C69) |
    | Failed | Red (#E06C75) |
  </Accordion>

  <Accordion title="3. SLOs">
    One card per SLO, sorted by name, plus an **Active SLO Alerts** table listing the SLOs below target.

    **Components per SLO:**

    | Component | Description |
    | - | - |
    | **Card** | Service (or SLO name), SLI type, window, current value vs. target |
    | **State** | Badge from the SLO state the API derives (`Healthy`, `AtRisk`, `Breached`); before the SLO's first calculation it shows `Healthy` or `Firing` from current vs. target |
    | **Error Budget Remaining** | Bar with the remaining error budget: green (>50%), amber (>20%), red otherwise |
    | **Burn Rate Chips** | 1h, 6h, 24h, 72h chips with the burn rates the SLO controller computes (`burnRate1h` … `burnRate72h`): red above 14x, amber above 1x |
  </Accordion>

  <Accordion title="4. Approvals">
    Approval requests, newest first, with approve/reject actions. The tab badge shows the number of pending requests.

    **Features:**

    | Feature | Description |
    | - | - |
    | **Filter** | State (Pending, Approved, Rejected, Expired) |
    | **Context** | Each card shows: name, namespace, requested by, requested action (the first one), resource (the Issue), reason (the policy that raised it, e.g. `decision-engine`) |
    | **Your name** | Required field on pending cards: the approver name recorded on the decision, next to the identity of the API key. It is remembered in the browser (`localStorage` `chatcli_approver`) for the next decision |
    | **Decision reason** | Optional text field on pending cards |
    | **Quorum progress** | On requests that need more than one approver: "*n* of *required* required approvals" |
    | **Approve / Reject** | Buttons on pending cards; require the `operator` role. Each click records one decision on the request; the list reloads so the operator's verdict (quorum, change window) shows up |
    | **Decided cards** | Show approver, reason and decision time |

    <Note>
      A decision made here is the same as the REST call: it is stored in `status.decisions` as `<your name> (api-key: <key identity>)`, and the approval controller then evaluates quorum and change window. A quorum counts each API key once, so each approver needs their own key; a request that is no longer pending, or that your key already decided, returns an error. See [Approval Workflow](/kubernetes/aiops/approval-workflow#via-rest-api). When the request carries the annotation `platform.chatcli.io/blast-risk-level`, the card shows a blast-radius badge colored by level.
    </Note>
  </Accordion>

  <Accordion title="5. AI Insights">
    View all AI-generated analyses to understand how the AI reasoned about each incident.

    **Features:**

    | Feature | Description |
    | - | - |
    | **Filter** | Filter by incident name to see insights for a specific issue |
    | **Table** | Columns: Incident, Provider, Model, Confidence, Recommendations, Actions, Generated |
    | **Confidence** | Color-coded confidence score: green (≥85%), amber (70-84%), orange/red (\<70%) |
    | **Expansion** | Click row to expand: full AI analysis text, recommendations list, suggested actions with parameters |
    | **Log Analysis** | The log analysis text the operator attached to the analysis, when present |
    | **Cascade Analysis** | The cross-service cascade analysis, when present |
    | **GitOps Context** | Helm/ArgoCD/Flux context at the time of analysis, when present |
    | **Blast Radius** | The blast-radius prediction for the suggested actions, when present |

    This view is essential when an incident is **escalated to human action** -- it shows what the AI found, why it recommended specific actions, and what enrichment data informed its analysis.

    **API endpoint:** [`GET /api/v1/aiinsights`](/reference/api/list-aiinsights)
  </Accordion>

  <Accordion title="6. Remediations">
    Track all remediation plans with execution details, both runbook-based and agentic.

    **Features:**

    | Feature | Description |
    | - | - |
    | **Filters** | State dropdown (Pending/Executing/Verifying/Completed/Failed/RolledBack), incident name filter |
    | **Table** | Columns: Name, Incident, Attempt, State, Mode (Runbook/Agentic), Actions/Steps, Started, Duration |
    | **Expansion** | Click row to expand: strategy, planned actions, result, start and completion time |
    | **Agentic details** | For agentic plans: step count shown in the table and the detail view |
    | **Duration** | Calculated from start to completion time |

    **Remediation modes explained:**

    * **Runbook mode**: Displays the action sequence from the runbook the plan was built from
    * **Agentic mode**: Shows the step count; use the [Get Remediation Plan API](/reference/api/get-remediation-detail) for the full AI conversation history

    **API endpoint:** [`GET /api/v1/remediations`](/reference/api/list-remediations)
  </Accordion>

  <Accordion title="7. Runbooks">
    View all runbooks — both manually created and generated by the AI.

    **Features:**

    | Feature | Description |
    | - | - |
    | **Table** | Columns: Name, Signal Type, Severity, Resource Kind, Steps, Max Attempts, Created |
    | **Signal badge** | Color-coded badge showing the trigger signal type (oom\_kill, pod\_not\_ready, deploy\_failing, etc.) |
    | **Expansion** | Click row to expand: full description, trigger match criteria, and ordered step list |
    | **Step details** | Each step shows: action type badge, description, and parameters as JSON |

    Non-agentic RemediationPlans are built from a runbook. When an Issue is analyzed and no runbook matching its trigger (signal type, severity, resource kind) is accepted, the operator generates one from the AI's suggested actions; after a successful agentic remediation it also saves the actions that worked as an `agentic-…` runbook. Generated runbooks carry the label `platform.chatcli.io/auto-generated: "true"` and are reused by the next similar incident instead of starting from scratch.

    **API endpoint:** [`GET /api/v1/runbooks`](/reference/api/list-runbooks)
  </Accordion>

  <Accordion title="8. PostMortems">
    Table of post-mortems with expandable details.

    **Features:**

    | Feature | Description |
    | - | - |
    | **Table** | Columns: Name (with `CHAOS` and `NEEDS HUMAN` badges), Severity, State (Open/InReview/Closed), Resource, Duration, Issue, Created, actions |
    | **Expansion** | Click to expand: summary, root cause, impact, timeline, executed actions |
    | **Lessons Learned** | Section with lessons learned (AI-generated) |
    | **Prevention Actions** | List of suggested preventive actions |
    | **Developer Feedback** | Inline form: your name (required), remediation accuracy (1-5 stars, required), root cause override, comments. Once submitted, displays the feedback with visual rating |
    | **Review** | Marks the post-mortem `InReview` |
    | **Close** | Marks it `Closed` |
    | **Ack Human Action** | Replaces Review and Close while the post-mortem requires human action (a contained incident); asks for an optional note, then clears the flag so Close appears |

    Review, Close, feedback and Ack Human Action need the `operator` role.
  </Accordion>

  <Accordion title="9. Clusters">
    Federation overview and one card per registered cluster.

    **Summary tiles** (from `/clusters/global-status`): Total Clusters, Healthy (connected, every node Ready), Degraded (connected, with the `Degraded` condition: some nodes not Ready), Offline.

    **Federation Panel:**

    | Component | Description |
    | - | - |
    | **Federation Status** | Connected/disconnected cluster counts and active local Issues |
    | **Cross-Cluster Correlations** | Correlated Issues with severity badge, signal type, CASCADE/ELEVATED flags and the number of clusters the correlation spans, one entry per correlation id (see [Federation](/kubernetes/aiops/federation#cross-cluster-correlation)) |

    **Components per cluster:**

    | Component | Description |
    | - | - |
    | **Card** | Display name, `Healthy` or `Offline` badge (from `status.connected`) |
    | **Details** | Region, environment, tier, nodes, namespaces, Kubernetes version, active issues, active remediations |
    | **Last Health Check** | Time since the last health check |

    **API endpoints:** [`GET /api/v1/clusters/global-status`](/reference/api/global-status), [`GET /api/v1/federation/status`](/reference/api/federation-status), [`GET /api/v1/federation/correlations`](/reference/api/federation-correlations)
  </Accordion>

  <Accordion title="10. Policies">
    Read-only view of the four policy kinds the controllers read: ApprovalPolicy, NotificationPolicy, EscalationPolicy and IncidentSLA. Policies are written through `kubectl` or GitOps, where they are reviewed; the tab shows what is in force, per namespace or across all.

    | Panel | Columns |
    | - | - |
    | **Approval Policies** | Name, namespace, enabled, rule count, approved and rejected totals |
    | **Notification Policies** | Name, namespace, channel count, rule count, age |
    | **Escalation Policies** | Name, namespace, enabled (and `default`), level count, severities |
    | **Incident SLAs** | Name, namespace, severity, response and resolution targets, violations, compliance |

    Every column sorts. **API endpoints:** [`GET /api/v1/policies/{kind}`](/reference/api/list-policies), [`GET /api/v1/policies/{kind}/{name}`](/reference/api/get-policy)
  </Accordion>

  <Accordion title="11. Audit">
    Searchable audit log (AuditEvent CRs) with export.

    **Features:**

    | Feature | Description |
    | - | - |
    | **Search** | Client-side text search over name, event type, resource, actor, correlation id and detail |
    | **Filters** | Event type (grouped: Issues, Remediation, Approvals, Operations, SLO/SLA, Clusters, System), severity, plus the header time range |
    | **Table** | Columns: Timestamp, Event Type, Severity, Actor, Resource, Correlation, Detail. Shows the 100 most recent matching events |
    | **Export JSON** | Downloads `/api/v1/audit/export` with the event type and severity filters (not the time range) as a JSON file. The `viewer` role is enough |
  </Accordion>
</AccordionGroup>

## Grafana Dashboards

The repository ships **4 Grafana dashboards** as JSON in [`deploy/grafana/`](https://github.com/diillson/chatcli/tree/main/deploy/grafana), ready for import. Panels built on `chatcli_operator_*` metrics need the operator's metrics endpoint scraped; panels on `chatcli_grpc_*`, `chatcli_llm_*`, `chatcli_session_*`, `chatcli_server_*` and `chatcli_watcher_*` need the ChatCLI server's metrics port (9090) scraped too.

### 1. AIOps Overview (`aiops-overview.json`)

| Row | Panels |
| - | - |
| **AIOps Platform Overview** | Active Issues, Total Issues (24h), MTTR (p50 of resolution time), Remediation Success Rate, Managed Instances, Notifications Sent (24h) |
| **Issues by Severity** | Issues Created Over Time (by severity), Issues by State |
| **Remediation Performance** | Remediation Actions by Type, Remediation Duration (p50/p95/p99), Issue Resolution Duration (p50/p95) |
| **Watcher & Node Health** | Watcher Targets Monitored, Pods Ready vs Desired, Collection Errors, Pod Restarts, Watcher Alerts by Type, Collection Duration |
| **LLM & AI Performance** | LLM Requests Rate, LLM Latency (p50/p95/p99), Tokens Used, LLM Error Rate |
| **gRPC & Server** | Server Uptime, gRPC Request Rate, In-Flight Requests, gRPC Latency (p95), gRPC Error Rate |
| **Chaos Engineering** | Chaos Experiments (24h), Chaos Success Rate, Recovery Time (p50), Pods Affected by Chaos |

### 2. SLO Burn Rate (`slo-burn-rate.json`)

Template variables `service` and `slo_name`.

| Row | Panels |
| - | - |
| **Error Budget Status** | Error Budget Remaining, Current SLI Value, Burn Rate (1h), SLO Violations (24h) |
| **Multi-Window Burn Rates** | Burn rate over the 1h, 6h, 24h and 72h windows, each with its threshold line (14.4x, 6x, 3x, 1x) |
| **SLA Compliance** | SLA Compliance %, SLA Response Time by Severity, SLA Violations Rate |
| **SLO Detail by Service** | Current SLI by Service, Error Budget by Service, SLO Violations by Service |
| **Burn Rate Analysis** | Burn rate across all windows, Burn Rate vs Threshold, Time Until Budget Exhaustion |
| **SLA Deep Dive** | SLA Response/Resolution Time Distribution, SLA Compliance Trend, Violations by Type (24h) |

### 3. Incident Timeline (`incident-timeline.json`)

| Row | Panels |
| - | - |
| **Incident Lifecycle** | Critical Issues Detected (24h), Escalated Issues, Resolved (24h), Anomalies Processed (24h), Issues from Correlation, Escalation Levels Reached |
| **Notification Delivery** | Notifications by Channel, Notification Failures |
| **Approval Workflow** | Approval Decisions (by mode and result), Approval Decision Latency |
| **Federation** | Connected Clusters, Cross-Cluster Issues, Cascade Detected |
| **Incident Intelligence** | Anomaly Processing Rate, Issues Created by Correlation, Active Issues Trend, Issue Resolution Duration Distribution |
| **Notification & Escalation Analytics** | Notification Success Rate, Notification Latency by Channel, Escalation Level Distribution, Notification Failures by Reason |
| **Federation & Multi-Cluster** | Clusters by Status, Cross-Cluster Issues Rate, Cascade Events Rate |

### 4. Remediation Stats (`remediation-stats.json`)

| Row | Panels |
| - | - |
| **Remediation Performance** | Overall Success Rate, Total Remediations (24h), Median Remediation Duration, Auto-Approved (24h) |
| **Success Rate by Action Type** | Remediations by Action Type & Result, Success Rate per Action Type (24h), Remediation Duration Distribution |
| **Operator Health** | Operator Reconciliation Rate, Reconciliation Duration, Instance Health |
| **Approval Workflow Analytics** | Expired Approvals (24h), Approval Decision Distribution, Approval Duration by Mode, Auto-Approve Rate |
| **SLA Performance** | SLA Compliance by Severity, Response/Resolution Time by Severity (p95), SLA Violations Rate |
| **Session & Platform** | Active Sessions, Session Operations, Server Info, Server Uptime |

<Note>
  Every panel queries only metrics, labels and values the operator or the server exports; a test in the repository holds the four files to that. A few panels worth knowing:

  * **Expired Approvals (24h)** counts `chatcli_operator_approvals_total{result="expired"}`. For the approvals waiting right now, use the dashboard's Pending Approvals tile or `/api/v1/approvals?state=Pending`.
  * **Auto-Approved (24h)** and **Auto-Approve Rate** read `mode="auto", result="approved"`; **Notification Success Rate** reads `result="success"`.
  * **Notification Latency by Channel** reads `chatcli_operator_notification_duration_seconds` (p50/p95 per `channel_type`).
  * **Clusters by Status** / **Connected Clusters** read `chatcli_operator_federation_clusters_total`, a gauge set per state (`connected`, `degraded`, `disconnected`) from the registrations that exist.
  * **Critical Issues Detected (24h)** counts the critical Issues that entered `Analyzing` in the last 24 hours.
</Note>

## Grafana Dashboard Installation

### Via Grafana Sidecar (Recommended)

If you use the [Grafana Helm chart](https://github.com/grafana/helm-charts) with the dashboard sidecar enabled, create one ConfigMap with the four files and the label `grafana_dashboard: "1"`, from a checkout of the repository:

```bash theme={"system"}
kubectl create configmap chatcli-grafana-dashboards -n monitoring \
  --from-file=deploy/grafana/aiops-overview.json \
  --from-file=deploy/grafana/incident-timeline.json \
  --from-file=deploy/grafana/remediation-stats.json \
  --from-file=deploy/grafana/slo-burn-rate.json \
  --dry-run=client -o yaml | kubectl apply -f -

# Label for sidecar discovery, and a folder for the four dashboards
kubectl label configmap chatcli-grafana-dashboards -n monitoring grafana_dashboard=1 --overwrite
kubectl annotate configmap chatcli-grafana-dashboards -n monitoring grafana-folder="ChatCLI AIOps" --overwrite
```

<Note>
  The Grafana sidecar detects ConfigMaps with the label `grafana_dashboard: "1"` and imports the dashboards without a restart. `deploy/grafana/dashboards-configmap.yaml` holds only the two `ServiceMonitor`s Prometheus needs to scrape the operator and the server (the server one selects `app.kubernetes.io/name: chatcli`, the label Instance Services and the server chart carry), with these exact commands in its comments; it carries no dashboard ConfigMap, so applying it never overwrites the one created above. Name each JSON file: pointing `--from-file` at the directory would also pack that YAML into the ConfigMap.
</Note>

<Warning>
  The panels reference the datasource as `${DS_PROMETHEUS}`, which only the import dialog resolves; sidecar provisioning does not, and the panels show `Datasource ${DS_PROMETHEUS} was not found`. Substitute your Prometheus datasource UID in the files before creating the ConfigMap, for example `sed 's/\${DS_PROMETHEUS}/prometheus/g'` over each JSON file (see the [production setup](/cookbook/aiops-production-setup#11-grafana-dashboards)).
</Warning>

### Via Manual Import

1. Go to Grafana > **Dashboards** > **Import**
2. Upload the JSON file or paste the content
3. Select the Prometheus datasource
4. Click **Import**

### ServiceMonitor for Prometheus Operator

With the operator Helm chart, set `serviceMonitor.enabled: true` (plus `serviceMonitor.interval`, `scrapeTimeout` and `labels` as needed); the chart renders a ServiceMonitor for the `metrics` port (8080, path `/metrics`). Written by hand, for a release named `chatcli-operator`:

```yaml theme={"system"}
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: chatcli-operator
  namespace: chatcli-system
spec:
  selector:
    matchLabels:
      app.kubernetes.io/name: chatcli-operator
      app.kubernetes.io/instance: chatcli-operator
  endpoints:
    - port: metrics
      interval: 30s
      path: /metrics
  namespaceSelector:
    matchNames:
      - chatcli-system
```

The metrics endpoint is plain HTTP without authentication.

## Prometheus Metrics Reference

The operator exposes these metrics (in addition to the controller-runtime defaults) on its metrics port:

| Metric | Type | Labels | Description |
| - | - | - | - |
| `chatcli_operator_reconciliations_total` | Counter | `result` | Instance reconciliations by result |
| `chatcli_operator_reconciliation_duration_seconds` | Histogram | - | Instance reconciliation duration |
| `chatcli_operator_managed_instances` | Gauge | - | Instances managed |
| `chatcli_operator_instance_ready` | Gauge | `name`, `namespace` | 1 when the Instance is ready |
| `chatcli_operator_anomalies_processed_total` | Counter | `result` | Anomalies processed by result |
| `chatcli_operator_issues_created_by_correlation_total` | Counter | - | Issues created by anomaly correlation |
| `chatcli_operator_issues_total` | Counter | `severity`, `state` | Issue state transitions by severity and state |
| `chatcli_operator_issue_resolution_duration_seconds` | Histogram | - | Detection to resolution |
| `chatcli_operator_active_issues` | Gauge | - | Issues not yet resolved |
| `chatcli_operator_remediations_total` | Counter | `action_type`, `result` | Remediation actions by type and result (`success`, `failed`, ...) |
| `chatcli_operator_remediation_duration_seconds` | Histogram | - | Remediation plan execution duration |
| `chatcli_operator_decision_engine_evaluations_total` | Counter | `mode` | Decision engine verdicts |
| `chatcli_operator_decision_engine_circuit_breaker_state` | Gauge | `namespace` | 1 while the circuit breaker is open |
| `chatcli_operator_agentic_convergence_stops_total` | Counter | `reason` | Agentic loops stopped by the convergence detector |
| `chatcli_operator_approvals_total` | Counter | `mode`, `result` | Approval outcomes (`approved`, `rejected`, `expired`) by mode |
| `chatcli_operator_approval_duration_seconds` | Histogram | `mode` | Approval request lifecycle |
| `chatcli_operator_sla_response_time_seconds` | Histogram | `severity` | Detection to first analysis |
| `chatcli_operator_sla_resolution_time_seconds` | Histogram | `severity` | Detection to resolution |
| `chatcli_operator_sla_violations_total` | Counter | `severity`, `type` | SLA violations |
| `chatcli_operator_sla_compliance_percentage` | Gauge | `severity` | SLA compliance percentage |
| `chatcli_operator_slo_current_value` | Gauge | `service`, `slo_name` | Current SLI value |
| `chatcli_operator_slo_error_budget_remaining` | Gauge | `service`, `slo_name` | Remaining error budget fraction (0.0-1.0) |
| `chatcli_operator_slo_burn_rate` | Gauge | `service`, `slo_name`, `window` | Burn rate per window (`1h`, `6h`, `24h`, `72h`) |
| `chatcli_operator_slo_violations_total` | Counter | `service`, `slo_name`, `severity` | SLO violations |
| `chatcli_operator_notifications_sent_total` | Counter | `channel_type`, `severity`, `result` | Notifications by channel type and result |
| `chatcli_operator_notifications_failed_total` | Counter | `channel_type`, `reason` | Notification failures |
| `chatcli_operator_escalation_level_reached` | Counter | `policy`, `level` | Escalation levels reached |
| `chatcli_operator_notification_duration_seconds` | Histogram | `channel_type` | Time to deliver one notification |
| `chatcli_operator_federation_clusters_total` | Gauge | `status` | Registered clusters per state: `connected`, `degraded`, `disconnected` (see [Federation](/kubernetes/aiops/federation#prometheus-metrics)) |
| `chatcli_operator_federation_cross_cluster_issues_total` | Counter | - | Cross-cluster correlations |
| `chatcli_operator_federation_cascade_detected_total` | Counter | - | Cascades detected |
| `chatcli_operator_chaos_experiments_total` | Counter | `type`, `result` | Chaos experiments |
| `chatcli_operator_chaos_recovery_time_seconds` | Histogram | `type` | Target recovery time, measured from the end of the injection |
| `chatcli_operator_chaos_pods_affected_total` | Counter | `type` | Pods affected by experiments |

The operator exports no LLM token, cost or capacity metrics; LLM cost per incident is available from [`GET /api/v1/analytics/cost`](/reference/api/analytics-cost) and capacity forecasts from [`GET /api/v1/analytics/capacity`](/reference/api/analytics-capacity). LLM request and token metrics (`chatcli_llm_*`) come from the ChatCLI server.

## Useful Prometheus Queries

PromQL query examples for dashboards or alerts:

<AccordionGroup>
  <Accordion title="Median resolution time (last 24h)">
    ```promql theme={"system"}
    histogram_quantile(0.5,
      rate(chatcli_operator_issue_resolution_duration_seconds_bucket[24h])
    )
    ```
  </Accordion>

  <Accordion title="Remediation success rate">
    ```promql theme={"system"}
    sum(rate(chatcli_operator_remediations_total{result="success"}[24h]))
    /
    sum(rate(chatcli_operator_remediations_total[24h]))
    * 100
    ```
  </Accordion>

  <Accordion title="SLO burn rate (multi-window alert)">
    ```promql theme={"system"}
    # Alert: burn rate 1h > 14.4x AND burn rate 6h > 6x (Google SRE model)
    chatcli_operator_slo_burn_rate{window="1h"} > 14.4
    and
    chatcli_operator_slo_burn_rate{window="6h"} > 6.0
    ```
  </Accordion>

  <Accordion title="Plans parked for a human by the decision engine">
    ```promql theme={"system"}
    sum(rate(chatcli_operator_decision_engine_evaluations_total{mode=~"approval|manual|blocked"}[1h]))
    /
    sum(rate(chatcli_operator_decision_engine_evaluations_total[1h]))
    ```
  </Accordion>

  <Accordion title="Anomalies by processing result">
    ```promql theme={"system"}
    sum by (result) (rate(chatcli_operator_anomalies_processed_total[1h]))
    ```
  </Accordion>

  <Accordion title="Notification failure rate by channel">
    ```promql theme={"system"}
    sum by (channel_type) (rate(chatcli_operator_notifications_sent_total{result="failure"}[1h]))
    /
    sum by (channel_type) (rate(chatcli_operator_notifications_sent_total[1h]))
    ```
  </Accordion>
</AccordionGroup>

## Accessing and publishing the dashboard

<Steps>
  <Step title="Verify the operator">
    Confirm the operator is running:

    ```bash theme={"system"}
    kubectl get pods -n chatcli-system -l app.kubernetes.io/name=chatcli-operator
    ```
  </Step>

  <Step title="Local preview (no cluster)">
    To review the dashboard itself, or to try a change to it, serve it over a fake client seeded with synthetic data:

    ```bash theme={"system"}
    cd operator
    make dash-preview
    ```

    Open `http://127.0.0.1:8085/` and log in with the API key `preview`, which carries the admin role so approvals, reviews and post-mortem feedback can be exercised. The seed covers every tab: issues in every state, an AI insight, remediation plans, a pending approval, a post-mortem, SLOs, clusters, audit events and runbooks with steps. Set `DASHPREVIEW_ADDR` to change the address. The tool lives in `operator/hack/dashpreview` and is never part of the operator image.
  </Step>

  <Step title="Configure API Keys">
    The operator reads its API keys from the Secret `chatcli-operator-secrets` (key `api-keys`) in its own namespace, falling back to the ConfigMap `chatcli-operator-config` (same key). Create it yourself as below, or let the operator chart render it (`apiKeys.create: true` with `apiKeys.entries`; the keys are then stored in the Helm release, so do not do both). The value is a YAML list; roles are `viewer`, `operator` and `admin` (any other role is denied):

    ```bash theme={"system"}
    # Generate each key with e.g. `openssl rand -hex 32`
    cat > api-keys.yaml <<'EOF'
    - key: "replace-with-a-long-random-string"
      role: admin
      description: dashboard admin
    - key: "replace-with-another-random-string"
      role: viewer
      description: read-only dashboards
    EOF

    kubectl -n chatcli-system create secret generic chatcli-operator-secrets \
      --from-file=api-keys=api-keys.yaml
    ```

    Changes are picked up within about 30 seconds, without a restart. An edit that leaves invalid YAML keeps the last valid key set in force (the operator logs the error). With no keys configured, every `/api/` call returns `401` (the page loads but cannot log in).

    To read the keys back, for example to log in from a new browser:

    ```bash theme={"system"}
    kubectl -n chatcli-system get secret chatcli-operator-secrets \
      -o jsonpath='{.data.api-keys}' | base64 -d
    ```
  </Step>

  <Step title="Port-forward (development)">
    For local access during development:

    ```bash theme={"system"}
    kubectl port-forward -n chatcli-system svc/chatcli-operator 8090:8090
    ```

    Access: `http://localhost:8090/`
  </Step>

  <Step title="Publish it with an Ingress (production)">
    Serve the dashboard at the **root path of a host of its own**. The page calls `/api/v1/...` and `/healthz` with absolute paths, so under a sub-path (`/chatcli` with a rewrite or strip-path) the page loads but every panel fails; routing `/api/v1` separately on the shared host only hides the problem and publishes the API on every host of that controller.

    ```yaml theme={"system"}
    apiVersion: networking.k8s.io/v1
    kind: Ingress
    metadata:
      name: chatcli-dashboard
      namespace: chatcli-system
    spec:
      ingressClassName: nginx        # your controller's class: nginx, kong, traefik...
      tls:
        - hosts: ["aiops.example.com"]
          secretName: aiops-tls
      rules:
        - host: aiops.example.com
          http:
            paths:
              - path: /
                pathType: Prefix
                backend:
                  service:
                    name: chatcli-operator   # <release>-chatcli-operator for another release name
                    port:
                      name: api              # 8090
    ```

    * **Any controller:** the rule above needs no controller annotation. Do not add path rewriting (`nginx.ingress.kubernetes.io/rewrite-target`, `konghq.com/strip-path`, a Traefik `StripPrefix` middleware): with `path: /` there is nothing to strip.
    * **TLS:** terminate it at the Ingress as above, or serve HTTPS (TLS 1.3) from the operator itself with the chart's `security.apiTLS.certFile`/`keyFile` (`CHATCLI_AIOPS_TLS_CERT`/`CHATCLI_AIOPS_TLS_KEY`), mounting the certificate with `extraVolumes`/`extraVolumeMounts`. A backend that serves TLS needs the controller's backend-protocol setting (for ingress-nginx, `nginx.ingress.kubernetes.io/backend-protocol: "HTTPS"`).
    * **Local cluster (kind, k3d, Docker Desktop):** when the ingress controller answers on `localhost`, use a name under `.localhost`, for example `host: aiops.localhost`, and leave out the `tls` block. Browsers, `curl`, macOS and `systemd-resolved` send every `*.localhost` name to the loopback address, so no DNS or `/etc/hosts` entry is needed.
    * **Metrics:** do not publish the metrics port (8080) through the same Ingress: it is plain HTTP without authentication.
    * **Rate limits:** 30 requests per minute per client host for requests without a valid API key, 600 per minute per valid key. Behind an Ingress every unauthenticated request comes from the controller, so they share one bucket; dashboard traffic carries the key, so logged-in browsers do not.
    * **CORS:** cross-origin calls are denied unless allowed with `security.corsAllowedOrigins` (`CHATCLI_CORS_ALLOWED_ORIGINS`); the dashboard itself is same-origin and needs no CORS setting.

    Check it: `curl -s https://aiops.example.com/healthz` answers `{"status":"ok",...}`, and `curl -s -o /dev/null -w '%{http_code}' https://aiops.example.com/api/v1/incidents` answers `401` without a key.
  </Step>
</Steps>

<Warning>
  Never expose the dashboard without API keys in production. Dev mode (`CHATCLI_OPERATOR_DEV_MODE=true`, or `TRUE`, `1`, `t`; chart `security.devMode`) with no keys configured gives **every caller the admin role** without a key, including write operations such as acknowledge, resolve, approve and reject. Use it only locally.
</Warning>

## Next Steps

<CardGroup cols={2}>
  <Card title="REST API Reference" icon="code" href="/reference/api/overview">
    Complete reference of all endpoints consumed by the dashboard.
  </Card>

  <Card title="Capacity & Costs" icon="chart-line" href="/kubernetes/aiops/capacity-cost">
    Details on the Capacity Planner, Noise Reducer, and Cost Tracker.
  </Card>

  <Card title="AIOps Platform" icon="brain" href="/kubernetes/aiops-platform">
    Complete architecture of the autonomous operations pipeline.
  </Card>

  <Card title="K8s Operator" icon="dharmachakra" href="/kubernetes/k8s-operator">
    Kubernetes operator configuration and deployment.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.