Skip to main content
The ChatCLI AIOps platform includes a built-in web dashboard (self-contained SPA served by the operator) and 4 Grafana dashboards (JSON in the repository) for observability of the autonomous operations pipeline.

Web Dashboard

Overview

The Web Dashboard is a Single Page Application embedded directly in the operator binary via Go embed.FS — it does not require Node.js, npm, or any separate frontend build.
The dashboard consumes the same REST API documented in the API Reference. Every operation available in the dashboard (acknowledge, snooze, resolve, approve, reject, post-mortem review, close, feedback and human-action acknowledgement) is an authenticated REST call and needs at least the operator role; a viewer key sees everything but its actions return 403.

Theme and language

Both switchers live in the header, next to the auto-refresh control, and are keyboard accessible. The choice is stored in the browser and applied before the first paint, so a reload never flashes the wrong theme. The query parameters ?theme= and ?lang= set the preference once and persist it, so a bookmark such as http://localhost:8090/?lang=pt-BR&theme=dark opens the dashboard in Portuguese and dark for every visit that follows. The <html lang> attribute tracks the language, and values that come from the API (severity, incident and plan states, approval decisions) keep their raw value in the markup while the visible label is translated.
Product names stay in English in both languages: Issue, SLO, Runbook, PostMortem, AIInsight and RemediationPlan are the names of the CRDs the operator manages.

Architecture

The page itself (/) is served without authentication; every data call it makes goes to /api/v1/ with the API key. On first load the dashboard asks for the key, checks it against /api/v1/incidents, and keeps it in the browser’s localStorage (chatcli_api_key); Change API key in the header clears it.
The dashboard is served by the operator, on the api port (8090) of the Service chatcli-operator in the operator namespace, not by the ChatCLI Instance’s Service. Every operator replica serves the dashboard and the API, so a Service with several replicas can route to any of them.

Dashboard Views

The dashboard has 11 views accessible via tab navigation. A time-range filter in the header (All time, or From/To) is passed as from/to to the lists that support it.
Platform overview with aggregated metrics, built from /analytics/summary, /analytics/compliance, /analytics/capacity, /analytics/remediation-stats, /analytics/mttd, /analytics/mttr and the incident list.Components:The panels can be rearranged by dragging their headers; the order is remembered in localStorage (panelOrder_overviewPanels).
Interactive table of the incidents with filters and actions.Features:Ack, Snooze and Resolve need the operator role. Ack stops the incident’s escalation at its current level (no further level, no repeat) and stamps who acknowledged on the EscalationPolicy status. Snooze holds every notification except Resolved, and the escalation, until the snooze ends; a held escalation page goes out then and the level’s timer restarts. See acknowledgement and snooze.Severity badges (dark theme):State badges (dark theme):
One card per SLO, sorted by name, plus an Active SLO Alerts table listing the SLOs below target.Components per SLO:
Approval requests, newest first, with approve/reject actions. The tab badge shows the number of pending requests.Features:
A decision made here is the same as the REST call: it is stored in status.decisions as <your name> (api-key: <key identity>), and the approval controller then evaluates quorum and change window. A quorum counts each API key once, so each approver needs their own key; a request that is no longer pending, or that your key already decided, returns an error. See Approval Workflow. When the request carries the annotation platform.chatcli.io/blast-risk-level, the card shows a blast-radius badge colored by level.
View all AI-generated analyses to understand how the AI reasoned about each incident.Features:This view is essential when an incident is escalated to human action — it shows what the AI found, why it recommended specific actions, and what enrichment data informed its analysis.API endpoint: GET /api/v1/aiinsights
Track all remediation plans with execution details, both runbook-based and agentic.Features:Remediation modes explained:
  • Runbook mode: Displays the action sequence from the runbook the plan was built from
  • Agentic mode: Shows the step count; use the Get Remediation Plan API for the full AI conversation history
API endpoint: GET /api/v1/remediations
View all runbooks — both manually created and generated by the AI.Features:Non-agentic RemediationPlans are built from a runbook. When an Issue is analyzed and no runbook matching its trigger (signal type, severity, resource kind) is accepted, the operator generates one from the AI’s suggested actions; after a successful agentic remediation it also saves the actions that worked as an agentic-… runbook. Generated runbooks carry the label platform.chatcli.io/auto-generated: "true" and are reused by the next similar incident instead of starting from scratch.API endpoint: GET /api/v1/runbooks
Table of post-mortems with expandable details.Features:Review, Close, feedback and Ack Human Action need the operator role.
Federation overview and one card per registered cluster.Summary tiles (from /clusters/global-status): Total Clusters, Healthy (connected, every node Ready), Degraded (connected, with the Degraded condition: some nodes not Ready), Offline.Federation Panel:Components per cluster:API endpoints: GET /api/v1/clusters/global-status, GET /api/v1/federation/status, GET /api/v1/federation/correlations
Read-only view of the four policy kinds the controllers read: ApprovalPolicy, NotificationPolicy, EscalationPolicy and IncidentSLA. Policies are written through kubectl or GitOps, where they are reviewed; the tab shows what is in force, per namespace or across all.Every column sorts. API endpoints: GET /api/v1/policies/{kind}, GET /api/v1/policies/{kind}/{name}
Searchable audit log (AuditEvent CRs) with export.Features:

Grafana Dashboards

The repository ships 4 Grafana dashboards as JSON in deploy/grafana/, ready for import. Panels built on chatcli_operator_* metrics need the operator’s metrics endpoint scraped; panels on chatcli_grpc_*, chatcli_llm_*, chatcli_session_*, chatcli_server_* and chatcli_watcher_* need the ChatCLI server’s metrics port (9090) scraped too.

1. AIOps Overview (aiops-overview.json)

2. SLO Burn Rate (slo-burn-rate.json)

Template variables service and slo_name.

3. Incident Timeline (incident-timeline.json)

4. Remediation Stats (remediation-stats.json)

Every panel queries only metrics, labels and values the operator or the server exports; a test in the repository holds the four files to that. A few panels worth knowing:
  • Expired Approvals (24h) counts chatcli_operator_approvals_total{result="expired"}. For the approvals waiting right now, use the dashboard’s Pending Approvals tile or /api/v1/approvals?state=Pending.
  • Auto-Approved (24h) and Auto-Approve Rate read mode="auto", result="approved"; Notification Success Rate reads result="success".
  • Notification Latency by Channel reads chatcli_operator_notification_duration_seconds (p50/p95 per channel_type).
  • Clusters by Status / Connected Clusters read chatcli_operator_federation_clusters_total, a gauge set per state (connected, degraded, disconnected) from the registrations that exist.
  • Critical Issues Detected (24h) counts the critical Issues that entered Analyzing in the last 24 hours.

Grafana Dashboard Installation

If you use the Grafana Helm chart with the dashboard sidecar enabled, create one ConfigMap with the four files and the label grafana_dashboard: "1", from a checkout of the repository:
The Grafana sidecar detects ConfigMaps with the label grafana_dashboard: "1" and imports the dashboards without a restart. deploy/grafana/dashboards-configmap.yaml holds only the two ServiceMonitors Prometheus needs to scrape the operator and the server (the server one selects app.kubernetes.io/name: chatcli, the label Instance Services and the server chart carry), with these exact commands in its comments; it carries no dashboard ConfigMap, so applying it never overwrites the one created above. Name each JSON file: pointing --from-file at the directory would also pack that YAML into the ConfigMap.
The panels reference the datasource as ${DS_PROMETHEUS}, which only the import dialog resolves; sidecar provisioning does not, and the panels show Datasource ${DS_PROMETHEUS} was not found. Substitute your Prometheus datasource UID in the files before creating the ConfigMap, for example sed 's/\${DS_PROMETHEUS}/prometheus/g' over each JSON file (see the production setup).

Via Manual Import

  1. Go to Grafana > Dashboards > Import
  2. Upload the JSON file or paste the content
  3. Select the Prometheus datasource
  4. Click Import

ServiceMonitor for Prometheus Operator

With the operator Helm chart, set serviceMonitor.enabled: true (plus serviceMonitor.interval, scrapeTimeout and labels as needed); the chart renders a ServiceMonitor for the metrics port (8080, path /metrics). Written by hand, for a release named chatcli-operator:
The metrics endpoint is plain HTTP without authentication.

Prometheus Metrics Reference

The operator exposes these metrics (in addition to the controller-runtime defaults) on its metrics port: The operator exports no LLM token, cost or capacity metrics; LLM cost per incident is available from GET /api/v1/analytics/cost and capacity forecasts from GET /api/v1/analytics/capacity. LLM request and token metrics (chatcli_llm_*) come from the ChatCLI server.

Useful Prometheus Queries

PromQL query examples for dashboards or alerts:

Accessing and publishing the dashboard

1

Verify the operator

Confirm the operator is running:
2

Local preview (no cluster)

To review the dashboard itself, or to try a change to it, serve it over a fake client seeded with synthetic data:
Open http://127.0.0.1:8085/ and log in with the API key preview, which carries the admin role so approvals, reviews and post-mortem feedback can be exercised. The seed covers every tab: issues in every state, an AI insight, remediation plans, a pending approval, a post-mortem, SLOs, clusters, audit events and runbooks with steps. Set DASHPREVIEW_ADDR to change the address. The tool lives in operator/hack/dashpreview and is never part of the operator image.
3

Configure API Keys

The operator reads its API keys from the Secret chatcli-operator-secrets (key api-keys) in its own namespace, falling back to the ConfigMap chatcli-operator-config (same key). Create it yourself as below, or let the operator chart render it (apiKeys.create: true with apiKeys.entries; the keys are then stored in the Helm release, so do not do both). The value is a YAML list; roles are viewer, operator and admin (any other role is denied):
Changes are picked up within about 30 seconds, without a restart. An edit that leaves invalid YAML keeps the last valid key set in force (the operator logs the error). With no keys configured, every /api/ call returns 401 (the page loads but cannot log in).To read the keys back, for example to log in from a new browser:
4

Port-forward (development)

For local access during development:
Access: http://localhost:8090/
5

Publish it with an Ingress (production)

Serve the dashboard at the root path of a host of its own. The page calls /api/v1/... and /healthz with absolute paths, so under a sub-path (/chatcli with a rewrite or strip-path) the page loads but every panel fails; routing /api/v1 separately on the shared host only hides the problem and publishes the API on every host of that controller.
  • Any controller: the rule above needs no controller annotation. Do not add path rewriting (nginx.ingress.kubernetes.io/rewrite-target, konghq.com/strip-path, a Traefik StripPrefix middleware): with path: / there is nothing to strip.
  • TLS: terminate it at the Ingress as above, or serve HTTPS (TLS 1.3) from the operator itself with the chart’s security.apiTLS.certFile/keyFile (CHATCLI_AIOPS_TLS_CERT/CHATCLI_AIOPS_TLS_KEY), mounting the certificate with extraVolumes/extraVolumeMounts. A backend that serves TLS needs the controller’s backend-protocol setting (for ingress-nginx, nginx.ingress.kubernetes.io/backend-protocol: "HTTPS").
  • Local cluster (kind, k3d, Docker Desktop): when the ingress controller answers on localhost, use a name under .localhost, for example host: aiops.localhost, and leave out the tls block. Browsers, curl, macOS and systemd-resolved send every *.localhost name to the loopback address, so no DNS or /etc/hosts entry is needed.
  • Metrics: do not publish the metrics port (8080) through the same Ingress: it is plain HTTP without authentication.
  • Rate limits: 30 requests per minute per client host for requests without a valid API key, 600 per minute per valid key. Behind an Ingress every unauthenticated request comes from the controller, so they share one bucket; dashboard traffic carries the key, so logged-in browsers do not.
  • CORS: cross-origin calls are denied unless allowed with security.corsAllowedOrigins (CHATCLI_CORS_ALLOWED_ORIGINS); the dashboard itself is same-origin and needs no CORS setting.
Check it: curl -s https://aiops.example.com/healthz answers {"status":"ok",...}, and curl -s -o /dev/null -w '%{http_code}' https://aiops.example.com/api/v1/incidents answers 401 without a key.
Never expose the dashboard without API keys in production. Dev mode (CHATCLI_OPERATOR_DEV_MODE=true, or TRUE, 1, t; chart security.devMode) with no keys configured gives every caller the admin role without a key, including write operations such as acknowledge, resolve, approve and reject. Use it only locally.

Next Steps

REST API Reference

Complete reference of all endpoints consumed by the dashboard.

Capacity & Costs

Details on the Capacity Planner, Noise Reducer, and Cost Tracker.

AIOps Platform

Complete architecture of the autonomous operations pipeline.

K8s Operator

Kubernetes operator configuration and deployment.