Web Dashboard
Overview
The Web Dashboard is a Single Page Application embedded directly in the operator binary via Goembed.FS — it does not require Node.js, npm, or any separate frontend build.
The dashboard consumes the same REST API documented in the API Reference. Every operation available in the dashboard (acknowledge, snooze, resolve, approve, reject, post-mortem review, close, feedback and human-action acknowledgement) is an authenticated REST call and needs at least the
operator role; a viewer key sees everything but its actions return 403.Theme and language
Both switchers live in the header, next to the auto-refresh control, and are keyboard accessible. The choice is stored in the browser and applied before the first paint, so a reload never flashes the wrong theme.
The query parameters
?theme= and ?lang= set the preference once and persist it, so a bookmark such as http://localhost:8090/?lang=pt-BR&theme=dark opens the dashboard in Portuguese and dark for every visit that follows. The <html lang> attribute tracks the language, and values that come from the API (severity, incident and plan states, approval decisions) keep their raw value in the markup while the visible label is translated.
Architecture
/) is served without authentication; every data call it makes goes to /api/v1/ with the API key. On first load the dashboard asks for the key, checks it against /api/v1/incidents, and keeps it in the browser’s localStorage (chatcli_api_key); Change API key in the header clears it.
The dashboard is served by the operator, on the
api port (8090) of the Service chatcli-operator in the operator namespace, not by the ChatCLI Instance’s Service. Every operator replica serves the dashboard and the API, so a Service with several replicas can route to any of them.Dashboard Views
The dashboard has 11 views accessible via tab navigation. A time-range filter in the header (All time, or From/To) is passed asfrom/to to the lists that support it.
1. Overview
1. Overview
Platform overview with aggregated metrics, built from
/analytics/summary, /analytics/compliance, /analytics/capacity, /analytics/remediation-stats, /analytics/mttd, /analytics/mttr and the incident list.Components:The panels can be rearranged by dragging their headers; the order is remembered in
localStorage (panelOrder_overviewPanels).2. Incidents
2. Incidents
Interactive table of the incidents with filters and actions.Features:
Ack, Snooze and Resolve need the
operator role. Ack stops the incident’s escalation at its current level (no further level, no repeat) and stamps who acknowledged on the EscalationPolicy status. Snooze holds every notification except Resolved, and the escalation, until the snooze ends; a held escalation page goes out then and the level’s timer restarts. See acknowledgement and snooze.Severity badges (dark theme):State badges (dark theme):
3. SLOs
3. SLOs
One card per SLO, sorted by name, plus an Active SLO Alerts table listing the SLOs below target.Components per SLO:
4. Approvals
4. Approvals
Approval requests, newest first, with approve/reject actions. The tab badge shows the number of pending requests.Features:
A decision made here is the same as the REST call: it is stored in
status.decisions as <your name> (api-key: <key identity>), and the approval controller then evaluates quorum and change window. A quorum counts each API key once, so each approver needs their own key; a request that is no longer pending, or that your key already decided, returns an error. See Approval Workflow. When the request carries the annotation platform.chatcli.io/blast-risk-level, the card shows a blast-radius badge colored by level.5. AI Insights
5. AI Insights
View all AI-generated analyses to understand how the AI reasoned about each incident.Features:
This view is essential when an incident is escalated to human action — it shows what the AI found, why it recommended specific actions, and what enrichment data informed its analysis.API endpoint:
GET /api/v1/aiinsights6. Remediations
6. Remediations
Track all remediation plans with execution details, both runbook-based and agentic.Features:
Remediation modes explained:
- Runbook mode: Displays the action sequence from the runbook the plan was built from
- Agentic mode: Shows the step count; use the Get Remediation Plan API for the full AI conversation history
GET /api/v1/remediations7. Runbooks
7. Runbooks
View all runbooks — both manually created and generated by the AI.Features:
Non-agentic RemediationPlans are built from a runbook. When an Issue is analyzed and no runbook matching its trigger (signal type, severity, resource kind) is accepted, the operator generates one from the AI’s suggested actions; after a successful agentic remediation it also saves the actions that worked as an
agentic-… runbook. Generated runbooks carry the label platform.chatcli.io/auto-generated: "true" and are reused by the next similar incident instead of starting from scratch.API endpoint: GET /api/v1/runbooks8. PostMortems
8. PostMortems
Table of post-mortems with expandable details.Features:
Review, Close, feedback and Ack Human Action need the
operator role.9. Clusters
9. Clusters
Federation overview and one card per registered cluster.Summary tiles (from
/clusters/global-status): Total Clusters, Healthy (connected, every node Ready), Degraded (connected, with the Degraded condition: some nodes not Ready), Offline.Federation Panel:Components per cluster:
API endpoints:
GET /api/v1/clusters/global-status, GET /api/v1/federation/status, GET /api/v1/federation/correlations10. Policies
10. Policies
Read-only view of the four policy kinds the controllers read: ApprovalPolicy, NotificationPolicy, EscalationPolicy and IncidentSLA. Policies are written through
kubectl or GitOps, where they are reviewed; the tab shows what is in force, per namespace or across all.Every column sorts. API endpoints:
GET /api/v1/policies/{kind}, GET /api/v1/policies/{kind}/{name}11. Audit
11. Audit
Searchable audit log (AuditEvent CRs) with export.Features:
Grafana Dashboards
The repository ships 4 Grafana dashboards as JSON indeploy/grafana/, ready for import. Panels built on chatcli_operator_* metrics need the operator’s metrics endpoint scraped; panels on chatcli_grpc_*, chatcli_llm_*, chatcli_session_*, chatcli_server_* and chatcli_watcher_* need the ChatCLI server’s metrics port (9090) scraped too.
1. AIOps Overview (aiops-overview.json)
2. SLO Burn Rate (slo-burn-rate.json)
Template variables service and slo_name.
3. Incident Timeline (incident-timeline.json)
4. Remediation Stats (remediation-stats.json)
Every panel queries only metrics, labels and values the operator or the server exports; a test in the repository holds the four files to that. A few panels worth knowing:
- Expired Approvals (24h) counts
chatcli_operator_approvals_total{result="expired"}. For the approvals waiting right now, use the dashboard’s Pending Approvals tile or/api/v1/approvals?state=Pending. - Auto-Approved (24h) and Auto-Approve Rate read
mode="auto", result="approved"; Notification Success Rate readsresult="success". - Notification Latency by Channel reads
chatcli_operator_notification_duration_seconds(p50/p95 perchannel_type). - Clusters by Status / Connected Clusters read
chatcli_operator_federation_clusters_total, a gauge set per state (connected,degraded,disconnected) from the registrations that exist. - Critical Issues Detected (24h) counts the critical Issues that entered
Analyzingin the last 24 hours.
Grafana Dashboard Installation
Via Grafana Sidecar (Recommended)
If you use the Grafana Helm chart with the dashboard sidecar enabled, create one ConfigMap with the four files and the labelgrafana_dashboard: "1", from a checkout of the repository:
The Grafana sidecar detects ConfigMaps with the label
grafana_dashboard: "1" and imports the dashboards without a restart. deploy/grafana/dashboards-configmap.yaml holds only the two ServiceMonitors Prometheus needs to scrape the operator and the server (the server one selects app.kubernetes.io/name: chatcli, the label Instance Services and the server chart carry), with these exact commands in its comments; it carries no dashboard ConfigMap, so applying it never overwrites the one created above. Name each JSON file: pointing --from-file at the directory would also pack that YAML into the ConfigMap.Via Manual Import
- Go to Grafana > Dashboards > Import
- Upload the JSON file or paste the content
- Select the Prometheus datasource
- Click Import
ServiceMonitor for Prometheus Operator
With the operator Helm chart, setserviceMonitor.enabled: true (plus serviceMonitor.interval, scrapeTimeout and labels as needed); the chart renders a ServiceMonitor for the metrics port (8080, path /metrics). Written by hand, for a release named chatcli-operator:
Prometheus Metrics Reference
The operator exposes these metrics (in addition to the controller-runtime defaults) on its metrics port:
The operator exports no LLM token, cost or capacity metrics; LLM cost per incident is available from
GET /api/v1/analytics/cost and capacity forecasts from GET /api/v1/analytics/capacity. LLM request and token metrics (chatcli_llm_*) come from the ChatCLI server.
Useful Prometheus Queries
PromQL query examples for dashboards or alerts:Median resolution time (last 24h)
Median resolution time (last 24h)
Remediation success rate
Remediation success rate
SLO burn rate (multi-window alert)
SLO burn rate (multi-window alert)
Plans parked for a human by the decision engine
Plans parked for a human by the decision engine
Anomalies by processing result
Anomalies by processing result
Notification failure rate by channel
Notification failure rate by channel
Accessing and publishing the dashboard
1
Verify the operator
Confirm the operator is running:
2
Local preview (no cluster)
To review the dashboard itself, or to try a change to it, serve it over a fake client seeded with synthetic data:Open
http://127.0.0.1:8085/ and log in with the API key preview, which carries the admin role so approvals, reviews and post-mortem feedback can be exercised. The seed covers every tab: issues in every state, an AI insight, remediation plans, a pending approval, a post-mortem, SLOs, clusters, audit events and runbooks with steps. Set DASHPREVIEW_ADDR to change the address. The tool lives in operator/hack/dashpreview and is never part of the operator image.3
Configure API Keys
The operator reads its API keys from the Secret Changes are picked up within about 30 seconds, without a restart. An edit that leaves invalid YAML keeps the last valid key set in force (the operator logs the error). With no keys configured, every
chatcli-operator-secrets (key api-keys) in its own namespace, falling back to the ConfigMap chatcli-operator-config (same key). Create it yourself as below, or let the operator chart render it (apiKeys.create: true with apiKeys.entries; the keys are then stored in the Helm release, so do not do both). The value is a YAML list; roles are viewer, operator and admin (any other role is denied):/api/ call returns 401 (the page loads but cannot log in).To read the keys back, for example to log in from a new browser:4
Port-forward (development)
For local access during development:Access:
http://localhost:8090/5
Publish it with an Ingress (production)
Serve the dashboard at the root path of a host of its own. The page calls
/api/v1/... and /healthz with absolute paths, so under a sub-path (/chatcli with a rewrite or strip-path) the page loads but every panel fails; routing /api/v1 separately on the shared host only hides the problem and publishes the API on every host of that controller.- Any controller: the rule above needs no controller annotation. Do not add path rewriting (
nginx.ingress.kubernetes.io/rewrite-target,konghq.com/strip-path, a TraefikStripPrefixmiddleware): withpath: /there is nothing to strip. - TLS: terminate it at the Ingress as above, or serve HTTPS (TLS 1.3) from the operator itself with the chart’s
security.apiTLS.certFile/keyFile(CHATCLI_AIOPS_TLS_CERT/CHATCLI_AIOPS_TLS_KEY), mounting the certificate withextraVolumes/extraVolumeMounts. A backend that serves TLS needs the controller’s backend-protocol setting (for ingress-nginx,nginx.ingress.kubernetes.io/backend-protocol: "HTTPS"). - Local cluster (kind, k3d, Docker Desktop): when the ingress controller answers on
localhost, use a name under.localhost, for examplehost: aiops.localhost, and leave out thetlsblock. Browsers,curl, macOS andsystemd-resolvedsend every*.localhostname to the loopback address, so no DNS or/etc/hostsentry is needed. - Metrics: do not publish the metrics port (8080) through the same Ingress: it is plain HTTP without authentication.
- Rate limits: 30 requests per minute per client host for requests without a valid API key, 600 per minute per valid key. Behind an Ingress every unauthenticated request comes from the controller, so they share one bucket; dashboard traffic carries the key, so logged-in browsers do not.
- CORS: cross-origin calls are denied unless allowed with
security.corsAllowedOrigins(CHATCLI_CORS_ALLOWED_ORIGINS); the dashboard itself is same-origin and needs no CORS setting.
curl -s https://aiops.example.com/healthz answers {"status":"ok",...}, and curl -s -o /dev/null -w '%{http_code}' https://aiops.example.com/api/v1/incidents answers 401 without a key.Next Steps
REST API Reference
Complete reference of all endpoints consumed by the dashboard.
Capacity & Costs
Details on the Capacity Planner, Noise Reducer, and Cost Tracker.
AIOps Platform
Complete architecture of the autonomous operations pipeline.
K8s Operator
Kubernetes operator configuration and deployment.