Quick start
Is the deployment healthy?, Why is pod myapp-7d9f-abcde restarting?. One-shot:
Commands and flags
chatcli watch
The single-target flags always watch a Deployment. StatefulSets, DaemonSets, Jobs and CronJobs are watched through
kind in a config file (with the Helm chart, watcher.targets[].kind).
Inside ChatCLI
/watch start accepts --deployment, --namespace, --interval, --window, --max-log-lines and --kubeconfig.
chatcli server
On the server the context is added to prompts server-side, so every
chatcli connect client gets it without setup, and the alerts are served by GetAlerts / StreamAlerts. The server prints K8s watcher active: N targets (interval: 30s) at startup; a config file that fails to load stops the server with failed to load watch config: … (in the log, see Logs).
Multi-target configuration file
The file is checked at load:
watch config file has no targets, target[0]: deployment (resource name) is required, target[2]: invalid kind "Pod" (must be Deployment, StatefulSet, DaemonSet, Job, or CronJob), invalid interval "30". The kubeconfig is not a file field; pass --kubeconfig / --watch-kubeconfig.
Kubeconfig and cluster access
The watcher uses the kubeconfig given by--kubeconfig / --watch-kubeconfig / CHATCLI_KUBECONFIG. Without one it tries the in-cluster service account, then ~/.kube/config, always with the file’s current context. The KUBECONFIG environment variable is not read.
What is collected
Every interval, per target:
The store keeps
window / interval + 1 snapshots (at least 10) and up to 10 × maxLogLines log lines per target; older data falls out of the window.
Prometheus scraping
- Scrapes one pod: the first
Runningpod with an IP among the resource’s pods, athttp://<podIP>:<metricsPort><metricsPath>(plain HTTP, 5 s timeout). The watcher must be able to reach pod IPs (run it in the cluster, or from a network that routes to pods). - Parses the text exposition format, skips comments,
NaNand±Inf. - Keeps one value per metric name: labels are dropped and the last sample of a name wins, so
http_requests_total{code="200"}and{code="500"}collapse into one number. Filter to metrics that make sense unlabeled, or expose aggregates. metricsFilterglobs use*as the only wildcard:http_*,*_errors_total,go_goroutines.
Alerts
An alert with the same type and object as one already in the window is not raised again. Alerts appear in the context (
## Active Alerts), drive the context budget, and on the server feed GetAlerts / StreamAlerts.
Context and budget
With one target the full context is sent (resource status, pods, HPA, nodes, recent events, application metrics, active alerts, recent error logs):maxContextChars:
- Each target is scored:
2with a critical alert;1with ready replicas below desired, a warning alert or error logs;0otherwise. - Targets are sorted by score, then by alert count.
- Targets scoring 1 or 2 get their full context; healthy ones a one-line summary.
- Over budget, the healthiest detailed targets are reduced to one line, then the healthiest lines are dropped.
/watch status in chatcli watch shows K8s Watcher: Watching 12 targets: 10 healthy, 1 warning, 1 critical.
Prometheus metrics of the watcher
When the watcher runs insidechatcli server with metrics on (--metrics-port, default 9090):
chatcli_watcher_alerts_total counts an alert when the watcher stores it as new. A condition that persists is seen on every collection, but a repeat of the same type on the same object is dropped while the earlier alert is still held in the window (--watch-window), so a standing problem counts once rather than once per collection. It is a counter: use increase() or rate() over a range.
RBAC
Read-only access the watcher uses. The node section (cluster-scoped) is optional: without it the node data and node alerts are missing and everything else works.chatcli watch, the pod’s ServiceAccount for the server.
- Helm chart:
rbac.create: true(default) creates the RBAC; it becomes a ClusterRole automatically when targets are outside the release namespace or span several, or when a single-targetwatcher.namespaceis not the release namespace (an empty one meansdefault).watcher.targets[].kindpicksDeployment(default),StatefulSet,DaemonSet,JoborCronJob, as in the config file. The chart’s rules are broader than the watcher needs (they also cover the AIOps remediation resources) but grant no access to Secrets; reviewrbac.*before installing in a sensitive cluster, and userbac.additionalRulesfor a plugin that needs a Secret. - Operator: for a watcher reading outside the Instance namespace (through
targetsor a legacy singlewatcher.namespace) the operator binds the pre-provisionedchatcli-watcherClusterRole; in the Instance namespace it creates a namespaced Role. Both read Jobs and CronJobs as well as the other workload kinds; see K8s Operator.
AIOps integration
The operator reads the server’s watcher alerts overStreamAlerts (falling back to polling GetAlerts against a server without the stream) and turns them into Anomaly resources, which it correlates into Issues, analyzes (AIInsight) and remediates (RemediationPlan). See AIOps Platform.
Troubleshooting
Next steps
Server Mode
Share the watcher with a team
K8s Monitoring recipe
Step by step
K8s Operator
Managed Instances and AIOps
Docker & Kubernetes
Deploy the server