Skip to main content
Multi-Cluster Federation lets a ChatCLI operator see beyond its own cluster. You register the other clusters with a ClusterRegistration; the operator health-checks them, correlates new Issues with the Issues running there, flags staging-to-production cascades, and applies a remediation policy based on its own cluster’s tier.
Federation does not require a service mesh or external tools. The operator connects directly to each registered cluster with a kubeconfig stored in a Secret.
Federation reads other clusters; it does not remediate them. Detection and remediation stay local: each cluster that should heal itself runs its own operator (and its own ChatCLI Instance; only one Instance per cluster drives AIOps). Through the registered kubeconfig the operator only lists nodes, namespaces, Issues and RemediationPlans, and updates remote Issues when it correlates them.

Why Multi-Cluster Federation?

In modern production environments, infrastructure is rarely limited to a single cluster:

Multi-Region

Clusters in us-east-1, eu-west-1, and ap-southeast-1 for latency and regional compliance.

Multi-Environment

Staging, production, and DR in separate clusters with different security policies.

Multi-Tenant

Dedicated clusters per team or product with strong workload isolation.
Without federation, each cluster is a silo. AIOps loses the ability to:
  • Detect that the same problem affects 5 clusters simultaneously
  • Correlate a staging deploy with a production failure
  • Apply differentiated remediation policies by cluster importance
  • Aggregate cluster health into a global view

ClusterRegistration CRD

The ClusterRegistration CRD is the entry point for adding clusters to the federation. Create it in the cluster where the operator runs, in any namespace; the kubeconfig Secret must live in the same namespace.

Complete Specification

Spec Fields

Status Fields

How Federation Works

Kubeconfig and client cache

On each reconcile the FederationReconciler reads data.kubeconfig from the Secret named in kubeconfigSecretRef, in the ClusterRegistration’s namespace, builds a controller-runtime client from it, and caches that client by registration name. The cache is dropped when a health check fails or the registration is deleted, so a rotated kubeconfig is picked up after the next failed check (or an operator restart). The kubeconfig identity needs, on the remote cluster: list on nodes and namespaces for the health check; list on issues and remediationplans (group platform.chatcli.io) for the counters and correlation; and update on issues for correlation to annotate them.

Health Check Loop

1

List Nodes

Lists the remote nodes to verify connectivity, count them and read the first node’s kubelet version.
2

List Namespaces

Lists the remote namespaces and counts them.
3

Count AIOps work

Lists remote Issues and RemediationPlans, when those CRDs exist there, to fill activeIssues and activeRemediations.
4

Update Status

Writes connected, lastHealthCheck, kubernetesVersion, nodeCount, namespaceCount, activeIssues, activeRemediations and the Connected/Degraded conditions, refreshes the chatcli_operator_federation_clusters_total gauge, then requeues after healthCheckInterval. A failure at any step before this one sets connected: false, refreshes lastHealthCheck and records the failure in the Connected condition.

Cross-Cluster Correlation

When an Issue is detected, the IssueReconciler asks the federation controller to look for the same problem elsewhere. Both checks below run only against ClusterRegistrations that are connected and whose kubeconfig client is cached; with no registrations they cost two cached lists and change nothing.

Automatic Severity Elevation

The controller lists the active (non-terminal) Issues with the same signalType on the local cluster and on every connected cluster. When they span 3 or more clusters (the local one counts as one), every matching Issue, local and remote, is annotated and raised to critical:
An Issue raised to critical also gets platform.chatcli.io/elevated-severity: "true". One correlation keeps one id: when a new matching Issue arrives and one of the matching Issues already carries a correlation id, the group reuses it instead of relabeling everything, and affected-clusters is refreshed. The counter chatcli_operator_federation_cross_cluster_issues_total increments once per new correlation id. An Issue that changed underneath the update (a conflict) is skipped and not retried; the next correlation of the group annotates it.
The correlation id is an annotation, so find the related Issues with:
GET /api/v1/federation/correlations lists the correlations from the local Issues, one entry per correlation id, with correlationId, issue, namespace, severity, signalType, correlatedClusters (from affected-clusters), elevated (from elevated-severity) and cascade. It also reads the older platform.chatcli.io/correlation-id / correlated-clusters keys, so Issues annotated with them still show up.

Cascade Detection

Cascade detection flags a newly detected Issue whose resource (same name) already failed earlier in a staging cluster: the same problem is rolling through the environments.

Staging to Production

The check needs at least one connected registration with environment: staging and one with environment: prod. It runs for every Issue the local operator detects (the local cluster’s own environment is not consulted); for that Issue, the controller lists the Issues on every staging cluster and looks for one on a resource with the same name whose detectedAt is earlier than the local Issue’s. The first match annotates the local Issue:
and increments chatcli_operator_federation_cascade_detected_total.
The annotation is the whole effect: severity, notifications and the AI prompt are not changed by a cascade. Route on it from a NotificationPolicy or a dashboard filter if a cascade should page differently.

Global Status Aggregation

There is no federation status CRD. The aggregated view is computed on request by the operator’s REST API from the ClusterRegistrations and the local Issues:

Remediation Policy per Tier

Each cluster has a remediation policy based on the tier of its ClusterRegistration, which controls how much autonomy the operator running in that cluster grants before the Decision Engine evaluates confidence.

Policy Definitions

Wiring

Tell the operator which registration is its own cluster:
or CHATCLI_OPERATOR_CLUSTER_NAME on the Deployment. The operator looks the value up among the ClusterRegistrations of its own cluster (any namespace), matching metadata.name or spec.displayName, so register the local cluster too; only the tier is read, so the gate works even while that registration is not connected. When no registration matches, the tier is unknown and the gate fails closed: every plan that reaches it waits for manual approval under the synthetic cluster-tier policy, with the reason Cluster name "<name>" (CHATCLI_OPERATOR_CLUSTER_NAME) is not registered: ... on the plan, until the registration exists. If the registrations cannot be listed, the plan stays Pending and the reconcile is retried. Empty (the default) disables the tier gate. With it set, every plan that reaches Pending and was not parked by an ApprovalPolicy is checked against the table: when the tier says “manual”, the plan waits in WaitingApproval under the synthetic policy cluster-tier (one approver, 30-minute timeout, decided like any ApprovalRequest) and its annotations carry decision-mode: approval and the reason, for example Cluster prod-us-east-1 tier policy "manual" requires approval for high severity.
The tier gate runs before the Decision Engine and independently of it: a plan the tier parks is not evaluated for confidence, and a plan the tier lets through is still subject to the engine when the engine is enabled.

YAML Examples

Register a Production Cluster

Register a Staging Cluster

Complete Multi-Region Setup

Federation Monitoring

Prometheus Metrics

Served on the operator metrics port (8080, /metrics):
The three states are exclusive: degraded is a cluster that answers the health check with some nodes not Ready; it is not counted as connected in the gauge (the REST connectedClusters does include it).

Network Architecture

The operator needs network connectivity to the API server of each registered cluster. In restricted network environments, consider using a bastion host or dedicated VPN for management traffic.

Next Steps

Decision Engine

Understand how confidence is calculated and how it combines with the tier gate.

Chaos Engineering

Run controlled chaos experiments to validate remediation.

Audit and Compliance

Audit trail of Issues, approvals and remediations.

AIOps Platform

Return to the AIOps platform overview.