ClusterRegistration; the operator health-checks them, correlates new Issues with the Issues running there, flags staging-to-production cascades, and applies a remediation policy based on its own cluster’s tier.
Federation does not require a service mesh or external tools. The operator
connects directly to each registered cluster with a kubeconfig stored in a Secret.
Why Multi-Cluster Federation?
In modern production environments, infrastructure is rarely limited to a single cluster:Multi-Region
Clusters in us-east-1, eu-west-1, and ap-southeast-1 for latency and
regional compliance.
Multi-Environment
Staging, production, and DR in separate clusters with different security
policies.
Multi-Tenant
Dedicated clusters per team or product with strong workload isolation.
- Detect that the same problem affects 5 clusters simultaneously
- Correlate a staging deploy with a production failure
- Apply differentiated remediation policies by cluster importance
- Aggregate cluster health into a global view
ClusterRegistration CRD
TheClusterRegistration CRD is the entry point for adding clusters to the federation. Create it in the cluster where the operator runs, in any namespace; the kubeconfig Secret must live in the same namespace.
Complete Specification
Spec Fields
Status Fields
How Federation Works
Kubeconfig and client cache
On each reconcile the FederationReconciler readsdata.kubeconfig from the Secret named in kubeconfigSecretRef, in the ClusterRegistration’s namespace, builds a controller-runtime client from it, and caches that client by registration name. The cache is dropped when a health check fails or the registration is deleted, so a rotated kubeconfig is picked up after the next failed check (or an operator restart).
The kubeconfig identity needs, on the remote cluster: list on nodes and namespaces for the health check; list on issues and remediationplans (group platform.chatcli.io) for the counters and correlation; and update on issues for correlation to annotate them.
Health Check Loop
1
List Nodes
Lists the remote nodes to verify connectivity, count them and read the first node’s kubelet version.
2
List Namespaces
Lists the remote namespaces and counts them.
3
Count AIOps work
Lists remote Issues and RemediationPlans, when those CRDs exist there, to fill
activeIssues and activeRemediations.4
Update Status
Writes
connected, lastHealthCheck, kubernetesVersion, nodeCount, namespaceCount, activeIssues, activeRemediations and the Connected/Degraded conditions, refreshes the chatcli_operator_federation_clusters_total gauge, then requeues after healthCheckInterval. A failure at any step before this one sets connected: false, refreshes lastHealthCheck and records the failure in the Connected condition.Cross-Cluster Correlation
When an Issue is detected, the IssueReconciler asks the federation controller to look for the same problem elsewhere. Both checks below run only against ClusterRegistrations that areconnected and whose kubeconfig client is cached; with no registrations they cost two cached lists and change nothing.
Automatic Severity Elevation
The controller lists the active (non-terminal) Issues with the samesignalType on the local cluster and on every connected cluster. When they span 3 or more clusters (the local one counts as one), every matching Issue, local and remote, is annotated and raised to critical:
critical also gets platform.chatcli.io/elevated-severity: "true". One correlation keeps one id: when a new matching Issue arrives and one of the matching Issues already carries a correlation id, the group reuses it instead of relabeling everything, and affected-clusters is refreshed. The counter chatcli_operator_federation_cross_cluster_issues_total increments once per new correlation id. An Issue that changed underneath the update (a conflict) is skipped and not retried; the next correlation of the group annotates it.
The correlation id is an annotation, so find the related Issues with:
GET /api/v1/federation/correlations lists the correlations from the local Issues, one entry per correlation id, with correlationId, issue, namespace, severity, signalType, correlatedClusters (from affected-clusters), elevated (from elevated-severity) and cascade. It also reads the older platform.chatcli.io/correlation-id / correlated-clusters keys, so Issues annotated with them still show up.
Cascade Detection
Cascade detection flags a newly detected Issue whose resource (same name) already failed earlier in a staging cluster: the same problem is rolling through the environments.Staging to Production
The check needs at least one connected registration withenvironment: staging and one with environment: prod. It runs for every Issue the local operator detects (the local cluster’s own environment is not consulted); for that Issue, the controller lists the Issues on every staging cluster and looks for one on a resource with the same name whose detectedAt is earlier than the local Issue’s. The first match annotates the local Issue:
chatcli_operator_federation_cascade_detected_total.
The annotation is the whole effect: severity, notifications and the AI prompt are not changed by a cascade. Route on it from a NotificationPolicy or a dashboard filter if a cascade should page differently.
Global Status Aggregation
There is no federation status CRD. The aggregated view is computed on request by the operator’s REST API from the ClusterRegistrations and the local Issues:- kubectl
- REST API
Remediation Policy per Tier
Each cluster has a remediation policy based on thetier of its ClusterRegistration, which controls how much autonomy the operator running in that cluster grants before the Decision Engine evaluates confidence.
Policy Definitions
Wiring
Tell the operator which registration is its own cluster:CHATCLI_OPERATOR_CLUSTER_NAME on the Deployment. The operator looks the value up among the ClusterRegistrations of its own cluster (any namespace), matching metadata.name or spec.displayName, so register the local cluster too; only the tier is read, so the gate works even while that registration is not connected. When no registration matches, the tier is unknown and the gate fails closed: every plan that reaches it waits for manual approval under the synthetic cluster-tier policy, with the reason Cluster name "<name>" (CHATCLI_OPERATOR_CLUSTER_NAME) is not registered: ... on the plan, until the registration exists. If the registrations cannot be listed, the plan stays Pending and the reconcile is retried. Empty (the default) disables the tier gate. With it set, every plan that reaches Pending and was not parked by an ApprovalPolicy is checked against the table: when the tier says “manual”, the plan waits in WaitingApproval under the synthetic policy cluster-tier (one approver, 30-minute timeout, decided like any ApprovalRequest) and its annotations carry decision-mode: approval and the reason, for example Cluster prod-us-east-1 tier policy "manual" requires approval for high severity.
The tier gate runs before the Decision Engine and independently of it: a plan the tier parks is not evaluated for confidence, and a plan the tier lets through is still subject to the engine when the engine is enabled.
YAML Examples
Register a Production Cluster
Register a Staging Cluster
Complete Multi-Region Setup
Federation Monitoring
Prometheus Metrics
Served on the operator metrics port (8080,/metrics):
The three states are exclusive:
degraded is a cluster that answers the health check with some nodes not Ready; it is not counted as connected in the gauge (the REST connectedClusters does include it).Recommended Dashboards
Grafana Dashboard: Federation Overview
Grafana Dashboard: Federation Overview
Recommended Alerts
Network Architecture
Next Steps
Decision Engine
Understand how confidence is calculated and how it combines with the tier gate.
Chaos Engineering
Run controlled chaos experiments to validate remediation.
Audit and Compliance
Audit trail of Issues, approvals and remediations.
AIOps Platform
Return to the AIOps platform overview.