Each experiment is a
ChaosExperiment custom resource reconciled by the
operator’s chaos controller. Guardrails (namespace allow/block lists,
minimum healthy pods, concurrency limit, abort on new Issue) are declarative
fields on the CR. Nothing is protected by default: every guardrail is
opt-in.Chaos Engineering in the AIOps Context
Remediation Validation
After fixing an incident, re-create the failure with a new experiment and
watch whether the platform detects and remediates it again.
Resilience Testing
Check that workloads come back to full readiness after pods are killed.
Game Days
Run experiments on demand. Recurring schedules are not supported (an
experiment with
schedule fails); use an external scheduler to create
experiments periodically.Drill-aware Analytics
Chaos-induced Issues are counted separately, so drills do not inflate
production incident numbers.
Automatic Correlation with Issues
When the anomaly controller creates an Issue, it looks for aChaosExperiment in the same namespace as the affected resource whose spec.target has the same kind, name and namespace, and that is either Running or finished (Completed/Aborted) less than 2 minutes ago. If it finds one, the Issue gets two labels:
The REST
analytics/mttd and analytics/mttr endpoints leave chaos drills
out, like the Prometheus resolution histogram. NotificationPolicy
rules cannot match on labels, so drills still reach every channel whose rule
matches the Issue’s severity, namespace, kind or state. Run drills in a
dedicated namespace if you want to route them separately.ChaosExperiment CRD
Short name:chaos. Printer columns: Type, Target, State, Duration, Age.
Complete Specification
The controller never writes
status.conditions. Use status.state and
status.result (a human-readable summary) to follow an experiment.7 Experiment Types
The failure is injected once, on the first reconcile after the experiment entersRunning. duration is the observation window: the controller waits for it to elapse (checking every 10 seconds at most) and then runs the post-experiment steps. Target pods are the pods owned by the Deployment (through its ReplicaSets) or by the StatefulSet. Any other target.kind makes the experiment fail.
Parameters that are missing or not valid integers fall back to their defaults silently.
1. Pod Kill
Deletes randomly chosen target pods withgracePeriodSeconds: 0. The random selection is a Fisher-Yates shuffle using crypto/rand.
2. Pod Failure
Same selection aspod_kill, but deletes with a grace period taken from the experiment (it overrides the pod’s own terminationGracePeriodSeconds).
3. CPU Stress
Creates astress-ng pod pinned to the node of the first target pod. It stresses the node, not the target container.
4. Memory Stress
Creates astress-ng pod that allocates memory on the target’s node: stress-ng --vm <workers> --vm-bytes <bytes> --timeout <duration>.
Because of the 512Mi limit, asking for more memory gets the stress container OOM-killed rather than putting pressure on the node.
5. Network Delay
Not supported. The type is accepted by the CRD for compatibility, but injecting latency needs a privileged tc/netem injector that the operator does not deploy. The experiment fails as soon as it is picked up, while stillPending and before any safety check, approval or dry run; no pod is touched:
6. Network Loss
Not supported, for the same reason asnetwork_delay. An experiment of this type fails in Pending with result: "network_loss is not supported: ... nothing was injected", and no packet is dropped.
7. Disk Stress
Creates astress-ng pod that generates disk I/O on the target’s node: stress-ng --hdd <workers> --hdd-bytes <size> --timeout <duration>. The I/O goes to the stress container’s own filesystem.
Type Summary
Granting Pod Permissions
The chart (rbac.create: true) and operator/config/rbac/role.yaml already grant these verbs. If you run the operator with RBAC you manage yourself (rbac.create: false), the stress types need create on pods; without it the stress types fail (execution failed: creating ... stress pod: ... forbidden). For example, cluster-wide:
Role/RoleBinding in the namespaces you run experiments in instead.
Safety Checks
Safety checks run in this order while the experiment isPending. All of them are evaluated once, before the failure is injected, except abortOnIssueDetected, which is checked while the experiment runs.
AllowedNamespaces / BlockedNamespaces
Checked againstspec.target.namespace. A blocked namespace, or a namespace missing from a non-empty allow list, moves the experiment to Failed (namespace "<ns>" is not allowed by safety checks).
- AllowedNamespaces
- BlockedNamespaces
Allow list. If it is set, only these namespaces can be targeted. If it
is empty, every namespace that is not blocked is allowed.
MaxConcurrentExperiments
Counts experiments inRunning state across all namespaces and compares the count with this experiment’s own maxConcurrentExperiments (default 1; 0 or less is treated as 1). When the limit is reached, the experiment stays Pending and is retried every 15 seconds. It does not fail.
MinHealthyPods
If greater than0, the controller counts the target pods whose Ready condition is True, subtracts the count parameter (default 1, whatever the experiment type), and moves the experiment to Failed when the result is below minHealthyPods:
RequireApproval
Whentrue, the controller looks for an ApprovalRequest named chaos-<experiment-name> in the experiment’s namespace, creates it if missing (owned by the experiment, so it is deleted with it), and re-checks every 30 seconds. The experiment starts only when that request’s status.state is Approved. If the request is Rejected or Expired, the experiment moves to Failed (approval rejected by <approver>: <reason> or approval request chaos-<experiment-name> expired without a decision).
The request the controller builds looks like this:
policyRef: chaos-safety is not a built-in policy. When an ApprovalPolicy named chaos-safety with a rule named chaos-experiment-approval exists in the experiment’s namespace, that rule decides the request (for example a quorum or a change window). Without it, the request is evaluated with its own settings: one human approval, expiring after 30 minutes. Approve or reject it like any other request, with the platform.chatcli.io/approve / reject annotation, the REST API or the dashboard:AbortOnIssueDetected
While the experiment isRunning and its duration has not elapsed, each reconcile (every 10 seconds at most) lists the Issues in the target namespace. If any Issue has exactly the same spec.resource kind, name and namespace as the target and was created after status.startedAt, the experiment moves to Aborted with result: "aborted: new issue detected: <issue>". Stress pods are then deleted.
Post-Experiment Verification
Whenduration has elapsed, the controller deletes the stress pods, stores postExperimentSnapshot, and runs the steps below. The experiment always ends in Completed from here, even if recovery was not verified.
VerifyRecovery
The target counts as healthy when a Deployment hasreadyReplicas and availableReplicas at least equal to spec.replicas, or a StatefulSet has readyReplicas at least equal to spec.replicas. If it is not healthy yet, the controller re-checks every 5 seconds.
RecoveryTimeout
How long to keep re-checking afterduration has elapsed. Default 5m. An invalid value silently falls back to 5m. When the timeout passes, the experiment is marked Completed with recoveryVerified: false, recoveryTime: "timeout", and a result ending in recovery NOT verified (timeout). It is not marked Failed.
recoveryTime (and the chatcli_operator_chaos_recovery_time_seconds histogram) is measured from the end of the injection (startedAt + duration) to the first health check that passes. Checks run every 5 seconds, so the value has that granularity; a target that is already healthy when the experiment completes records the few seconds between the end of the injection and that check.
RunRemediationTest
This option does not re-inject the failure. At completion, the controller reads the Issue named inlinkedIssueRef (in the experiment’s namespace) and only checks whether its status.state is Resolved:
Resolved:result: "completed: <type> experiment finished; remediation validation passed"- anything else, or no
linkedIssueRef:result: "completed: <type> experiment finished; no remediation was applied during experiment"
Resolved before the experiment started, the check passes regardless of what happened during the experiment. When runRemediationTest is true, the result text does not mention recovery; read recoveryVerified for that.
State Machine
Terminal experiments are never run again. To repeat an experiment, delete it and create it again (resetting the status by hand does not re-inject the failure, because the controller skips injection when
podsAffected is already set).
There is no field to abort a running experiment. Deleting the CR or setting enabled: false makes the controller stop handling it, but it does not clean up: stress pods keep running until stress-ng’s --timeout.
duration is only parsed once the experiment is Running, so an invalid
value (for example 1d) is caught after the safety checks and approval have
passed.DryRun Mode
WithdryRun: true, the controller runs the Pending checks (namespaces, concurrency, minHealthyPods, and approval if required) and then moves the experiment straight to Completed, without touching any pod. It does not select pods or plan individual actions.
chatcli_operator_chaos_experiments_total{result="dry_run"}.
Schedule (Recurring Experiments)
For recurring game days, create a freshChaosExperiment on a schedule from outside the operator, for example with a Kubernetes CronJob that runs kubectl create -f on a manifest that uses metadata.generateName (so each run gets a new name and a new status). Clean up old experiments yourself; the operator does not prune them.
LinkedIssueRef
linkedIssueRef takes only a name; the Issue is looked up in the experiment’s namespace. It is used in two places:
- Approval: when
requireApprovalistrue, it becomes thespec.issueRefof the generated ApprovalRequest. runRemediationTest: at completion, the controller checks whether this Issue isResolvedand writes the outcome intostatus.result.
Complete YAML Examples
Pod Kill with Safety Checks
Pod Kill with Safety Checks
CPU Stress (requires pod create permission)
CPU Stress (requires pod create permission)
Post-Remediation Validation
Post-Remediation Validation
DryRun for Configuration Validation
DryRun for Configuration Validation
Metrics
The chaos controller registers three metrics on the operator’s metrics endpoint (port 8080 by default):result="failed" means a safety check or the injection failed. A target that did not recover in time is still counted as completed.
Alert Examples
status.recoveryVerified: false on Completed experiments); there is no metric for it.
Best Practices
1
Start with DryRun
Confirm that the namespace lists and
minHealthyPods accept the target
before running a real experiment.2
Staging First
Run experiments in staging before higher-tier environments. Set
allowedNamespaces and blockedNamespaces explicitly; nothing is
blocked by default.3
Conservative Safety Checks
Configure
minHealthyPods with margin. If the deployment has 5 replicas
and needs 3 to operate, set minHealthyPods: 3, and keep count low:
a count at or above the replica count deletes every pod.4
Choose abortOnIssueDetected Deliberately
Turn it on to stop a drill as soon as AIOps reacts; turn it off when the
point of the drill is to let AIOps detect and remediate.
5
Recreate for Every Run
Experiments run once. Use an external scheduler that creates a new CR
(with
generateName) for recurring game days.Next Steps
Approval Workflow
How ApprovalRequests are decided, used by
requireApproval.Incident Lifecycle
The Issue pipeline that chaos-induced Issues go through.
Web Dashboard
Filter chaos drills in or out of the incident views.
AIOps Platform
Return to the AIOps platform overview.