Skip to main content
The Chaos Engineering module lets you inject failures into Kubernetes workloads and watch how the AIOps platform reacts. Unlike standalone chaos tools, experiments here are wired into the AIOps pipeline: Issues raised while an experiment runs are tagged as chaos drills, so they do not page anyone and do not pollute the production MTTR histogram.
Each experiment is a ChaosExperiment custom resource reconciled by the operator’s chaos controller. Guardrails (namespace allow/block lists, minimum healthy pods, concurrency limit, abort on new Issue) are declarative fields on the CR. Nothing is protected by default: every guardrail is opt-in.
What works out of the box. The RBAC shipped by the operator Helm chart and operator/config/rbac/role.yaml lets the operator get, list, watch, create, update and delete pods, so every supported experiment type can run:
  • pod_kill and pod_failure delete target pods.
  • cpu_stress, memory_stress, disk_stress create a stress pod that passes the restricted Pod Security Standard.
  • network_delay and network_loss are not supported. The CRD accepts them, but an experiment of either type fails at once in Pending with a result that says nothing was injected: network faults need a privileged tc/netem injector the operator does not deploy.
If you manage the operator’s RBAC yourself, see Granting pod permissions below.

Chaos Engineering in the AIOps Context

Remediation Validation

After fixing an incident, re-create the failure with a new experiment and watch whether the platform detects and remediates it again.

Resilience Testing

Check that workloads come back to full readiness after pods are killed.

Game Days

Run experiments on demand. Recurring schedules are not supported (an experiment with schedule fails); use an external scheduler to create experiments periodically.

Drill-aware Analytics

Chaos-induced Issues are counted separately, so drills do not inflate production incident numbers.

Automatic Correlation with Issues

When the anomaly controller creates an Issue, it looks for a ChaosExperiment in the same namespace as the affected resource whose spec.target has the same kind, name and namespace, and that is either Running or finished (Completed/Aborted) less than 2 minutes ago. If it finds one, the Issue gets two labels:
The Issue then goes through the normal AIOps pipeline, with these differences:
The REST analytics/mttd and analytics/mttr endpoints leave chaos drills out, like the Prometheus resolution histogram. NotificationPolicy rules cannot match on labels, so drills still reach every channel whose rule matches the Issue’s severity, namespace, kind or state. Run drills in a dedicated namespace if you want to route them separately.
Create the ChaosExperiment in the same namespace as its target. The correlation lookup only searches the target’s namespace, while linkedIssueRef and the approval request are resolved in the experiment’s namespace.

ChaosExperiment CRD

Short name: chaos. Printer columns: Type, Target, State, Duration, Age.

Complete Specification

The controller never writes status.conditions. Use status.state and status.result (a human-readable summary) to follow an experiment.

7 Experiment Types

The failure is injected once, on the first reconcile after the experiment enters Running. duration is the observation window: the controller waits for it to elapse (checking every 10 seconds at most) and then runs the post-experiment steps. Target pods are the pods owned by the Deployment (through its ReplicaSets) or by the StatefulSet. Any other target.kind makes the experiment fail. Parameters that are missing or not valid integers fall back to their defaults silently.

1. Pod Kill

Deletes randomly chosen target pods with gracePeriodSeconds: 0. The random selection is a Fisher-Yates shuffle using crypto/rand.
Pod kill simulates an abrupt failure (e.g., node crash). Use safetyChecks.minHealthyPods to avoid taking down every replica.

2. Pod Failure

Same selection as pod_kill, but deletes with a grace period taken from the experiment (it overrides the pod’s own terminationGracePeriodSeconds).

3. CPU Stress

Creates a stress-ng pod pinned to the node of the first target pod. It stresses the node, not the target container.
Generated stress pod (same shape for the three stress types; only the name prefix and command change):
The stress pod’s resources are fixed: it can use at most 500m CPU and 512Mi memory, whatever cores, loadPercent or bytes say. The image (alexeiled/stress-ng:latest, from Docker Hub) cannot be changed. The pod passes the restricted Pod Security Standard: it runs as UID/GID 65534, non-root, with the RuntimeDefault seccomp profile, no privilege escalation, all capabilities dropped, a read-only root filesystem (an emptyDir at /tmp is the only writable path) and no ServiceAccount token.

4. Memory Stress

Creates a stress-ng pod that allocates memory on the target’s node: stress-ng --vm <workers> --vm-bytes <bytes> --timeout <duration>.
Because of the 512Mi limit, asking for more memory gets the stress container OOM-killed rather than putting pressure on the node.

5. Network Delay

Not supported. The type is accepted by the CRD for compatibility, but injecting latency needs a privileged tc/netem injector that the operator does not deploy. The experiment fails as soon as it is picked up, while still Pending and before any safety check, approval or dry run; no pod is touched:
Use a dedicated network-chaos tool for latency tests.

6. Network Loss

Not supported, for the same reason as network_delay. An experiment of this type fails in Pending with result: "network_loss is not supported: ... nothing was injected", and no packet is dropped.

7. Disk Stress

Creates a stress-ng pod that generates disk I/O on the target’s node: stress-ng --hdd <workers> --hdd-bytes <size> --timeout <duration>. The I/O goes to the stress container’s own filesystem.

Type Summary

Granting Pod Permissions

The chart (rbac.create: true) and operator/config/rbac/role.yaml already grant these verbs. If you run the operator with RBAC you manage yourself (rbac.create: false), the stress types need create on pods; without it the stress types fail (execution failed: creating ... stress pod: ... forbidden). For example, cluster-wide:
To limit the blast radius, bind the same rules with a Role/RoleBinding in the namespaces you run experiments in instead.

Safety Checks

Safety checks run in this order while the experiment is Pending. All of them are evaluated once, before the failure is injected, except abortOnIssueDetected, which is checked while the experiment runs.

AllowedNamespaces / BlockedNamespaces

Checked against spec.target.namespace. A blocked namespace, or a namespace missing from a non-empty allow list, moves the experiment to Failed (namespace "<ns>" is not allowed by safety checks).
Allow list. If it is set, only these namespaces can be targeted. If it is empty, every namespace that is not blocked is allowed.
blockedNamespaces takes precedence over allowedNamespaces. There are no built-in blocked namespaces: kube-system and chatcli-system are only protected if you list them.

MaxConcurrentExperiments

Counts experiments in Running state across all namespaces and compares the count with this experiment’s own maxConcurrentExperiments (default 1; 0 or less is treated as 1). When the limit is reached, the experiment stays Pending and is retried every 15 seconds. It does not fail.

MinHealthyPods

If greater than 0, the controller counts the target pods whose Ready condition is True, subtracts the count parameter (default 1, whatever the experiment type), and moves the experiment to Failed when the result is below minHealthyPods:
This check runs once, before injection. It does not watch pod health during the experiment.

RequireApproval

When true, the controller looks for an ApprovalRequest named chaos-<experiment-name> in the experiment’s namespace, creates it if missing (owned by the experiment, so it is deleted with it), and re-checks every 30 seconds. The experiment starts only when that request’s status.state is Approved. If the request is Rejected or Expired, the experiment moves to Failed (approval rejected by <approver>: <reason> or approval request chaos-<experiment-name> expired without a decision). The request the controller builds looks like this:
policyRef: chaos-safety is not a built-in policy. When an ApprovalPolicy named chaos-safety with a rule named chaos-experiment-approval exists in the experiment’s namespace, that rule decides the request (for example a quorum or a change window). Without it, the request is evaluated with its own settings: one human approval, expiring after 30 minutes. Approve or reject it like any other request, with the platform.chatcli.io/approve / reject annotation, the REST API or the dashboard:

AbortOnIssueDetected

While the experiment is Running and its duration has not elapsed, each reconcile (every 10 seconds at most) lists the Issues in the target namespace. If any Issue has exactly the same spec.resource kind, name and namespace as the target and was created after status.startedAt, the experiment moves to Aborted with result: "aborted: new issue detected: <issue>". Stress pods are then deleted.
The check does not distinguish Issues caused by the experiment itself. If AIOps detects the failure you injected (the usual goal of a drill), the experiment aborts. Leave abortOnIssueDetected: false when you expect AIOps to raise an Issue for the target.

Post-Experiment Verification

When duration has elapsed, the controller deletes the stress pods, stores postExperimentSnapshot, and runs the steps below. The experiment always ends in Completed from here, even if recovery was not verified.

VerifyRecovery

The target counts as healthy when a Deployment has readyReplicas and availableReplicas at least equal to spec.replicas, or a StatefulSet has readyReplicas at least equal to spec.replicas. If it is not healthy yet, the controller re-checks every 5 seconds.

RecoveryTimeout

How long to keep re-checking after duration has elapsed. Default 5m. An invalid value silently falls back to 5m. When the timeout passes, the experiment is marked Completed with recoveryVerified: false, recoveryTime: "timeout", and a result ending in recovery NOT verified (timeout). It is not marked Failed. recoveryTime (and the chatcli_operator_chaos_recovery_time_seconds histogram) is measured from the end of the injection (startedAt + duration) to the first health check that passes. Checks run every 5 seconds, so the value has that granularity; a target that is already healthy when the experiment completes records the few seconds between the end of the injection and that check.

RunRemediationTest

This option does not re-inject the failure. At completion, the controller reads the Issue named in linkedIssueRef (in the experiment’s namespace) and only checks whether its status.state is Resolved:
  • Resolved: result: "completed: <type> experiment finished; remediation validation passed"
  • anything else, or no linkedIssueRef: result: "completed: <type> experiment finished; no remediation was applied during experiment"
If the linked Issue was already Resolved before the experiment started, the check passes regardless of what happened during the experiment. When runRemediationTest is true, the result text does not mention recovery; read recoveryVerified for that.

State Machine

Terminal experiments are never run again. To repeat an experiment, delete it and create it again (resetting the status by hand does not re-inject the failure, because the controller skips injection when podsAffected is already set). There is no field to abort a running experiment. Deleting the CR or setting enabled: false makes the controller stop handling it, but it does not clean up: stress pods keep running until stress-ng’s --timeout.
duration is only parsed once the experiment is Running, so an invalid value (for example 1d) is caught after the safety checks and approval have passed.

DryRun Mode

With dryRun: true, the controller runs the Pending checks (namespaces, concurrency, minHealthyPods, and approval if required) and then moves the experiment straight to Completed, without touching any pod. It does not select pods or plan individual actions.
DryRun result:
A dry run increments chatcli_operator_chaos_experiments_total{result="dry_run"}.
Run a dry run first to confirm that the namespace lists and minHealthyPods accept the target. If a safety check fails, the dry run ends in Failed with the same message a real run would get.

Schedule (Recurring Experiments)

The schedule field exists in the CRD but recurring experiments are not supported. An experiment that sets schedule fails at once in Pending, without injecting anything, with result: "schedule is not supported: recurring experiments are not implemented; remove spec.schedule and create one ChaosExperiment per run".
For recurring game days, create a fresh ChaosExperiment on a schedule from outside the operator, for example with a Kubernetes CronJob that runs kubectl create -f on a manifest that uses metadata.generateName (so each run gets a new name and a new status). Clean up old experiments yourself; the operator does not prune them.

LinkedIssueRef

linkedIssueRef takes only a name; the Issue is looked up in the experiment’s namespace. It is used in two places:
  1. Approval: when requireApproval is true, it becomes the spec.issueRef of the generated ApprovalRequest.
  2. runRemediationTest: at completion, the controller checks whether this Issue is Resolved and writes the outcome into status.result.
The controller does not modify the Issue or its RemediationPlan, and does not record the link anywhere else in the experiment status.

Complete YAML Examples

Metrics

The chaos controller registers three metrics on the operator’s metrics endpoint (port 8080 by default): result="failed" means a safety check or the injection failed. A target that did not recover in time is still counted as completed.

Alert Examples

To alert on unverified recovery, query the CRs (status.recoveryVerified: false on Completed experiments); there is no metric for it.

Best Practices

1

Start with DryRun

Confirm that the namespace lists and minHealthyPods accept the target before running a real experiment.
2

Staging First

Run experiments in staging before higher-tier environments. Set allowedNamespaces and blockedNamespaces explicitly; nothing is blocked by default.
3

Conservative Safety Checks

Configure minHealthyPods with margin. If the deployment has 5 replicas and needs 3 to operate, set minHealthyPods: 3, and keep count low: a count at or above the replica count deletes every pod.
4

Choose abortOnIssueDetected Deliberately

Turn it on to stop a drill as soon as AIOps reacts; turn it off when the point of the drill is to let AIOps detect and remediate.
5

Recreate for Every Run

Experiments run once. Use an external scheduler that creates a new CR (with generateName) for recurring game days.

Next Steps

Approval Workflow

How ApprovalRequests are decided, used by requireApproval.

Incident Lifecycle

The Issue pipeline that chaos-induced Issues go through.

Web Dashboard

Filter chaos drills in or out of the incident views.

AIOps Platform

Return to the AIOps platform overview.