Skip to main content
The ChatCLI AIOps platform records the key steps of its pipeline — Issue lifecycle, remediation execution, approval gates, notification deliveries, SLO burn-rate alerts and SLA breaches — as AuditEvent resources. Combined with the platform ClusterRoles and an on-demand compliance report, this gives you a queryable trail of what the automation did and when.
An AuditEvent is a CRD with only a spec (no status subresource). The operator creates AuditEvents and never updates or deletes them, but nothing in the cluster enforces immutability: there is no admission webhook, and the shipped chatcli-role-admin ClusterRole can update/patch AuditEvents (chatcli-role-superadmin can also delete them). If you need a tamper-proof trail, see Immutability and ship the events to an external system.

Why Audit Trail for AIOps

When a platform makes autonomous decisions on production infrastructure, traceability is no longer optional:

Accountability

When was a remediation started, parked for approval, approved, rejected or expired? Which controller did it? Every one of those steps leaves a record.

Post-Incident Investigation

Every event carries a correlationId (the Issue name), so the trail of an incident — creation, remediation, notifications, resolution — can be pulled with one filter.

Regulatory Compliance

Evidence for change-control audits (SOC 2, ISO 27001, PCI-DSS and similar): a record of automated actions plus documented RBAC. The platform provides the records; it is not certified against any framework.

Continuous Improvement

MTTD, MTTR, remediation success rate and approval outcomes, computed on demand from the platform CRDs.

AuditEvent CRD

The AuditEvent has only spec, no status. Short name: ae.

Complete Specification

A real event, as written by the remediation controller when a plan starts executing:
Field reference (operator/api/v1alpha1/auditevent_types.go):

Event Types (EventType)

The operator emits 14 event types, the same list the CRD’s eventType field comment documents. All are written with actor.type: controller, except an approval or rejection decided by people (see the Governance tab):
No other event type is written. Names that older material listed, such as approval_granted (the real name is approval_approved), pattern_learned, config_changed, cluster_connected, cluster_disconnected, escalation_triggered, postmortem_created and runbook_generated, never appear. Chaos experiments, the decision engine’s confidence verdicts, AI analysis, anomaly detection, federation and REST API calls produce no AuditEvents of their own. Do not build alerts or reports on those names.

AuditActor

The actor field identifies who or what performed the action: Other human actions (acknowledging or snoozing an incident, kubectl edits) are not turned into AuditEvents. Use the Kubernetes API server audit log for who changed which object.

AuditResource

The resource field identifies the affected Kubernetes resource:

Name Format and Namespace

Each AuditEvent is named:
Example: audit-1773930600123456789-a7f3b2. The event is created in the namespace of the affected resource (the Issue’s, RemediationPlan’s, ApprovalRequest’s or SLO’s namespace), not in the operator namespace. Query with -A or with the workload’s namespace. Only two labels are set: platform.chatcli.io/event-type and platform.chatcli.io/severity. There is no correlation label; filter on spec.correlationId instead (examples below).

Immutability Annotation

Every AuditEvent is created with the annotation platform.chatcli.io/immutable: "true". The operator ships no admission webhook, so the annotation is only a marker. To make the trail tamper-resistant:
  • enforce it with a policy engine rule (Kyverno, Gatekeeper) that rejects UPDATE and DELETE on resources carrying the annotation, except for your retention job;
  • review who holds update/patch/delete on auditevents (the operator’s own ServiceAccount, chatcli-role-admin and chatcli-role-superadmin all do);
  • export the events to a SIEM or write-once storage.
The operator writes AuditEvents best-effort: if a create fails, the controller logs the error and carries on, so a gap in the trail does not block remediation.

Audit Recorder

The AuditRecorder (operator/controllers/audit_recorder.go) is the internal component the controllers call to write AuditEvents. It is not a public API or extension point: one recorder is created at startup and shared by the Issue, Remediation, Notification, SLO and SLA reconcilers. The Approval, Chaos, Federation, AIInsight, Anomaly and PostMortem reconcilers do not hold one.

Which Controller Writes What

Generated Event Example

An SLA breach, as written by the SLA controller:

Compliance Reporter

The ComplianceReporter computes an on-demand report and is served by the REST API at GET /api/v1/analytics/compliance (viewer role). Nothing is scheduled or stored: each call lists the platform CRDs and computes the figures.

Requesting a Report

The report covers objects created inside the period (by metadata.creationTimestamp). It reads Issues, RemediationPlans, ApprovalRequests, IncidentSLAs and AuditEvents — AuditEvents only feed the audit summary; the other figures come from the resources themselves. The response wraps the report in spec. Keys are PascalCase (pinned by explicit JSON tags), and durations are integers in nanoseconds:

Report Metrics

Computed from Issues created in the window.

Audit Summary

Counts of AuditEvents created in the window, by severity and by event type:

Kubernetes RBAC Roles

The platform ships 4 ClusterRoles for people and tooling that work with the platform CRDs through kubectl. They are created by the operator Helm chart (rbac.create: true, the default) or by make deploy (operator/config/rbac/role.yaml), never by the operator at runtime (H5 hardening: the operator holds no permission to create ClusterRoles, and nothing in the operator binds these roles — you bind them yourself).

Role Definitions

chatcli-role-viewer — read-only access to the incident-facing CRDs.
Not included: policies (Approval, Notification, Escalation), ChaosExperiments, ClusterRegistrations, SourceRepositories, Instances.

Granting a Role

Bind a role to a user or group the way you bind any ClusterRole:
Use a namespaced RoleBinding to the same ClusterRole to limit access to one namespace. Revoking is deleting the binding. Role changes are visible in the Kubernetes audit log; the operator’s AuditEvents cover what the controllers do, not who was granted what.
The REST API and the dashboard use their own role model, API keys with viewer, operator or admin, described in Web Dashboard. The ClusterRoles above are for people and tooling that reach the CRDs through kubectl.

Audit REST API

The operator’s REST API (port 8090, header X-API-Key, any role from viewer up) exposes two read-only endpoints. Both accept only GET.

GET /api/v1/audit

Lists events, filtered and paginated in memory, newest first by timestamp (the creation time when it is missing; ties by name), so each page is a stable slice of the trail. Query parameters: There is no actor or correlation-ID filter; to get all events of one incident, filter correlationId client-side (see the kubectl and jq examples below). Request example:
Response example:
The REST view flattens the record: details becomes one detail string of key=value pairs joined by ; (in no fixed order), and actor.controller and resource.uid are dropped. Use kubectl get auditevent <name> -o yaml for the full object.

GET /api/v1/audit/export

Takes the same filters (namespace, type, severity, resource, from, to) but no pagination, and returns every matching event as a downloadable JSON document (Content-Disposition: attachment; filename=audit-events-<timestamp>.json):
Export format — a single indented JSON object (not NDJSON), whose items have the same shape as the list endpoint:
Use jq -c '.items[]' to turn it into one event per line.
The REST API is rate-limited to 600 requests per minute per valid API key (30 per minute per client host without one). Access to it is only logged to the operator’s stdout ([REST] method path status duration role=...); REST calls do not create AuditEvents.

SIEM Integration

There is no built-in SIEM push. Pull the export endpoint on a schedule and forward it to Splunk, Elastic, Datadog or any other collector. The examples below use alpine with curl and jq, and an API key stored in a Secret (chatcli-audit-exporter, key api-key, holding a viewer key).

Splunk

1

Configure HEC (HTTP Event Collector)

Create an HEC token in Splunk to receive events from the AIOps platform.
2

Create export CronJob

3

Create index and dashboards

Configure a dedicated chatcli_audit index in Splunk and create dashboards to visualize events by type, severity and namespace.

Elasticsearch

If the operator’s REST API runs with TLS (CHATCLI_AIOPS_TLS_CERT / CHATCLI_AIOPS_TLS_KEY), switch the URLs to https:// and pass the CA with --cacert.

kubectl Commands

Event Retention

The operator never deletes AuditEvents — there is no TTL, retention setting or owner reference, so they are not garbage-collected with their Issue. Every event is an object in etcd; configure a retention job to keep the count bounded.
Export to your SIEM before the retention window closes if you need to keep events longer.

Server-Side Audit Trail

The AuditEvents above cover the operator. The ChatCLI server (chatcli server, the pod an Instance runs) has a separate, file-based trail: set spec.server.security.auditLogPath on the Instance (it becomes CHATCLI_AUDIT_LOG_PATH) to an absolute path, and every gRPC call is appended as a hash-chained JSON line (kind: "grpc": action, actor, role, caller IP, result, duration), interleaved with the LLM request entries in the same verifiable chain. Details and verification (/config security verify-audit) are in Security. The Instance pod has a read-only root filesystem; only /tmp and /home/chatcli/.chatcli are writable, both emptyDir volumes lost on restart. To keep the file across restarts, enable spec.persistence and point the path into the sessions volume, e.g. /home/chatcli/.chatcli/sessions/audit.jsonl.
The operator itself writes no file audit log. It does not read CHATCLI_AUDIT_LOG_PATH; the operator chart value security.auditLogPath still renders that variable but has no effect.

Next Steps

Approval Workflow

How plans are parked and decided — the source of the approval_* events, including gates raised by the decision engine and the cluster tier.

SLOs and SLAs

Burn-rate alerts and SLA timers behind slo_violation and sla_breach, and the real per-severity SLA compliance.

Web Dashboard

The audit view, API keys and REST roles.

AIOps Platform

Return to the AIOps platform overview.