AuditEvent resources. Combined with the platform ClusterRoles and an on-demand compliance report, this gives you a queryable trail of what the automation did and when.
An AuditEvent is a CRD with only a
spec (no status subresource). The
operator creates AuditEvents and never updates or deletes them, but nothing
in the cluster enforces immutability: there is no admission webhook, and
the shipped chatcli-role-admin ClusterRole can update/patch AuditEvents
(chatcli-role-superadmin can also delete them). If you need a
tamper-proof trail, see Immutability and ship the
events to an external system.Why Audit Trail for AIOps
When a platform makes autonomous decisions on production infrastructure, traceability is no longer optional:Accountability
When was a remediation started, parked for approval, approved, rejected or
expired? Which controller did it? Every one of those steps leaves a record.
Post-Incident Investigation
Every event carries a
correlationId (the Issue name), so the trail of an
incident — creation, remediation, notifications, resolution — can be
pulled with one filter.Regulatory Compliance
Evidence for change-control audits (SOC 2, ISO 27001, PCI-DSS and similar):
a record of automated actions plus documented RBAC. The platform provides
the records; it is not certified against any framework.
Continuous Improvement
MTTD, MTTR, remediation success rate and approval outcomes, computed on
demand from the platform CRDs.
AuditEvent CRD
TheAuditEvent has only spec, no status. Short name: ae.
Complete Specification
A real event, as written by the remediation controller when a plan starts executing:operator/api/v1alpha1/auditevent_types.go):
Event Types (EventType)
The operator emits 14 event types, the same list the CRD’seventType field comment documents. All are written with actor.type: controller, except an approval or rejection decided by people (see the Governance tab):
- Incidents
- Remediation
- Governance
- Alerts and delivery
AuditActor
Theactor field identifies who or what performed the action:
Other human actions (acknowledging or snoozing an incident,
kubectl edits) are not turned into AuditEvents. Use the Kubernetes API server audit log for who changed which object.
AuditResource
Theresource field identifies the affected Kubernetes resource:
Name Format and Namespace
Each AuditEvent is named:audit-1773930600123456789-a7f3b2.
The event is created in the namespace of the affected resource (the Issue’s, RemediationPlan’s, ApprovalRequest’s or SLO’s namespace), not in the operator namespace. Query with -A or with the workload’s namespace.
Only two labels are set: platform.chatcli.io/event-type and platform.chatcli.io/severity. There is no correlation label; filter on spec.correlationId instead (examples below).
Immutability Annotation
Every AuditEvent is created with the annotationplatform.chatcli.io/immutable: "true". The operator ships no admission webhook, so the annotation is only a marker. To make the trail tamper-resistant:
- enforce it with a policy engine rule (Kyverno, Gatekeeper) that rejects
UPDATEandDELETEon resources carrying the annotation, except for your retention job; - review who holds
update/patch/deleteonauditevents(the operator’s own ServiceAccount,chatcli-role-adminandchatcli-role-superadminall do); - export the events to a SIEM or write-once storage.
Audit Recorder
TheAuditRecorder (operator/controllers/audit_recorder.go) is the internal component the controllers call to write AuditEvents. It is not a public API or extension point: one recorder is created at startup and shared by the Issue, Remediation, Notification, SLO and SLA reconcilers. The Approval, Chaos, Federation, AIInsight, Anomaly and PostMortem reconcilers do not hold one.
Which Controller Writes What
Generated Event Example
An SLA breach, as written by the SLA controller:Compliance Reporter
TheComplianceReporter computes an on-demand report and is served by the REST API at GET /api/v1/analytics/compliance (viewer role). Nothing is scheduled or stored: each call lists the platform CRDs and computes the figures.
Requesting a Report
The report covers objects created inside the period (by
metadata.creationTimestamp). It reads Issues, RemediationPlans, ApprovalRequests, IncidentSLAs and AuditEvents — AuditEvents only feed the audit summary; the other figures come from the resources themselves.
The response wraps the report in spec. Keys are PascalCase (pinned by explicit JSON tags), and durations are integers in nanoseconds:
Report Metrics
- Incident Metrics
- Remediation Metrics
- SLA Metrics
- Approval Metrics
Computed from Issues created in the window.
Audit Summary
Counts of AuditEvents created in the window, by severity and by event type:Kubernetes RBAC Roles
The platform ships 4 ClusterRoles for people and tooling that work with the platform CRDs throughkubectl. They are created by the operator Helm chart (rbac.create: true, the default) or by make deploy (operator/config/rbac/role.yaml), never by the operator at runtime (H5 hardening: the operator holds no permission to create ClusterRoles, and nothing in the operator binds these roles — you bind them yourself).
Role Definitions
- Viewer
- Operator
- Admin
- SuperAdmin
chatcli-role-viewer — read-only access to the incident-facing CRDs.Granting a Role
Bind a role to a user or group the way you bind any ClusterRole:RoleBinding to the same ClusterRole to limit access to one namespace. Revoking is deleting the binding. Role changes are visible in the Kubernetes audit log; the operator’s AuditEvents cover what the controllers do, not who was granted what.
The REST API and the dashboard use their own role model, API keys with
viewer, operator or admin, described in Web Dashboard. The ClusterRoles above are for people and tooling that reach the CRDs through kubectl.Audit REST API
The operator’s REST API (port 8090, headerX-API-Key, any role from viewer up) exposes two read-only endpoints. Both accept only GET.
GET /api/v1/audit
Lists events, filtered and paginated in memory, newest first bytimestamp (the creation time when it is missing; ties by name), so each page is a stable slice of the trail.
Query parameters:
There is no actor or correlation-ID filter; to get all events of one incident, filter
correlationId client-side (see the kubectl and jq examples below).
Request example:
details becomes one detail string of key=value pairs joined by ; (in no fixed order), and actor.controller and resource.uid are dropped. Use kubectl get auditevent <name> -o yaml for the full object.
GET /api/v1/audit/export
Takes the same filters (namespace, type, severity, resource, from, to) but no pagination, and returns every matching event as a downloadable JSON document (Content-Disposition: attachment; filename=audit-events-<timestamp>.json):
items have the same shape as the list endpoint:
jq -c '.items[]' to turn it into one event per line.
The REST API is rate-limited to 600 requests per minute per valid API key (30 per minute per client host without one).
Access to it is only logged to the operator’s stdout (
[REST] method path status duration role=...); REST calls do not create AuditEvents.SIEM Integration
There is no built-in SIEM push. Pull the export endpoint on a schedule and forward it to Splunk, Elastic, Datadog or any other collector. The examples below usealpine with curl and jq, and an API key stored in a Secret (chatcli-audit-exporter, key api-key, holding a viewer key).
Splunk
1
Configure HEC (HTTP Event Collector)
Create an HEC token in Splunk to receive events from the AIOps platform.
2
Create export CronJob
3
Create index and dashboards
Configure a dedicated
chatcli_audit index in Splunk and create dashboards
to visualize events by type, severity and namespace.Elasticsearch
kubectl Commands
Common audit queries via kubectl
Common audit queries via kubectl
Event Retention
The operator never deletes AuditEvents — there is no TTL, retention setting
or owner reference, so they are not garbage-collected with their Issue.
Every event is an object in etcd; configure a retention job to keep the
count bounded.
Server-Side Audit Trail
The AuditEvents above cover the operator. The ChatCLI server (chatcli server, the pod an Instance runs) has a separate, file-based trail: set spec.server.security.auditLogPath on the Instance (it becomes CHATCLI_AUDIT_LOG_PATH) to an absolute path, and every gRPC call is appended as a hash-chained JSON line (kind: "grpc": action, actor, role, caller IP, result, duration), interleaved with the LLM request entries in the same verifiable chain. Details and verification (/config security verify-audit) are in Security.
The Instance pod has a read-only root filesystem; only /tmp and /home/chatcli/.chatcli are writable, both emptyDir volumes lost on restart. To keep the file across restarts, enable spec.persistence and point the path into the sessions volume, e.g. /home/chatcli/.chatcli/sessions/audit.jsonl.
Next Steps
Approval Workflow
How plans are parked and decided — the source of the
approval_*
events, including gates raised by the decision engine and the cluster tier.SLOs and SLAs
Burn-rate alerts and SLA timers behind
slo_violation and sla_breach,
and the real per-severity SLA compliance.Web Dashboard
The audit view, API keys and REST roles.
AIOps Platform
Return to the AIOps platform overview.