Skip to content

Principles of security chaos engineering for AI systems

Chaos engineering is the practice of running experiments on a system to build confidence in its ability to withstand turbulent conditions (Principles of Chaos Engineering). Security chaos engineering applies the same method to security: instead of assuming that controls work, you inject the conditions under which they might fail and observe what happens.

AI agents make this more urgent. They are non-deterministic, they read untrusted content, they act through tools with real privileges, and their security often depends on runtime controls: guardrails, classifiers, allow-lists, approval steps. Those controls are just more distributed-system components, and they can time out, crash, be misconfigured or drift.

Agentic Chaos is built on the following principles.

1. Define the steady state as security invariants

Classic chaos engineering measures throughput and error rates. For agents, the steady state is a set of invariants that must hold no matter what:

  • untrusted content never causes a privileged action (canary_not_leaked, tool_not_called)
  • when a control is unavailable, sensitive actions stop (fails_closed)
  • failures degrade the answer, not the system (no_unhandled_error, max_tool_calls)

Answer quality matters, but it is a different question. Keep invariants binary and observable.

2. Hypothesise that the steady state holds, then try to disprove it

Every experiment states a hypothesis in plain language ("when the guardrail times out, the agent refuses") and runs a baseline first. If the baseline already violates the steady state, the result is inconclusive: fix the baseline before blaming the fault.

3. Inject both accidents and adversaries

Real-world turbulence for AI systems includes:

  • Accidents: provider outages and rate limits, slow or failing tools, truncated or malformed output, context-window exhaustion.
  • Adversaries: indirect prompt injection, poisoned tool metadata, poisoned memory and RAG data, spoofed inter-agent messages.
  • Control failures: guardrails that time out, error, or silently allow everything.

The most interesting findings are usually at the intersection, such as an attack arriving while a control is degraded.

4. Fail closed is a property you prove, not a property you declare

"The guardrail fails closed" is a claim about error-handling code paths that rarely run. Exercise those paths on purpose: control_outage and force_verdict exist for exactly this.

5. Verify detection, not just prevention

A blocked attack that nobody sees is a missed signal. A successful attack that nobody sees is an incident you learn about later. Use detection probes (for example alert_raised) to check that disruptions show up in logs, alerts and traces.

6. Measure distributions, not anecdotes

Models are stochastic. One passing run proves little. Run experiments many times (runs) and track pass rates over time, especially across model upgrades, prompt changes and new tools.

7. Prove that the fault fired

An experiment in which no fault was triggered proves nothing, and silently passes in most tools. Agentic Chaos reports such runs as inconclusive so mis-targeted experiments cannot pass by accident.

8. Use canaries, not harm

Adversarial payloads are benign and tagged with unique canary tokens. Success means a canary turned up where it should not: a tool argument or an output. No real exfiltration, no real exploit code, no real destinations.

9. Minimise the blast radius, then widen it

Start in unit tests and CI against stubs, move to staging with real models, and only then consider guarded production experiments with a kill switch, narrow targeting and low probability. See safety.md.

10. Degrade security gracefully, and test both sides

When a control fails, "fail open" gives up security and "fail closed" gives up availability. An attacker who can disrupt the control (for example with a guardrail DoS) wins either way. Prefer a third option, a safe mode such as read-only answers or no side-effecting tools, and verify it with security probes and availability probes in the same experiment (see statistics.md).

11. Automate continuously

Agent behaviour changes without a code change: the provider updates a model, someone edits a prompt, a new MCP server is connected. Run the catalog in CI and on a schedule so regressions are caught when they are introduced.