Guardian Autonomy Assurance Center

AI agent incident response and recovery

What to prepare before an agent does something you did not authorize — how to contain it, what you can and cannot undo, and the evidence we require before autonomy resumes.

“Autonomy must be earned, bounded, and revocable.”

Proposed standard · v0.1

Every scenario on this page is hypothetical. They are vendor-neutral failure modes we consider credible for agents with real access, written to be rehearsed. None of them describes an actual incident, a named product, or a Guardian customer, and no incident reports, breach claims, or response-time figures are presented here.

Six credible failure scenarios

Rehearse the failures that would actually hurt. For each one, there are two separate questions: could the deployment prevent or detect it, and could you contain and recover from it once it happened. Teams routinely have an answer to the first and none to the second.

Agent sends an unauthorized external message

Prevention & detection

Can the gateway enforce approved recipients and workflow?

Containment & recovery

Can queued messages be halted, and who manages correction of messages already sent?

Related controls: GCAS-11 · GCAS-15 · GCAS-20

State crosses a user or tenant boundary

Prevention & detection

Do retrieval and session isolation tests fail closed?

Containment & recovery

Can affected sessions be isolated and exposure scope established?

Related controls: GCAS-13 · GCAS-14 · GCAS-16

Agent deletes or corrupts records

Prevention & detection

Are write scopes, batch limits, and approvals enforced?

Containment & recovery

Can writes stop and records be restored or reconciled?

Related controls: GCAS-10 · GCAS-18 · GCAS-23

Malicious tool output changes the task

Prevention & detection

Can untrusted content cause a prohibited tool action?

Containment & recovery

Can the connector be disabled and affected memory reviewed?

Related controls: GCAS-17 · GCAS-16 · GCAS-20

Delegation multiplies cost or disruption

Prevention & detection

Do aggregate budgets include retries and child agents?

Containment & recovery

Can the entire task tree be stopped and pending jobs canceled?

Related controls: GCAS-12 · GCAS-18 · GCAS-20

The monitoring or policy service fails

Prevention & detection

Does privileged execution stop when authorization is unavailable?

Containment & recovery

Can people regain control through an independent administrative path?

Related controls: GCAS-17 · GCAS-19 · GCAS-20

The same set as a planning table:

Hypothetical AI agent failure scenarios with prevention and recovery questions
ScenarioPrevention / detection questionContainment and recovery question
Agent sends an unauthorized external message Can the gateway enforce approved recipients and workflow? Can queued messages be halted, and who manages correction of messages already sent?
State crosses a user or tenant boundary Do retrieval and session isolation tests fail closed? Can affected sessions be isolated and exposure scope established?
Agent deletes or corrupts records Are write scopes, batch limits, and approvals enforced? Can writes stop and records be restored or reconciled?
Malicious tool output changes the task Can untrusted content cause a prohibited tool action? Can the connector be disabled and affected memory reviewed?
Delegation multiplies cost or disruption Do aggregate budgets include retries and child agents? Can the entire task tree be stopped and pending jobs canceled?
The monitoring or policy service fails Does privileged execution stop when authorization is unavailable? Can people regain control through an independent administrative path?

Response sequence

A workable order of operations. Treat it as a structure to adapt, not a rigid script.

  1. 01

    Detect

    An alert, a report, or an anomaly in monitored activity establishes that something outside the approved boundary may have happened.

  2. 02

    Triage

    Establish what system is involved, what authority it holds, and the maximum credible impact. Assign incident command.

  3. 03

    Revoke / contain

    Stop the authority: revoke credentials, block egress, disable connectors, halt queued and delegated work. Do not delay urgent containment to finish collecting logs.

  4. 04

    Preserve evidence

    Capture action records, authorization decisions, tool results, and affected state where feasible, without unnecessarily duplicating sensitive data.

  5. 05

    Assess impact and duties

    Determine what was read, changed, sent, or spent, and who decides on notification and external communication.

  6. 06

    Restore / reconcile

    Restore from protected backups, reconcile corrupted records, and operate manually where required. Accept that sent messages and disclosed information need consequence management, not restoration.

  7. 07

    Retest

    Repeat the relevant evaluations and control tests against the repaired configuration.

  8. 08

    Approve restart

    An accountable person decides whether autonomy resumes, on what scope, and with what added limits.

Sequencing overlaps. These stages overlap in practice. Containment frequently runs in parallel with evidence preservation, and impact assessment continues well after service is restored. Responders should not delay urgent containment merely to collect logs.

What a restore cannot fix

Backups address deleted and corrupted data. They do nothing about consequences that have already left your systems, and planning that treats recovery as purely technical will be caught out here.

Usually recoverable

  • Deleted records, from protected backups.
  • Corrupted records, through reconciliation against a source of truth.
  • Unauthorized configuration and permission changes.
  • Queued or scheduled work, if it can be cancelled before execution.

Requires consequence management

  • Messages already delivered to customers, partners, or regulators.
  • Information already disclosed outside its boundary.
  • Payments already settled.
  • External commitments a counterparty may treat as binding.
  • Content already published or indexed elsewhere.

This asymmetry is the argument for binding approval to the specific action and for restricting outbound disclosure: the controls that matter most are the ones guarding actions you cannot take back.

Recovery drill worksheet

Run this against one deployment and one scenario. Record what you measured, not what you hoped for.

Blank recovery drill worksheet for your team to complete
FieldWhat to recordYour result
System and owner Which deployment, and who is accountable for it To be completed by your team
Scenario exercised Which credible failure you rehearsed To be completed by your team
Maximum credible impact Worst plausible outcome if uncontained To be completed by your team
Stop mechanism The specific control used to remove authority To be completed by your team
Containment target Your organization's target time to remove authority To be completed by your team
Measured containment time What the drill actually achieved To be completed by your team
Recovery-time target Your target to restore verified operation To be completed by your team
Acceptable data-loss target How much loss the business accepts To be completed by your team
Measured recovery Observed restore or reconciliation result To be completed by your team
Residual effects Irreversible consequences remaining after recovery To be completed by your team
Corrective-action owner Who owns each follow-up item, with a date To be completed by your team

No targets are prefilled. Targets are organization-specific and depend on your risk appetite, contractual duties, and regulatory obligations. Guardian does not publish benchmark figures for these fields, and no results are prefilled.

Print this page for a blank worksheet, or use the Markdown checklist alongside it to record which controls the drill exercised.

Where agents affect physical equipment

Where an agent can affect physical equipment, software controls must sit alongside appropriate independent equipment safety controls — interlocks, emergency stops, and protective systems designed for that purpose. GCAS addresses software authority. It is not a functional-safety assessment and does not substitute for one.

If an agent can influence machinery, vehicles, building systems, or laboratory equipment, treat the software authority controls in this standard as one layer among several. The independent protective systems remain the control of record for personnel safety.

Common questions

How should organizations prepare for an agent incident?

Rehearse a specific credible scenario against a specific deployment, with named responders, and measure the result. GCAS-21 lists the scenarios we expect to be exercised, including control-plane failure. The first rehearsal should not be the actual incident.

Can a kill switch stop queued or delegated work?

Only if it was designed to reach them. Stopping generation ends the current response; it says nothing about scheduled jobs, queued messages, in-flight tool calls, or child agents already dispatched. GCAS-20 requires a control outside the agent’s reach covering all of those, and a timed drill that documents what remained irreversible.

Who decides that autonomy can resume?

An accountable person, on evidence. GCAS-24 requires that the failure path be removed, affected credentials rotated, compromised memory repaired or cleared, relevant evaluations repeated, and a restart decision recorded. Restoring service is not the same as restoring authority.

Should containment wait for evidence collection?

No. Preserve what you can while you contain, but do not delay removing authority in order to finish gathering logs. GCAS-22 asks for evidence handling that is compatible with urgent containment rather than in competition with it.

Rehearse it with us

An AI Incident Readiness engagement produces a response runbook, a tabletop exercise with your responders, and explicit restoration and restart criteria.

Request incident readiness Run the 24-control checklist