Every scenario on this page is hypothetical. They are vendor-neutral failure modes we consider credible for agents with real access, written to be rehearsed. None of them describes an actual incident, a named product, or a Guardian customer, and no incident reports, breach claims, or response-time figures are presented here.
Six credible failure scenarios
Rehearse the failures that would actually hurt. For each one, there are two separate questions: could the deployment prevent or detect it, and could you contain and recover from it once it happened. Teams routinely have an answer to the first and none to the second.
Agent sends an unauthorized external message
Can the gateway enforce approved recipients and workflow?
Can queued messages be halted, and who manages correction of messages already sent?
State crosses a user or tenant boundary
Do retrieval and session isolation tests fail closed?
Can affected sessions be isolated and exposure scope established?
Agent deletes or corrupts records
Are write scopes, batch limits, and approvals enforced?
Can writes stop and records be restored or reconciled?
Malicious tool output changes the task
Can untrusted content cause a prohibited tool action?
Can the connector be disabled and affected memory reviewed?
Delegation multiplies cost or disruption
Do aggregate budgets include retries and child agents?
Can the entire task tree be stopped and pending jobs canceled?
The monitoring or policy service fails
Does privileged execution stop when authorization is unavailable?
Can people regain control through an independent administrative path?
The same set as a planning table:
| Scenario | Prevention / detection question | Containment and recovery question |
|---|---|---|
| Agent sends an unauthorized external message | Can the gateway enforce approved recipients and workflow? | Can queued messages be halted, and who manages correction of messages already sent? |
| State crosses a user or tenant boundary | Do retrieval and session isolation tests fail closed? | Can affected sessions be isolated and exposure scope established? |
| Agent deletes or corrupts records | Are write scopes, batch limits, and approvals enforced? | Can writes stop and records be restored or reconciled? |
| Malicious tool output changes the task | Can untrusted content cause a prohibited tool action? | Can the connector be disabled and affected memory reviewed? |
| Delegation multiplies cost or disruption | Do aggregate budgets include retries and child agents? | Can the entire task tree be stopped and pending jobs canceled? |
| The monitoring or policy service fails | Does privileged execution stop when authorization is unavailable? | Can people regain control through an independent administrative path? |
Response sequence
A workable order of operations. Treat it as a structure to adapt, not a rigid script.
-
01
Detect
An alert, a report, or an anomaly in monitored activity establishes that something outside the approved boundary may have happened.
-
02
Triage
Establish what system is involved, what authority it holds, and the maximum credible impact. Assign incident command.
-
03
Revoke / contain
Stop the authority: revoke credentials, block egress, disable connectors, halt queued and delegated work. Do not delay urgent containment to finish collecting logs.
-
04
Preserve evidence
Capture action records, authorization decisions, tool results, and affected state where feasible, without unnecessarily duplicating sensitive data.
-
05
Assess impact and duties
Determine what was read, changed, sent, or spent, and who decides on notification and external communication.
-
06
Restore / reconcile
Restore from protected backups, reconcile corrupted records, and operate manually where required. Accept that sent messages and disclosed information need consequence management, not restoration.
-
07
Retest
Repeat the relevant evaluations and control tests against the repaired configuration.
-
08
Approve restart
An accountable person decides whether autonomy resumes, on what scope, and with what added limits.
Sequencing overlaps. These stages overlap in practice. Containment frequently runs in parallel with evidence preservation, and impact assessment continues well after service is restored. Responders should not delay urgent containment merely to collect logs.
What a restore cannot fix
Backups address deleted and corrupted data. They do nothing about consequences that have already left your systems, and planning that treats recovery as purely technical will be caught out here.
Usually recoverable
- Deleted records, from protected backups.
- Corrupted records, through reconciliation against a source of truth.
- Unauthorized configuration and permission changes.
- Queued or scheduled work, if it can be cancelled before execution.
Requires consequence management
- Messages already delivered to customers, partners, or regulators.
- Information already disclosed outside its boundary.
- Payments already settled.
- External commitments a counterparty may treat as binding.
- Content already published or indexed elsewhere.
This asymmetry is the argument for binding approval to the specific action and for restricting outbound disclosure: the controls that matter most are the ones guarding actions you cannot take back.
Recovery drill worksheet
Run this against one deployment and one scenario. Record what you measured, not what you hoped for.
| Field | What to record | Your result |
|---|---|---|
| System and owner | Which deployment, and who is accountable for it | To be completed by your team |
| Scenario exercised | Which credible failure you rehearsed | To be completed by your team |
| Maximum credible impact | Worst plausible outcome if uncontained | To be completed by your team |
| Stop mechanism | The specific control used to remove authority | To be completed by your team |
| Containment target | Your organization's target time to remove authority | To be completed by your team |
| Measured containment time | What the drill actually achieved | To be completed by your team |
| Recovery-time target | Your target to restore verified operation | To be completed by your team |
| Acceptable data-loss target | How much loss the business accepts | To be completed by your team |
| Measured recovery | Observed restore or reconciliation result | To be completed by your team |
| Residual effects | Irreversible consequences remaining after recovery | To be completed by your team |
| Corrective-action owner | Who owns each follow-up item, with a date | To be completed by your team |
No targets are prefilled. Targets are organization-specific and depend on your risk appetite, contractual duties, and regulatory obligations. Guardian does not publish benchmark figures for these fields, and no results are prefilled.
Print this page for a blank worksheet, or use the Markdown checklist alongside it to record which controls the drill exercised.
Where agents affect physical equipment
Where an agent can affect physical equipment, software controls must sit alongside appropriate independent equipment safety controls — interlocks, emergency stops, and protective systems designed for that purpose. GCAS addresses software authority. It is not a functional-safety assessment and does not substitute for one.
If an agent can influence machinery, vehicles, building systems, or laboratory equipment, treat the software authority controls in this standard as one layer among several. The independent protective systems remain the control of record for personnel safety.
Common questions
How should organizations prepare for an agent incident?
Rehearse a specific credible scenario against a specific deployment, with named responders, and measure the result. GCAS-21 lists the scenarios we expect to be exercised, including control-plane failure. The first rehearsal should not be the actual incident.
Can a kill switch stop queued or delegated work?
Only if it was designed to reach them. Stopping generation ends the current response; it says nothing about scheduled jobs, queued messages, in-flight tool calls, or child agents already dispatched. GCAS-20 requires a control outside the agent’s reach covering all of those, and a timed drill that documents what remained irreversible.
Who decides that autonomy can resume?
An accountable person, on evidence. GCAS-24 requires that the failure path be removed, affected credentials rotated, compromised memory repaired or cleared, relevant evaluations repeated, and a restart decision recorded. Restoring service is not the same as restoring authority.
Should containment wait for evidence collection?
No. Preserve what you can while you contain, but do not delay removing authority in order to finish gathering logs. GCAS-22 asks for evidence handling that is compatible with urgent containment rather than in competition with it.
Rehearse it with us
An AI Incident Readiness engagement produces a response runbook, a tabletop exercise with your responders, and explicit restoration and restart criteria.
