Guardian Autonomy Assurance Center

AI agent security checklist: 24 controls

A working assessment for one AI deployment. Read the control, decide whether you can evidence it, and record where you stand. Ungated — no form, no email required.

“Autonomy must be earned, bounded, and revocable.”

Proposed standard · v0.1 Derived from the Controlled Autonomy Standard.

Download checklist Share on X

“Share on X” opens a prepared post; you finish and send it yourself. Nothing is posted automatically and no social scripts run on this page.

How to use this

Pick one deployment — a named use case, a specific model version, a specific permission set. Work through the 24 controls and mark each one. A control you cannot evidence is unknown, not a pass, and reproducible failures in unauthorized access, tenant separation, high-impact approval, or the ability to stop authority should block promotion outright.

Your selections stay in this page only. They are held in memory for the current page session, are never transmitted anywhere, and are cleared when you reload or close the tab. The export is generated locally in your browser.

01 · GCAS-01 – GCAS-04

Evaluate models in the system where they will operate

Test the actual model, tool, and data configuration for the intended use case. A published benchmark score describes a model in isolation; it cannot tell you whether your deployment will misuse a connected tool or retrieve the wrong customer's record.

  1. GCAS-01 Assign an accountable owner.

    Record the use case, business owner, technical owner, affected people, data classes, tools, and maximum authority before deployment.

    Evidence A dated system record and approved risk assessment.

    Your assessment of GCAS-01
  2. GCAS-02 Test the actual deployment.

    Evaluate the model together with its prompts, retrieval, memory, connectors, and permissions against realistic tasks and adversarial inputs. Include sensitive-data disclosure, prompt injection, incorrect actions, and harmful decisions relevant to affected people.

    Evidence A versioned test suite with outcomes and unresolved failures.

    Your assessment of GCAS-02
  3. GCAS-03 Make approval specific and temporary.

    Define measurable acceptance thresholds and a reassessment date for a named use case, model version, and permission set. Retest material changes to models, prompts, tools, memory, or access; suspend affected authority if new behavior exceeds the approved boundary.

    Evidence A release decision, recorded residual risk, expiry, and change history.

    Your assessment of GCAS-03
  4. GCAS-04 Examine provider dependencies.

    Review applicable data handling, retention, training use, subprocessors, hosting locations, incident notification terms, and service continuity. Establish what the available assurance evidence actually covers.

    Evidence A supplier review and a documented fallback path.

    Your assessment of GCAS-04
02 · GCAS-05 – GCAS-08

Secure local and private models

Verify provenance, infrastructure hardening, data flows, performance, and fallback. Running inference locally is an architecture choice, not a privacy guarantee.

  1. GCAS-05 Control model provenance.

    Record the origin, license, version, and integrity of model weights, images, libraries, and serving components. Review executable model-loading code before use.

    Evidence An artifact inventory with hashes and approval records.

    Your assessment of GCAS-05
  2. GCAS-06 Harden the inference environment.

    Authenticate inference requests; restrict network reachability, service privileges, admin access, and resource consumption. Maintain patching and vulnerability review.

    Evidence Deployment configuration, access tests, and a maintenance owner.

    Your assessment of GCAS-06
  3. GCAS-07 Verify the data boundary.

    Document and test outbound telemetry, logs, embeddings, backups, remote dependencies, and cloud fallback. Local inference alone does not establish that all data stays local.

    Evidence A data-flow diagram and observed network behavior under normal and failure conditions.

    Your assessment of GCAS-07
  4. GCAS-08 Maintain a qualified fallback.

    Benchmark the selected local configuration for its authorized tasks and rehearse capacity exhaustion or model unavailability. Any fallback must have equal or narrower approved data access.

    Evidence Evaluation results, capacity limits, and a fallback exercise.

    Your assessment of GCAS-08
03 · GCAS-09 – GCAS-12

Give every agent limited, attributable authority

Inspect identities, tool authority, delegated access, and high-impact approvals. An agent that inherits a person's unrestricted account has no meaningful boundary and no attribution.

  1. GCAS-09 Use a distinct agent identity.

    Attribute access to a named agent and responsible owner. Use scoped service identities rather than a person's unrestricted account or shared administrator credentials.

    Evidence Identity inventory and attributable access records.

    Your assessment of GCAS-09
  2. GCAS-10 Issue minimum necessary permissions.

    Default to no access; separately authorize tools, actions, resources, and duration. Use short-lived credentials where supported and broker secrets outside model context.

    Evidence Permission policy and denied-access tests.

    Your assessment of GCAS-10
  3. GCAS-11 Bind approval to the action.

    Require independent human approval for defined high-impact actions such as payments, bulk deletion, permission changes, and consequential external commitments. Bind approval to recipient, resource, parameters, time limit, and scope; changed actions need new approval.

    Evidence An approval record linked to the executed action and rejection of altered or replayed approvals.

    Your assessment of GCAS-11
  4. GCAS-12 Constrain delegation.

    Child agents and connected tools must receive no more authority than the authorized parent task permits. Restrict delegated credentials, recursion, tool registration, and the ability to spawn more agents.

    Evidence An agent/tool dependency map and tests that delegation cannot expand privilege.

    Your assessment of GCAS-12
04 · GCAS-13 – GCAS-16

Restrict data access and prove isolation

Restrict retrieval, outbound data, memory, and tenant or session state. Enforce permissions before data reaches the model, and demonstrate separation with negative tests rather than assuming it.

  1. GCAS-13 Limit retrieval at the source.

    Enforce row, document, tenant, and purpose restrictions before data reaches the model. Start with synthetic or minimized data and explicitly authorize broader use.

    Evidence Retrieval authorization tests across users and data classes.

    Your assessment of GCAS-13
  2. GCAS-14 Test tenant and session separation.

    Isolate identities, workspaces, browser state, files, caches, queues, vector indexes, and persistent memory. Test both concurrent use and reuse after cleanup. Shared infrastructure requires demonstrated boundaries; dedicated infrastructure still requires controls.

    Evidence Cross-tenant and cross-session negative tests with synthetic markers.

    Your assessment of GCAS-14
  3. GCAS-15 Restrict outbound disclosure.

    Apply destination allowlists, export limits, and appropriate content checks to tools, messages, uploads, and network traffic. Prevent an agent from turning a new connector into an unauthorized exit path.

    Evidence Blocked-export tests and access-controlled egress records.

    Your assessment of GCAS-15
  4. GCAS-16 Govern memory and retention.

    Define purpose, access, retention, and deletion for prompts, outputs, logs, and persistent memory. Keep secrets out of model context where practical; verify cleanup on session termination and tenant teardown.

    Evidence Retention configuration and deletion tests, including documented backup limitations.

    Your assessment of GCAS-16
05 · GCAS-17 – GCAS-20

Enforce limits outside the model

Enforce independent authorization, budgets, monitoring, and revocation. A model cannot be trusted to police itself once its context contains untrusted content.

  1. GCAS-17 Put authorization at the execution boundary.

    Use independent enforcement to check each tool action against identity, scope, resource, approval, and current policy. Untrusted content cannot authorize action. Deny new privileged actions when authorization cannot be verified.

    Evidence Bypass, stale-policy, and enforcement-outage tests.

    Your assessment of GCAS-17
  2. GCAS-18 Cap the damage a task can cause.

    Set per-task and aggregate ceilings for spending, runtime, tool calls, concurrency, data exports, and records changed. Enforce limits across retries and child agents.

    Evidence Boundary tests showing the enforced stop condition.

    Your assessment of GCAS-18
  3. GCAS-19 Observe consequential activity.

    Record action requests, authorization decisions, approvals, tool results, state changes, and agent identity in protected logs. Minimize sensitive content. AI may assist detection and investigation, but cannot independently grant broader access or suppress oversight.

    Evidence Traceable action records and tested alerts with a human response owner.

    Your assessment of GCAS-19
  4. GCAS-20 Stop authority independently.

    Provide a control outside the agent's reach to revoke credentials, block egress, disable connectors, and stop scheduled, queued, in-flight where feasible, and delegated work. Stopping generation alone is insufficient.

    Evidence A timed containment drill documenting remaining irreversible effects.

    Your assessment of GCAS-20
06 · GCAS-21 – GCAS-24

Prepare for the worst credible failure

Rehearse cascading failures, preserve evidence, restore operations, and require an explicit restart decision. Some consequences — a message already sent, information already disclosed — cannot be restored from a backup.

  1. GCAS-21 Rehearse realistic incidents.

    Exercise unauthorized messages, bulk deletion, exfiltration, cross-tenant exposure, malicious tools, model/provider compromise, runaway spending, and cascading agent failures. Include control-plane failure and physical consequences where agents affect equipment.

    Evidence A tabletop or sandbox drill with named responders and recorded decisions.

    Your assessment of GCAS-21
  2. GCAS-22 Preserve evidence and coordinate response.

    Assign incident command, severity criteria, evidence handling, escalation, and notification decision owners. Contain the incident while preserving useful evidence where feasible; avoid duplicating sensitive data unnecessarily.

    Evidence An incident runbook and tested communication paths.

    Your assessment of GCAS-22
  3. GCAS-23 Restore verified business operations.

    Maintain protected backups, recovery targets, reconciliation procedures, and a manual operating mode. Measure recovery time and data loss during exercises. External disclosures and messages require consequence management; a restore cannot undo them.

    Evidence A successful restore or reconciliation drill and a verified fallback procedure.

    Your assessment of GCAS-23
  4. GCAS-24 Require evidence before restart.

    Remove the failure path, rotate affected credentials, repair or clear compromised memory, repeat relevant evaluations, and obtain accountable human approval before restoring autonomy.

    Evidence A remediation record, regression results, and a restart decision.

    Your assessment of GCAS-24

Want help closing the gaps?

Guardian runs deployment risk reviews, model evaluations, agent security assessments, containment engineering, and incident readiness exercises. Tell us what the system does and which controls you could not evidence.

Request an agent security assessment Read the full standard

Guardian Controlled Autonomy Standard — Proposed standard, v0.1. Published by Guardian Robotics, . Control identifiers are stable across revisions.