Guardian Autonomy Assurance Center

AI governance for systems that can act.

Our proposed requirements for evaluating frontier models, deploying private AI, limiting agent authority, and recovering when systems fail.

“Autonomy must be earned, bounded, and revocable.”

Get the 24-control checklist — ungated, no form.

Proposed standard · v0.1

AI can act on your business. Its authority must have limits. Guardian's proposed standard defines the evidence we require before models and agents receive access, take action, or return to service after an incident.

What this section is

This is Guardian Robotics’ own proposed policy for controlled autonomy — published so that the organizations we work with can read our expectations, challenge them, and use them to ask better questions of their own deployments.

It states requirements and the evidence we expect against each one. It does not assert that Guardian or any other organization has implemented or independently verified them, and it is not a certification, an audit report, or a statement of legal compliance.

Six domains

Each domain has a policy and a corresponding artifact.

A requirement without a named piece of evidence is an opinion. For every domain below, the standard specifies what we expect to see.

01

Frontier-model evaluation

Test the actual model, tool, and data configuration for the intended use case. A published benchmark score describes a model in isolation; it cannot tell you whether your deployment will misuse a connected tool or retrieve the wrong customer's record.

Evidence artifact
Deployment evaluation and acceptance decision
02

Local and private AI

Verify provenance, infrastructure hardening, data flows, performance, and fallback. Running inference locally is an architecture choice, not a privacy guarantee.

Evidence artifact
Private deployment boundary review
03

Agent security assessment

Inspect identities, tool authority, delegated access, and high-impact approvals. An agent that inherits a person's unrestricted account has no meaningful boundary and no attribution.

Evidence artifact
Permission and tool-control matrix
04

Data access and isolation

Restrict retrieval, outbound data, memory, and tenant or session state. Enforce permissions before data reaches the model, and demonstrate separation with negative tests rather than assuming it.

Evidence artifact
Access and isolation test results
05

Containment

Enforce independent authorization, budgets, monitoring, and revocation. A model cannot be trusted to police itself once its context contains untrusted content.

Evidence artifact
Measured containment exercise
06

Worst-case readiness

Rehearse cascading failures, preserve evidence, restore operations, and require an explicit restart decision. Some consequences — a message already sent, information already disclosed — cannot be restored from a backup.

Evidence artifact
Incident runbook and recovery exercise
Method

How we use AI to govern AI

Our proposed policy permits AI-assisted test generation, anomaly detection, evidence summarization, and incident triage — under the same access and data controls that apply to any other system.

Detection models can be wrong, and they can be compromised. Independent authorization controls enforce limits; accountable people approve authority changes and incident restart. The monitor is another system to evaluate, not an infallible judge.

  • AI may assist test generation, anomaly detection, evidence summarization, and incident triage.
  • Independent policy enforces authority — not the model being governed.
  • People own risk acceptance, authority changes, and the decision to restart after an incident.
  • The monitor is in scope. A detection model is another system to evaluate, not an infallible judge.
Enterprise services

Engagements with a defined output.

Each option below is a scoped engagement. Selecting one preselects the topic on our contact form; nothing is booked automatically.

AI Deployment Risk Review

A structured review of what AI you are actually running, what authority it holds, and what to fix first.

Output Model and agent inventory, threat assessment, prioritized remediation plan.

Discuss this engagement

Frontier Model Evaluation

Task and adversarial evaluation of the deployment you intend to ship, not the model in isolation.

Output Task and adversarial evaluation suite, results, deployment recommendation.

Discuss this engagement

Private AI Deployment

Design and acceptance testing for privately hosted models, including the data paths people forget.

Output Architecture, data-flow controls, hardening plan, acceptance testing.

Discuss this engagement

Agent Security Assessment

Identity, permission, connector, and delegation review of agents already in your environment.

Output Identity and access review, connector tests, delegation and approval findings.

Discuss this engagement

Agent Containment Engineering

Building the enforcement layer outside the model, then proving it stops what it claims to stop.

Output Authorization design, runtime limits, revocation and containment tests.

Discuss this engagement

AI Incident Readiness

Runbook, rehearsal, and restart criteria for the failures you would rather not meet unprepared.

Output Response runbook, tabletop exercise, restoration and restart criteria.

Discuss this engagement

Tell us what the system does.

Describe the deployment, the environment it runs in, and what you want to improve. We will tell you which controls we would look at first.