Federal AI compliance explained

Short answer

Deploying AI for a federal agency means fitting it inside an existing authorization boundary. The practical constraint is that a commercial model API outside your FedRAMP boundary is usually not usable for controlled data, which pushes designs toward FedRAMP-authorized services, GovCloud deployments, or self-hosted open-weight models. Decide the boundary before the architecture, because it determines every subsequent choice.

4 min readUpdated 2026-09-28Federal & Public Sector AI

Federal AI work is ordinary AI engineering inside an unusual constraint: every component must sit within an authorized boundary, and the boundary is decided before anything else.

Start with the authorization boundary

An agency system runs under an Authority to Operate granted against a defined boundary and a control baseline. Adding AI means either extending that boundary or keeping the AI inside it.

This single decision determines your architecture. A commercial model API sitting outside the boundary generally cannot process controlled data, no matter how good its commercial security posture is. The realistic options are:

  1. A FedRAMP-authorized AI service. Cleanest path where one exists at the required impact level and offers the capability you need.
  2. Deployment in a government-region cloud within the agency's existing boundary.
  3. Self-hosted open-weight models on agency infrastructure — maximum control, maximum operational burden.
  4. Keep controlled data out entirely — viable for some use cases and worth checking before assuming you need more.

Confirm authorization status and impact level against the official FedRAMP marketplace rather than vendor claims, since status changes.

The frameworks in play

FedRAMP authorizes cloud services for federal use at Low, Moderate, or High impact. Agencies can generally only use authorized services for in-scope data. Check the actual authorization scope — a vendor may have one authorized offering and several that are not.

NIST 800-53 is the control catalogue underlying federal authorizations. AI systems touch familiar control families in slightly unfamiliar ways: access control over retrieval, audit logging of model inputs and outputs, configuration management of model versions, and supply-chain risk for model provenance.

NIST AI RMF is voluntary but increasingly referenced in federal solicitations and agency policy. Treat it as expected for anything sold into government. See AI governance frameworks.

CMMC applies to defence contractors handling controlled unclassified information. If your AI pipeline touches CUI, it falls inside your CMMC scope — including, notably, the logs and traces your pipeline generates.

FISMA sets the overall framework agencies operate under.

The most commonly missed item: your observability stack is in scope. Traces and logs containing prompts and outputs hold the same data classification as the source, and they frequently end up in a tool nobody assessed. This finding comes up repeatedly.

AI-specific control considerations

Access control at retrieval. A RAG system that searches a whole corpus regardless of the requesting user's clearance is an access-control failure, even though no conventional permission was bypassed. This must be enforced at retrieval, and revocation must propagate.

Audit completeness. Logging that the model was called is insufficient. You need inputs, retrieved context, outputs, model version, and the decision taken — sufficient to reconstruct why the system produced what it produced.

Model version control. A model update changes system behaviour. Under configuration management that is a change requiring assessment, not a silent dependency bump. Provider-side model deprecation can force this on your timeline rather than yours.

Supply chain and provenance. Where did model weights come from, what is in the training data, what is the update path. Increasingly asked, rarely answerable in detail for closed models.

Human oversight on consequential decisions. Expect to document where a human decides and what evidence they saw.

Practical sequencing

  1. Classify the data the system will touch.
  2. Identify the authorization boundary and impact level.
  3. Choose components that fit inside it — this eliminates most options immediately, which is useful.
  4. Design retrieval access control and audit logging as first-class requirements, not additions.
  5. Build the evaluation set; it doubles as accuracy evidence for assessors.
  6. Document human oversight points.
  7. Engage the assessor early. Late-stage surprises are expensive and schedule-fatal.

Frequently asked questions

Can federal agencies use commercial AI APIs?

For public or non-sensitive data, often yes. For controlled data, generally only where the service is FedRAMP-authorized at the appropriate impact level and the agency has accepted it into their boundary. Verify current authorization status directly, and do not rely on a vendor's marketing description of their compliance posture.

Does FedRAMP cover AI specifically?

FedRAMP authorizes cloud services generally rather than AI as a category. An AI service can be FedRAMP-authorized like any other cloud offering. What AI adds is scrutiny of data flows, model provenance, and output auditability within that existing process.

What about CUI in an AI pipeline?

If the pipeline processes CUI, the whole pipeline is in scope — including vector stores, caches, logs, and traces. The derived artifacts are the part organizations most often overlook, and embeddings of CUI should be treated as CUI.

How long does federal AI authorization take?

Adding an AI capability inside an existing authorized boundary using already-authorized components is the fast path and can be a matter of months. Anything requiring new authorization of a service is substantially longer. The single biggest schedule driver is whether you can reuse existing authorizations.

Is self-hosting open-weight models a good option for agencies?

It gives maximum data control and avoids external dependency, which solves the boundary problem cleanly. The cost is real operational burden — serving infrastructure, scaling, security patching, and model upgrades all become yours. It makes most sense where data sensitivity rules out alternatives or where a narrow fine-tuned small model suffices.

Guardian Robotics is an AI consultancy.

We build the pipelines, agents, and automation this article describes — for commercial teams and federal agencies alike.