Paved roads, not roadblocks

An AI governance operating model

Most AI governance arrives as a policy nobody can act on. This one starts at the other end — the operating model that runs every day, the people accountable for it, and one register everything else hangs off. Govern AI like any other critical capability: proportionate, risk-based, automated where it can be, accountable always.

01The operating modelFive moves · everything below sits under one of them

Each move is a section below. Select one to jump to it.

Tools
ServiceNow AI Control TowerIBM watsonx.governanceCredo AIOneTrustHolistic AIVantaDrata
Accountable throughoutCross-cutting — applies to all five moves
0.1Who owns whatLane owners & the council

A model with no owner is an incident with a delay. Name the owner before the use case ships.

 Board / Risk Committee
Executive AI Risk Council
CISOSecurity
CIO / CTOTechnology
LegalInterpretation
PrivacyData
BusinessProduct owners
AI Security
Lane owners — one named owner per AI population in the inventory
AEnterprise AICopilots, assistants, productivity
BProduct AICustomer-facing, RAG, models
CAgentic AITools, actions, workflows
DVendor AISaaS, APIs, managed services
EShadow AIBrowser tools, free SaaS

Who owns what

Business / product owner owns the AI risk.
Security reduces it.
Legal interprets it.
Privacy governs the data.
Engineering builds it.
GRC provides governance and evidence.
Tools
Entra ID groupsOktaServiceNow assignment groupsConfluence RACI
1DiscoverFind it, register it, name an owner
1.1The AI inventory and its registerFive populations · A–E

You cannot govern what you don't know — and what you know needs different guardrails depending on how it reaches a person. What's in each population, and what you apply to it, in one view. Lettered A–E, and the same five carry through accountability (02) and tooling.

AEnterprise AI

  • Copilots (M365, Google)
  • ChatGPT / Claude
  • Coding assistants
  • HR / Finance AI
  • Productivity tools
Controls
  • Identity & access
  • Acceptable use
  • Data classification
  • Retention
  • Connectors / plugins
  • DLP & logging
  • Monitoring

BProduct AI

  • Product LLMs
  • RAG / vector DB
  • ML models
  • AI features
  • Customer facing AI
Controls
  • AI gateway
  • Input validation
  • Prompt protection
  • RAG security
  • Model management
  • Output controls
  • Observability

CAgentic AI

  • AI agents
  • Tool calling
  • API actions
  • Autonomous workflows
Controls
  • Agent identity
  • Tools & actions
  • Permissions
  • Human approval
  • Least privilege
  • Kill switch
  • Audit / logging

DVendor AI

  • SaaS with AI
  • Vendor copilots
  • APIs / models
  • Managed AI services
Controls
  • Model & data
  • Training / retention
  • Subprocessors
  • Security testing
  • AI change management
  • Incident notification
  • Contractual controls

EShadow AI

  • Browser tools
  • Free SaaS
  • Unsanctioned LLMs
  • Unapproved uploads
Controls
  • Discover
  • Classify risk
  • Coach / approve
  • Block if needed
  • Provide alternatives
  • Monitor usage
AI Register — the system of record
OwnerBusiness useUsersModel DataVendorRAG?Agent? Tools / actionsExternal exposureDecision impact Risk tierControlsApprovalMonitoring
Same principles across all five. Different controls, because the way the AI reaches a person is different.
Discovery tools
Wiz AI-SPMOrcaPrisma Cloud AI-SPMDefender for Cloud AppsNetskope OneNudge SecuritySailPointJFrog AI CatalogMLflow
What runs the register

The register is the one thing that must exist before anything else works. Three honest paths — start where you are, not where you want to be.

Under ~50 AI systems

Spreadsheet

SharePoint listExcelAirtable Notion DBGoogle Sheets
Live in a week. No workflow, no evidence trail, no audit history.
~50–300 · approvals needed

Your ITSM / GRC

ServiceNow CMDB class + IRMJira + Confluence Power Apps + DataverseArcherLogicGate Vanta / Drata custom objects
Reuses your approval engine and CMDB. Configuration effort, no AI-specific content.
Regulated · auditor coming

AI governance platform

ServiceNow AI Control TowerIBM watsonx.governance Credo AIOneTrust AI GovernanceHolistic AI ModulosCollibraDataiku Govern Databricks Unity Catalog
Pre-built EU AI Act / ISO 42001 / NIST evidence. Cost, and a new platform to run.
Choose byVolume — under 50, a list beats a platform.
Choose byWorkflow — if approvals must route, use ITSM.
Choose byEvidence — if an auditor is dated, buy the platform.
What fills it automatically

A hand-typed register is stale in a month. Wire the feeds first, the fields second.

CASB / SSE
usage
AI-SPM
cloud models
IdP app
registrations
Expense &
procurement
Model
registry
Repos &
artifacts
AI gateway
logs
Intake
form

Non-negotiable fields

One named human owner
Stable unique ID
Risk tier + date set
Link to CMDB / vendor record
Evidence attachments
Change history
If it has no owner, it isn't registered — it's just listed.
1.2Shadow AIWhat discovery turns up that nobody declared

People reach for shadow AI because the approved path is slower. Fix the path, not just the policy.

Signals

  • SSO / identity
  • CASB / SSE
  • Browser / DNS
  • Expense / cards
  • Cloud accounts
  • Git repos / code
  • Network / proxy

Discovery & classification

  • Identify tools
  • Assess risk
  • Map data flow
  • User impact

Decision

AllowLow risk
CoachMove to an approved tool
BlockHigh risk

Outcomes

  • Approved tools list
  • Better alternatives
  • Guidance & training
  • Usage monitoring
  • Periodic review
Enable the right AI. Give employees approved, secure alternatives — that is the control that actually works.
Tools
Defender for Cloud AppsNetskopeZscalerCloudflareNudge SecurityCisco AI GuardrailsJFrog Shadow AI DetectionIslandChrome Enterprise PremiumRamp / Brex / Coupa
2DecideTier the risk, approve or refuse
2.1The risk engineSeven factors in, one tier out

Seven factors in, one tier out. Meeting summaries and autonomous clinical decisions should never meet the same committee.

Data sensitivity
Customer exposure
Regulatory impact
Autonomous actions
Privileged access
Decision consequence
External model / vendor
AI
Risk engine
LowAuto-approve
MediumReview
High / CriticalAssurance + approval
Tools
ServiceNow IRMOneTrustCredo AI policy packsArcherLogicGatePower AppsNIST AI RMF profile
2.2Risk tiersWhat each tier actually demands

The tier determines the controls. Not just a colour on a dashboard.

TierExampleGovernance approachKey requirements
Tier 1 — Low
Meeting summaries, basic Q&A, non-sensitive data Register + baseline controls Acceptable use, data classification, logging, DLP, training
Tier 2 — Moderate
Internal RAG over company data, document analysis Security & privacy review + testing Access controls, data protection, prompt/RAG security, vendor review
Tier 3 — High
Customer chatbot, coding agent, HR AI Threat model + AI red team + monitoring + approval Comprehensive security testing, continuous monitoring, human oversight
Tier 4 — Critical
Medical or financial decisions, autonomous agents Executive & risk approval + continuous assurance Formal risk acceptance, continuous testing, human-in-the-loop, incident response
Tools
Register field + ITSM rulesOPA policy-as-codeDataiku Govern sign-offs
2.3The approval workflowIntake, tiering, decision

If asking permission is harder than going around it, people go around it.

I want to use AI
AI intake portal8–10 useful questions
Automated tiering
LowAuto-approve
MediumSecurity review
HighAI assurance review
AI register+ controls + monitoring
Simple intake. Fast decisions. Transparent outcomes.
Tools
ServiceNowJira Service ManagementFreshserviceMicrosoft FormsPower AutomateCredo AI
3ControlApply the guardrails the tier demands
Lane controls A–E live with the inventory in 1.1 — the three below are the ones that need their own design.
3.1Agentic control planeNever let the model authorize itself

Agents can do things — so they need guardrails, not just guidance.

User / application
AI agentTools · APIs · systems

Policy enforcement point

  • Is the agent authorized?
  • Is the user authorized?
  • Is the action allowed?
  • Is the amount / destination allowed?
  • Is human approval required?
  • Context and risk checks
Execute action
Never let the LLM decide authorization. Agents get their own identities and least-privilege permissions, the same as any other principal.
Tools
Entra Agent IDOktaSPIFFE/SPIREOPACedarPermit.ioHashiCorp VaultCloudflare MCP PortalsE2B sandboxTemporalLangfuseAgentOps
3.2Vendor AI and the AI-BOMThe chain behind the contract

Your contract is with one company. Your data reaches four. Map the chain, put the answers in writing, take delivery of an AI-BOM, then keep watching — because the vendor's model will change without asking you.

Map the chain
  • Who is the contracting party
  • Who is underneath them
  • Whose model is it really
  • Where does data land
Ask
  • Eight questions
  • Model, data, location
  • Training and retention
  • Agent capabilities
Require
  • Contract terms
  • DPA and BAA flow-down
  • AI-BOM at each release
  • Notification rights
Verify
  • Ingest and scan the AI-BOM
  • Certifications and reports
  • Subprocessor list
  • Test what you can
Monitor
  • Model or version change
  • New subprocessor
  • New agent actions
  • Re-assess or exit
Any material change sends the vendor back to Ask
The chain behind the contract
You
ControllerAccountable to the regulator and the patient or customer, whatever the contract says.
Tier 1
VendorThe SaaS you signed with.
DPA · BAA · AI-BOM · audit rights
Tier 2
SubprocessorsCloud, storage, analytics, support.
Named list · change notice · right to object
Tier 3
Model providerWhoever actually serves the inference.
No-training terms · region · retention
Tier 4
Data servicesLabelling, fine-tuning, RAG sources.
Human review disclosure · licence · consent
Most programmes stop at Tier 1. The questions that actually matter — is my data training a model, is a human reading it, which country is it in — are usually answered two or three tiers down, by a company you have no contract with. Ask the vendor to answer for the whole chain, and make the flow-down obligation contractual.
The AI-BOM

What it has to contain

ModelsName, version, provider, licence, weights hash
DatasetsTraining, fine-tune and RAG corpora, licence, consent basis
PromptsSystem instructions and their version
Tools & actionsWhat the agent can invoke, and against what
SoftwareOrdinary SBOM: libraries, frameworks, transitive deps
ServingInfrastructure, region, subprocessors
EvaluationsResults, known limitations, guardrails in place
AttestationSignature, build provenance, VEX for known CVEs

Formats

CycloneDX ML-BOMSPDX 3.0 AI profile SPDX dataset profileModel cardsOpenVEX / CSAF

Generate & sign

cdxgenSyftCycloneDX CLI Sigstore / cosignin-toto + SLSA

Ingest, scan & watch

Dependency-TrackManifest CyberFinite State JFrog CurationAnchoreReversingLabs HiddenLayerPrisma AIRS
CadenceAt onboarding, at every material release, and on demand after an incident. A one-off AI-BOM at procurement is a snapshot of a system that no longer exists.
Ask
What model?
What data?
Where processed?
Training on our data?
Retention & deletion?
Subprocessors?
Security testing?
Agent capabilities?

Require in the contract

No training on our data without authorization
Data retention & deletion requirements
Security incident notification
Material AI change notification
Audit / assurance rights
Data-location requirements
Exit & data deletion
AI-BOM at every material release
Named subprocessor list & right to object
Processor terms flow down the full chain
Disclosure of any human review
Model version pinning or notice before change
Tools
OneTrust TPRMWhisticPanoraysUpGuardIroncladCycloneDX ML-BOMcdxgenSyftDependency-TrackManifest CyberFinite StateJFrog CurationSubprocessor list monitoring
3.3The engineering lifecycleControls in the pipeline, not the policy

Governance that lives in a document gets skipped. Governance that lives in the pipeline gets run.

Plan / design
  • AI use case review
  • Threat modeling
  • Privacy assessment
  • Risk tiering
Code
  • Secure AI assistant
  • Secrets scanning
  • SAST / SCA / IaC
  • AI policy checks
Build / test
  • Model validation
  • Prompt tests
  • RAG tests
  • Agent tests
  • Red team tests
Deploy
  • Policy-as-code
  • Infrastructure
  • Security checks
  • Approval
Runtime
  • AI gateway
  • DLP
  • Abuse detection
  • Observability
  • SIEM / logging
Monitoring, telemetry, incidents and user feedback go back to plan
Tools
GitHub Advanced SecuritySemgrepSnykGitleaksCheckovpromptfooDeepEvalRagasGiskardInspectOPA GatekeeperKyvernoSigstore
4Monitor & RespondWatch, detect, contain, recover
4.1Runtime monitoring and responseTelemetry, detection, the AI SOC

A model that passed review in March is a different model in September. Watch it in production, and be able to reconstruct what happened.

Runtime monitoring

  • AI gateway logs
  • Agent actions
  • Model telemetry
  • IAM events
  • DLP events
  • User feedback
SIEM
/ AI SOC

Detect & respond

  • Prompt attacks
  • Data leakage
  • Abuse / anomalies
  • Excessive actions
  • Privilege escalation
  • Model behaviour drift
Tools
PyRITgarakpromptfoo redteamMindgardHiddenLayerCisco AI DefenseSplunkMicrosoft SentinelElasticMITRE ATLAS
Detection without a response path is an alert nobody can act on — see 4.2 Respond.
4.2RespondWhat to do when something goes wrong

Detection is not response. This part gets written before the afternoon an agent has already emailed four hundred customers the wrong price — or it does not get written at all.

DetectAlert, user report, vendor notice
ContainTurn it off at the gateway
PreserveFreeze the trace first
AssessBlast radius
NotifyWhoever has a clock
RecoverKnown-good version
LearnAn eval case, a detection
Containment before diagnosis. A non-deterministic system may never reproduce the exact failure. Stop the harm first and understand it second — the reverse order is how a two-hour incident becomes a two-day one.

Six ways it goes wrong

What went wrongWhat it looks likeFirst containment moveOwns the call
Harmful or wrong output at scaleFabricated fact, bad advice or offensive content reaching real usersDisable the route at the gateway and serve the fallbackProduct owner
Sensitive data exposureRegulated data in prompts, retrieved context, output or vendor logsQuarantine the retrieval index; suspend the connectorAI Security + Privacy
Prompt injection or tool abuseUntrusted content steering the model into calls nobody intendedQuarantine the content source; drop the tool from the manifestAI Security
Agent overreachActions beyond intent — spend, send, delete, commit, escalateRevoke the agent identity; freeze its entitlementsAI Security + Engineering
Upstream change or compromiseProvider silently changed the model, poisoned dependency, vendor incidentPin to the last known-good version; fail over to the fallback modelEngineering + vendor owner
Runaway loop or costRecursion, retry storms, denial of walletEnforce the spend cap; kill the sessionEngineering

Severity, and who wakes up

SeverityTriggerAcknowledgeIn the roomExternal clock
S1 — Critical
Real-world harm, regulated data confirmed out, or an agent committed a financial or safety-relevant action 15 min, 24×7 Incident commander, exec sponsor, Legal, Privacy, Comms Starts now
S2 — High
Exposure suspected but bounded, or wrong output to a limited, identifiable population 1 hour, on-call AI Security, product owner, engineering lead Legal on standby
S3 — Moderate
Policy breached, no confirmed exposure or harm Next business day AI Security, product owner Internal only
S4 — Near miss
Caught by an eval, a guardrail or a user before it reached anyone Backlog Owning team None — feeds Assure
Severity is about impact, not tier. A Tier 2 internal tool that leaks a payroll file is an S1.
SWITCHRoute disableAt the gateway, not the deployment
SWITCHIdentity revocationAgent credentials, in seconds
SWITCHIndex quarantineRetrieval source, isolated
SWITCHVersion rollbackModel, prompt and config pinned
These four have to exist before the incident, plus a hard spend cap. A kill switch you have to build during a response is not a kill switch — it is an outage with extra steps. Test each one in a tabletop, quarterly.

Preserve, before you fix

  • The prompt and system prompt exactly as sent
  • Retrieved context and the documents it came from
  • Every tool call, with arguments and result
  • Model name, version and hash at time of call
  • The identity that made the call, human or agent
  • Gateway request id, to stitch the chain together
A rollback destroys the evidence. Capture first, redeploy second.

Recover to known good

  • Restore the pinned model, prompt and config from the last signed release
  • Re-run the evaluation suite — a rollback is itself a change
  • Re-enable staged, behind a flag, internal users first
  • The route stays off until the product owner signs it back on
  • Watch the same signal that caught it, for a defined window
Reversal question, asked early: can the actions be undone at all — emails recalled, transactions reversed, commits reverted?

Who might have a clock

Affected users and customers Data protection authority Data subjects The model or platform provider Contractual notification terms Insurer Board or audit committee, at S1
Notification windows differ by jurisdiction, sector and contract. The point of this row is that somebody looked them up before the incident and wrote the numbers next to the names. Counsel makes the call, not the responder — and none of this is legal advice.
What makes an AI incident different

There is no stack trace. The failure lives in a prompt, a retrieved document and a model version, and the same input may not reproduce it tomorrow.

The fix is often a prompt change — which means the fix is a change nobody versioned, unless prompts were already release artefacts.

The vendor can change the model underneath you between the incident and the review, and take your ability to reproduce it with them.

The evidence is gone unless it was logged at the time. Nothing here is recoverable after the fact, which makes 4.1 the precondition for this whole page.

Every incident closes with something that stops it recurring

  • An evaluation case that reproduces the failure, added to the suite
  • A detection that would have caught it earlier
  • A guardrail, or a tier change, or a scope reduction on the agent
  • A named owner and a date, not an action item
Blameless on people, ruthless on the control that was missing. The output feeds 5 Assure; if nothing changed, the review did not happen.
Tools
PagerDutyincident.ioServiceNow SecOpsJiraSplunkMicrosoft SentinelLangfuseOpenTelemetry GenAIMITRE ATLASNIST SP 800-61
5AssureTest, evidence, learn, strengthen
5.1Pre-production assurance and provenanceWhat you prove before it ships, and how you know it hasn't changed

Assurance is a subscription, not a purchase. Prove it before it ships, then keep proving it.

Before production

  • Architecture review
  • AI threat model
  • Privacy assessment
  • Security testing
  • AI red team
  • Risk acceptance
Detecting change, proving provenance

Four things move under you: the data, where it came from, the model, and where it came from. Two need a signal. Two need a signed record.

Data change

Did the input distribution move?

Watch for
Schema driftVolume & freshnessDistribution shiftNull & range breaks
Tools
Great ExpectationsSoda Coredbt testsEvidently Monte CarloAnomaloBigeyeDatafoldDatabricks Lakehouse Monitoring
Evidence producedDated drift report and failed-check log, attached to the register entry.

Data provenance

Where did it come from, who touched it, what did it feed?

Capture
Source systemTransform chainConsent & licenceDownstream use
Tools
OpenLineage + MarquezDataHubApache Atlasdbt docsC2PA Content Credentials Unity Catalog lineageCollibraAlationAtlan
Evidence producedColumn-level lineage graph tracing source to model input.

Model change

Is it still the model that passed review?

Watch for
Version & weightsSystem prompt editsOutput driftEval regressionSilent vendor upgrade
Tools
MLflow Model RegistryEvidentlyNannyMLDeepcheckspromptfoo regression Weights & BiasesArizeFiddlerWhyLabs
Evidence producedGolden-set eval run per version, with a pass/fail gate before promotion.

Model provenance

Is this model what it claims to be?

Capture
Origin & licenceTraining data claimsFine-tune historySignature & hash
Tools
CycloneDX ML-BOMSPDX 3.0 AI profileSigstore / cosignin-toto + SLSAModel cards HiddenLayer Model ScannerPrisma AIRSJFrog ML
Evidence producedSigned AI-BOM plus a scan attestation, stored with the release.
You can't instrument a model you don't host. For vendor and API models, provenance is contractual and change detection is behavioural: pin the version, subscribe to the deprecation feed, and run a canary eval on a fixed golden set on a schedule. When the canary moves and you changed nothing, the vendor did.
5.2Where this mapsPractice → framework crosswalk

Build it once, evidence it many times.

PracticeNIST AI RMFISO/IEC 42001EU AI Act
01 Operating modelGOVERN 1Cl. 4–10 · A.2 · A.6Art. 17
02 AccountabilityGOVERN 2, 3A.3Art. 4 · 26
03 Inventory & lane controlsMAP 1, 4 · MANAGE 2A.4 · A.6 · A.7Art. 10 · 11 · 15 · 49
04 Risk engineMAP 5 · MEASURE 1A.5Art. 6 · Annex III
05 Risk tiersMANAGE 1A.5 · A.9Art. 9
06 Agentic control planeMANAGE 2.3 · MEASURE 2A.6.2 · A.9.3Art. 14 · 15
07 Vendor AIGOVERN 6 · MANAGE 3A.10Art. 25 · 26
08 Shadow AIGOVERN 1.6 · MAP 4.1A.4.2 · A.9.2Art. 4 · 26
09 Engineering lifecycleMEASURE 2 · MANAGE 4A.6Art. 15 · 17
10 Continuous assuranceMEASURE 3, 4A.6.2Art. 72 · 73
11 Approval workflowGOVERN 1.3 · MANAGE 1.1A.5.4 · A.9.2Art. 9 · 26
Indicative crosswalk for programme design and gap analysis — not a conformity assessment. Also credits NIST CSF 2.0, SSDF SP 800-218A, OWASP LLM Top 10, MITRE ATLAS and, in medtech, FDA §524B with ISO 14971.
5.3Read it the other wayNIST AI RMF → practice

An assessor starts from their framework, not yours. MAP, MEASURE and MANAGE run the lifecycle; GOVERN sits across all three, which is why it comes last here and first in a maturity conversation.

FunctionCategoryWhat it asks forDelivered by
MAP — understand the system in context
MAP 1Context, purpose and users established03 Inventory
MAP 2AI system categorized04 Risk engine · 05 Tiers
MAP 3Capabilities, targets, costs and benefits11 Approval intake
MAP 4Risks mapped across all components, third party included03 Inventory · 07 Vendor AI
MAP 5Impacts to individuals, groups and society04 Decision consequence factor
MEASURE — test it, and keep testing it
MEASURE 1Methods and metrics identified09 Lifecycle tests · 10 Assurance
MEASURE 2Evaluated for trustworthy characteristics09 Build / test · 10 Before production
MEASURE 3Mechanisms for tracking identified risks10 Runtime monitoring
MEASURE 4Feedback on whether the measurement works01 Assure
MANAGE — act on what you found
MANAGE 1Risks prioritized and responded to05 Tiers · 11 Approval
MANAGE 2Strategies to maximize benefit, minimize harm03 Lane controls · 06 Agentic control plane
MANAGE 3Third-party risks managed07 Vendor AI
MANAGE 4Treatments documented and monitored, incidents handled10 Detect & respond · 01 Monitor
GOVERN — cross-cutting, sits over all three
GOVERN 1Policies, processes and procedures in place01 Operating model · 11 Approval
GOVERN 2Accountability structures and named roles02 Accountability
GOVERN 3Workforce competency and team composition02 Accountability — thin
GOVERN 4Culture of risk awareness08 Shadow AI coaching
GOVERN 5Stakeholder engagement and feedback10 User feedback — thin
GOVERN 6Third-party and supply chain policy07 Vendor AI
GOVERN 3 and GOVERN 5 come out weak, and they are the same two every security-led programme misses. Workforce competency and stakeholder engagement have no real home here. An assessor working the GOVERN layer will stop there first.
5.4Which frameworks, and in what orderMust have, nice to have, and the agentic gap

They are not competing options. Buy the must-haves, borrow from the nice-to-haves, and know what none of them covers.

Must have
ISO/IEC 42001The management system. Certifiable.38 controls
ISO/IEC 27001Your security baseline underneath it.93 controls
NIST AI RMFThe language you use with the board.0 controls · 19 categories
EU AI ActIf you touch the EU. Not optional.Obligations, from 2 Aug 2026
Sector regulationFDA §524B, ISO 14971, IEC 62304 in medtech.Evidence packages
Nice to have
CSA AICMWhere the real AI security controls live.200+ objectives
OWASP LLM & AgenticWhat you red team against.Threat lists
MITRE ATLASAdversary techniques for scenario design.Knowledge base
NIST SP 800-218ASecure development for generative AI.Practices
ISO 23894 & 42005How to run risk and impact assessment.Guidance
Where the actual controls are
CSA AICM
200+
ISO/IEC 27001 Annex A
93
ISO/IEC 42001 Annex A
38
NIST AI RMF
0
EU AI Act
0
OWASP / ATLAS
0
The three most-cited frameworks contain no controls at all. NIST AI RMF gives outcomes, the AI Act gives obligations, OWASP and ATLAS give threats. If your programme cites only those, you have described your risk without treating it.
The agentic gap
Agent identityA non-human principal with its own credentialsNothing certifiable
Tool & action authorizationPolicy decision before the action runsAICM · OWASP, partial
Human-in-the-loopWhen approval is required, and by whomEU AI Act Art. 14, in principle
Kill switch & containmentStopping an agent mid-trajectoryNothing certifiable
Action audit trailReconstructing what an agent did and why27001 logging, not agent-aware
Multi-agent delegationAuthority passed between agentsNothing at all
Agentic red teamingAdversarial testing of an agent that plans, calls tools and keeps stateNothing at all
Trajectory & tool-misuse testingTesting the multi-step path, not the single promptNothing at all
Agent evaluation criteriaWhat "passed" even means for an autonomous systemNothing at all
No framework tells you how to test an agent

Every red team tool in this deck — PyRIT, garak, promptfoo — tests a prompt. Agents fail somewhere else: on the third tool call, in a delegation chain, when state from turn one poisons turn nine. Nothing in ISO 42001, the AI Act, NIST AI RMF or the SSDF tells you how to test that, what coverage looks like, or when to stop. Even AICM and the OWASP agentic list describe the risk without prescribing the test.

Which means you write the method yourself. Define your own trajectory test suite, your own tool-misuse cases, your own pass criteria — and document that you invented them, because an assessor will ask which standard you followed and there isn't one. Say so plainly rather than mapping it to a control that doesn't fit.

Two gaps survive the full stack. Agentic runtime authorization is addressed only by CSA AICM and the OWASP agentic list, neither of which you can certify against. And shadow AI discovery appears in no framework as a control — every one assumes you already know what you have. Practices 06 and 08 exist because the standards don't cover them, and that is worth saying out loud rather than mapping around.
Enable innovationMake safe and approved AI easy to use.
Protect what mattersProtect data, customers, and the business.
Build trustDemonstrate responsibility and transparency.
Continuously improveAI evolves. Governance must evolve too.
One enterpriseOne framework. Many AI. Consistent everywhere.