AI Governance

AI Governance and AI Management Systems

Governance and a management system are two different things. Governance directs and oversees; a management system delivers against that direction. Almost everything named AI governance is in fact AI management — real work, and most of what regulation requires. What it cannot do is decide what the AI is for. That decision belongs to the other half. Both are set out here, in the order you would take them on.

Written against
ISO/IEC 38500:2024 · ISO/IEC 38507:2022 · ISO/IEC 42001:2023 · ISO/IEC 42005:2025 · ISO/IEC 42006:2025 · ISO/IEC 23894:2023 · NIST AI RMF 1.0 · IEEE 7000-2021 · Regulation (EU) 2024/1689
Last updated
The principle underneath

Governance directs and controls. Management delivers.

Governance has a precise meaning and it is not a synonym for oversight, or for compliance, or for a committee. ISO/IEC 38500:2024 defines the governance of IT as the system by which its current and future use is directed and controlled. The governing body does three things: it evaluates the options, it directs by setting objectives and policy, and it monitors performance against that direction.

Management is the other half, and a different job. It plans, builds, runs and monitors the activities that deliver what the governing body directed. Governance sets the what and the why; management owns the how. The two collapse into one easily, and the cost of that is concrete: a steering committee that reviews project status is doing management, whatever the meeting is called.

What governance aims at is stated plainly in the same standard, and it makes a usable test: technology whose use is effective, efficient, and acceptable. Effective means it does what it was directed to do. Efficient means it is worth what it costs. Acceptable means it can be defended to the people it affects, to a regulator, and to the market. ISO/IEC 38507 carries that framing into AI specifically.

Governance — the governing body
EvaluateDirectMonitor

Evaluate the options and the pressures. Direct by setting objectives, policy and risk appetite. Monitor performance and conformance against that direction. Three verbs, none of which is project approval.

Management — the organization
PlanDoCheckAct

Plan against the direction received. Do — build, acquire, operate. Check through measurement, internal audit and management review. Act on what that showed. This is the loop ISO/IEC 42001 specifies, across clauses 4 to 10.

The two interlock, and the joins are where programmes fail. Direct is the input to Plan: management cannot align to a direction that was never set, so it substitutes its own. Check is the input to Monitor: a governing body with no reliable measurement is directing blind and usually knows it. And Monitor returns to Evaluate, which is how a direction gets revised rather than quietly outlived.

Consider what an AI governance programme usually contains. Inventories, policies, risk assessments, model documentation, monitoring, evidence, certification — most of that is plan, do, check, act. The platforms are tooling for it: workflow, registers, evidence capture. The committee is usually a review body reporting into the same layer.

ISO/IEC 42001 is the interesting case, because it is not purely a management standard. Its leadership and planning clauses require an AI policy, AI objectives compatible with the strategic direction of the organization, allocated responsibilities and authorities, and risk criteria that separate acceptable risk from unacceptable. That is direction, and it sits inside the standard. Management review in clause 9.3 is oversight. Both halves are present.

What the standard does not do is address the governing body, and it says so itself, at exactly the two points where the decision is properly the board’s. The note to clause 5.2 refers organizations developing an AI policy to ISO/IEC 38507. The note to clause 6.1.1 refers the determination of how much risk the organization is willing to take or retain to ISO/IEC 38507 and ISO/IEC 23894. The line does not fall between the two standards. It runs through the middle of ISO/IEC 42001, and the notes are where the standard draws it.

The other half is smaller and more specific, and nothing in a compliance programme forces it. What will we use AI for, and what will we not? How much judgement are we prepared to delegate to a system nobody can fully explain? How much opacity is acceptable where an outcome affects a person? What will we spend, and against what return? Who answers when it goes wrong? Five questions, and no management system can answer any of them on the organization’s behalf.

That is the failure mode: management running ahead of direction. Management without direction optimises for the only goal anyone actually stated, which is usually evidence of compliance, so the inventory gets built, the policy gets written, the assessment gets filed, and none of it is aligned to an intent, because no intent was set. The documents pile up and nothing anyone does is different.

None of which makes the management work optional. It is most of the effort, most of what regulation requires, and most of what an auditor will examine. The point is narrower: it is the second half of the job, and doing it first is expensive.

ISO/IEC 42001’s own introduction makes two points this practice works to. An AI management system is an extension of existing management structures rather than a parallel one, which is why the answer to “do we need a separate AI governance function” is usually no. And an organization has to find an appropriate balance between governance mechanisms and innovation. Direction that only constrains is a failure of governance, not an excess of it.

How this page uses the word

Governance is used here as an umbrella for both halves, because that is what the market means by it and pretending otherwise helps nobody. Where the difference matters, the two are named: ISO/IEC 38507 for the governing body’s work — direction, risk appetite, accountability, oversight of the estate — and ISO/IEC 42001 for the management system that delivers against it. Everything below covers both halves.

Part one

Direction

Evaluate · Direct · Monitor

The half the governing body owns. What it decides, and who answers for it.

This half consists of decisions rather than documents, so it is worth being concrete about what they are. Three verbs, and a small number of decisions nobody else in the organization can take on its behalf.

Evaluate · Direct · Monitor, in practice
Evaluate

What has to be in front of the governing body before it can decide anything: the estate and what it costs, what it returns, where exposure sits, and what has changed since last time. Most of that arrives from the inventory and the indicator sets. A board with no reliable measurement is directing blind, and usually knows it.

Direct

The decisions nobody else can take. What AI is for in this organization. Which uses are out of bounds regardless of return. How much judgement may be delegated to a machine, and where a human decision stays mandatory. Risk appetite, expressed as thresholds rather than adjectives. What will be spent, and against what expected return.

Monitor

Conformance against that direction, and performance against the objectives it set, reviewed on a fixed schedule. A board that looks only once something has broken is not monitoring — it is reacting, and by then the decision has already been made for it.

Risk appetite is the decision that makes everything downstream possible, and ISO/IEC 38507 places it here deliberately. How much bias, uncertainty, opacity and autonomy the organization will accept is not a judgement the team building the system should be making on its behalf. Setting that boundary is a governing-body decision and not a delivery one: it is where an organization states how much it is willing to be wrong about, and in which direction.

The second standing obligation is visibility: a maintained, high-level view of every AI system in use, which function owns it, what risk it carries, and how it serves the strategy. That is the inventory again, arriving from the governance direction rather than the risk one, and it is the input without which the first verb cannot be performed at all.

Cadence and forum are part of the decision rather than administration around it. Who meets, how often, and with what standing inputs determines whether direction gets revised as the estate changes, or is set once at launch and quietly outlived. Reviewing project status and asking whether the direction still holds are different meetings. The second is the governing one, and it needs its own place on the agenda.

Governance fails at the allocation more often than at the design. The controls exist, the policy names them, and when one does not operate there is no individual who has to answer for it.

The standards are explicit. ISO/IEC 42001 requires roles, responsibilities and authorities to be defined and allocated (clause 5.3, and A.3.2 for AI specifically), and extends the requirement across partners, suppliers and third parties (A.10.2). The EU AI Act adds a sharper version for deployers: human oversight assigned to people with the competence and the authority to intervene. Authority is the word doing the work there: oversight becomes a control at the point the overseer can stop the system.

The allocation, per system
Accountable

One named individual, senior enough to stop the system. Not a function, not a committee, not "the AI board". If naming who is accountable takes more than one name, nobody is.

Responsible

Who runs the monitoring, reviews the overrides, decides on retraining, and executes the retirement. Day-to-day operation of the controls, distinct from answering for them.

Consulted

Whose input is required before a decision is taken, and who therefore has to be asked rather than informed afterwards. For AI the list has to reach outside the delivery team, because changes that look purely technical often are not: swapping a model version, widening the data a system may reach, or extending it to an adjacent decision can move the organization’s legal position and not just the system’s behaviour. Under Article 25 of the EU AI Act, a new intended purpose can make a deployer into a provider.

Informed

Who hears about an incident, within what period, and in what form. Defined in advance, because the moment it is needed is the worst possible moment to design it.

Ordinary RACI is not difficult. The allocation is. Accountability for a decision no human made, taken by a system a vendor built, on data a third party supplied, is the allocation that takes the most work to settle, and an incident is a poor moment to be settling it for the first time. The same allocation then recurs at finer grain during delivery: every control has an operator, every lifecycle stage a sign-off, every incident a first responder and someone with the authority to stop the system. Those are set out with the controls themselves.

Part two

Delivery

Plan · Do · Check · Act

The half the management system runs. What must exist, where it lives, and how it is tested.

04

What controls are, and where they come from

Permalink to “What controls are, and where they come from

A control is a measure that maintains or modifies risk. Everything in this half of the page is either a control, a way of deciding which ones you need, or a way of checking that they work, so it is worth being exact about the term before sorting out where they originate.

The definition the risk standards use is deliberately broad. A control can be a process, a policy, a device, a practice, a condition, or an action. It does not have to be a document, and with AI it usually is not: an access boundary, an input validation step, a logging configuration, a threshold that halts a model, and a mandatory human checkpoint are all controls, and none of them is written prose.

The standards also separate the control from the control objective. The objective is the outcome you are trying to secure; the control is the measure taken to secure it. ISO/IEC 42001’s Annex A is a table of both, which is why two organizations can share an objective and implement quite different controls against it, and why an auditor asks what you were trying to achieve before asking what you built.

Four kinds, each catching a different moment
Directive

Establishes what should happen before anything runs: policy, standards, the intended purpose, and the action space a system is permitted to operate within.

Preventive

Stops the unwanted outcome from occurring: input validation, access restriction, a hard limit on which decisions a model may take unaided, refusal conditions.

Detective

Surfaces it once it has occurred: drift monitoring, sampled output review, an alert when an agent acts outside its boundary. Detective controls are what decide whether a problem is found internally or by a customer.

Corrective

Restores the position afterwards: rollback, retraining, suspension, notification, redress. Designing these before they are needed is what keeps an incident short.

One property is worth stating plainly, because the rest of this page depends on it. A control may not exert the effect assumed of it, the risk standards say so explicitly. A control can be present, documented, approved, and still modify nothing. That is why design effectiveness and operating effectiveness are tested as separate questions further down, and why a control inventory is not an assurance programme.

Reference controls in ISO/IEC 42001:2023

The standard is built in two parts. Clauses 4 to 10 carry the requirements — context, leadership, planning, support, operation, performance evaluation, improvement — and conformity is assessed against those. The controls reach the audit through the Statement of Applicability that clause 6.1.3 itself requires. The annexes then supply the control material: Annex A lists the reference control objectives and controls, Annex B gives implementation guidance under the same numbering, Annex C sets out candidate organizational objectives and risk sources, and Annex D covers running the management system across domains and sectors.

Annexes A and B are normative; C and D are informative. That distinction is narrower than it sounds, and it cuts both ways. Normative does not mean apply everything — Annex A is explicitly a reference set, selected from through the Statement of Applicability. What it does mean is that clause 6.1.3 requires Annex B’s guidance to be taken into account when implementing the controls you did select. Annex B is part of the standard, not commentary running alongside it, and working from the Annex A titles alone leaves half of what the standard provides unused.

Annex A · Table A.1

Nine control groups, ten control objectives, thirty-eight controls. A.6 states two — one for management guidance on development, one for the life cycle itself.

A.2Policies related to AI3
A.3Internal organization2
A.4Resources for AI systems5
A.5Assessing impacts of AI systems4
A.6AI system life cycle9
A.7Data for AI systems5
A.8Information for interested parties of AI systems4
A.9Use of AI systems3
A.10Third-party and customer relationships3
Controls in total38

Annex A is a reference set, and reading it as a mandatory checklist inverts how it is meant to work. Clause 6.1.3 introduces it as reference material; Annex A itself states that not all of its control objectives need be applied, and that an organization may design and implement its own. The mechanism runs the other way round: determine the controls your risk treatment actually requires, then compare against Annex A to verify that nothing necessary was left out.

The EU AI Act works differently again, and the difference matters in practice. It imposes requirements rather than controls: Articles 9 to 15 state what a high-risk system must achieve — risk management, data governance, documentation, logging, transparency, human oversight, accuracy, robustness and cybersecurity — not how to achieve it. Those requirements have to be mapped onto the controls you actually implement, so that every obligation has something concrete answering it, and so that one control can answer several obligations instead of each being built separately.

That mapping is the substance of an AI management system built for a specific organization. One control list, derived from your own risk treatment, cross-referenced to the legal obligations attaching to the systems in scope and to the Annex A objectives, with the gaps visible. Without it you get the common failure: a control set that satisfies an annex, a separate compliance workstream chasing the Act, and no single view of whether anything is actually covered.

Your own controls are the normal case rather than an exception. Clause 6.1.3 is explicit that the Annex A set is not exhaustive, that additional controls may be required, and that an organization may design them itself or adopt them from other sources. That matters most where published sets have not caught up: for agentic systems there is no adequate control catalogue anywhere yet, so controls have to be designed against the specific action space, and no annex will supply them.

The Statement of Applicability is where the choices get justified. It records the necessary controls together with a reason for each inclusion and each exclusion. An exclusion is defensible where the risk assessment did not find the control necessary and no external requirement demands it, and a justification for excluding a control objective may be given in general or for specific AI systems, whether the objective came from Annex A or was determined in-house. An auditor reads the SoA closely, because it is where judgement is visible.

NIST AI RMF is absent from this section by design. It specifies outcomes, not controls: an outcome taxonomy sits above any catalogue, so it can be satisfied with ISO/IEC 42001 controls, another set, or your own. Used as a control set it would leave a list of desired states with no measures underneath, nothing to exclude and therefore no Statement of Applicability. It belongs upstream, in the risk process, deciding which outcomes you are aiming at.

05

Responsible AI is a requirement, not a posture

Permalink to “Responsible AI is a requirement, not a posture

Ethics arrives in both standards as a requirement rather than an aspiration. ISO/IEC 42001 clause 6.1.4 requires a process for assessing the potential consequences of AI systems for individuals, groups of individuals and societies — that one sits in clauses 4 to 10, and an audit tests it. Annex A then supplies the reference controls that implement it: the impact-assessment process and its documentation (A.5.2, A.5.3), and the assessment of impacts on individuals and on society across the lifecycle (A.5.4, A.5.5). Responsible development and responsible use have their own reference controls (A.6.1.2, A.6.1.3 and A.9.2, A.9.3). Whether each belongs in your control set is a Statement of Applicability decision — but excluding all of them while claiming a responsible-AI posture is a position that will not survive an audit.

NIST is no softer. GOVERN 1.2 requires the seven characteristics of trustworthy AI — valid and reliable; safe; secure and resilient; accountable and transparent; explainable and interpretable; privacy-enhanced; and fair, with harmful bias managed — to be integrated into organizational policies, processes, procedures and practices. Not published as principles. Integrated.

Nor is the content left to interpretation. ISO/IEC 42001's Annex C is informative rather than binding, but it sets out the objectives the standard has in mind, and points to ISO/IEC 23894 for how each one relates to risk management.

Objectives named in ISO/IEC 42001 Annex C, which states it is not exhaustive
AccountabilityAI expertiseAvailability and quality of training and test dataEnvironmental impactFairnessMaintainabilityPrivacyRobustnessSafetySecurityTransparency and explainability

The gap is method, and that is where IEEE 7000 earns its place

A requirement to be fair does not tell a delivery team what to build. That is the gap most responsible-AI programmes fall into: an obligation everyone accepts, expressed as a values statement, handed over with no way to act on it. IEEE 7000 closes it by treating ethical values as requirements — elicited from stakeholders, and traced into the design the way a performance or security requirement would be.

It is a technique for satisfying the obligation the standards already create, not a separate obligation and not a certification anyone is asked to hold. Its provenance is largely European, developed out of Vienna and supported by German standards work, which tends to matter where the buyer expects ethics handled with method rather than assertion.

One distinction to keep clean: the value in value-based engineering is ethical, not financial. Conflating the two produces the worst of both — ethics defended only while it is profitable, and business cases padded with sentiment.

Specifying a control and enforcing one are different activities, and the gap between them is where most governance programmes are exposed. Written governance describes intent; production enforces behaviour. In between sits a policy requiring human oversight with no point in the workflow where a human can intervene, an approval requirement with no record that approval occurred, and a retention rule the pipeline never implemented.

The gap widens with autonomy. Agentic systems plan, call tools, and pass context between one another, so the decision path a control was written to constrain is assembled at runtime rather than at design time. Controls written for request-and-response do not bind a system that composes its own steps.

This is also where the data-side techniques earn their place — under the requirement they serve, rather than as a framework of their own. Provenance and lineage are what turn a claim about where training and retrieval data came from into a lookup instead of an assertion. Event logging and traceability are what turn a decision path into evidence instead of an intention. A defined intervention point, with authority attached and a record that it was available, is what stops human oversight being an entry on an org chart. Bounding the action space a system may operate within — which tools it may call, which data it may reach, which decisions it may take unaided — is what keeps an agentic workload inside something anyone can reason about afterwards.

Which of those you are legally obliged to produce depends on where you operate and what the system does. For high-risk systems within the EU AI Act’s scope the Act names them directly: data governance in Article 10, record-keeping in Article 12, human oversight in Article 14, accuracy, robustness and cybersecurity in Article 15. Outside that scope none of those articles binds, and the same capabilities are still what ISO/IEC 42001’s data and lifecycle controls ask for, and still what any usable answer to a customer, an insurer, a regulator in another jurisdiction or a court will require. The technique does not change. Only the reason you can be compelled to produce it does.

Designing that is architecture rather than paperwork, and it is the part that decides whether an audit is a formality or a problem.

What changes when you bought it rather than built it

Deployers control less of the stack and remain accountable for the outcome. The model is not yours to inspect, the training data is not yours to govern, and the vendor’s documentation is what you have to work with. That narrows the work. It does not remove it.

What stays in your hands is substantial: which decisions the system is allowed to touch, what data you feed it, where a human can intervene and whether that person has real authority to stop it, what you log and how long you keep it, and how you detect that behaviour has shifted. Those are configuration, integration, and process controls, and they are precisely what a deployer is asked to demonstrate.

The failure mode is assuming the vendor’s compliance covers yours. It does not. Their conformity assessment establishes that the system can be used lawfully. Nothing in it establishes that you are using it lawfully.

07

Testing that the controls still work

Permalink to “Testing that the controls still work

A control that was designed correctly and is never tested is an assumption. Assurance is what turns it back into a control.

There are two questions, not one. Design effectiveness asks whether this is the right control for the risk — whether a human review step placed after the decision has already taken effect can prevent anything at all. Operating effectiveness asks whether it is working now, at the volume it currently runs at, with the people currently doing it. A control can pass one and fail the other, and most audit findings live in exactly that gap.

Cadence comes from the risk tier rather than the calendar. A system informing credit decisions and a system drafting internal summaries do not warrant the same test frequency, and treating them alike is how an assurance budget gets spent on the wrong things.

Evidence is the third piece, and it decides whether an audit is routine or a project. If demonstrating that oversight happened means reconstructing it from memory and email, the control may well have operated, and it cannot be shown to have.

Controls drift too, and more quietly than models do. Permissions widen. An exception path gets added for a deadline and never removed. A threshold is relaxed after a noisy quarter and never restored. None of that surfaces in a policy review, because the policy still says exactly what it always said.

Assurance also has an allocation, and it is finer than the system-level one. Four roles per control, and the second is the one organizations skip.

Per control — who does what
Operates

The team running the control day to day: the reviewer exercising oversight, the engineer maintaining the logging, the analyst watching a drift threshold. Named per control, not per department.

Tests

Someone other than the operator. A control tested by the people who run it is a self-assessment, which cannot supply the independence the test exists for. That independence is what makes the result worth having.

Reviews

Who receives the findings, decides what they mean, and owns a remediation date. A named owner and a date are what close a finding; without them it ages, and then it gets reported as a trend.

Escalates

Who is told when a control fails, within what period, and who holds the authority to suspend the system rather than log the failure and carry on. Decided in advance, because the moment it is needed nobody is reading a policy.

Control indicators answer whether each safeguard is operating as designed. They are the earliest signal in the chain: controls degrade first, exposure rises next, value is lost last. Three groups — whether the controls exist, whether they work, and whether they are keeping up with an estate that changes faster than the review cycle.

A · Are the controls in place?
  • Share of production AI with complete logging and traceability, against the record-keeping obligations that apply.
  • Human-oversight coverage: high-impact decisions with a defined and exercised override point.
  • Share of systems carrying a named accountable individual, not a function, not a committee.
  • Share of AI systems classified against the frameworks that apply to them.
  • High-risk systems with a completed impact or conformity assessment.
B · Are they working?
  • Control-test pass rate by control type, rather than in aggregate — an average hides the one control that never passes.
  • Open findings against AI controls, and their age.
  • Incidents where a control existed and did not operate. The most informative number in the set, and the one that requires somebody to have been looking.
  • Exception volume, and how long exceptions persist past their stated expiry.
C · Are they keeping up?
  • Time from a system entering the estate to its classification.
  • Time from a purpose change to re-assessment — the EU AI Act Article 25 trigger that can turn a deployer into a provider, controlled rather than discovered.
  • Share of controls tested within their required cadence.
  • Share of use cases that cleared the pre-implementation gate before development began.

Group B takes the most work of the three, and tells you the most once it exists. Presence is straightforward to evidence and operation is not, which is why a control inventory and a control assurance programme are different things, and why an audit that samples operation reliably finds what a self-assessment did not.

09

One management system, made auditable

Permalink to “One management system, made auditable

ISO/IEC 42001 is where the guardrails, the board's risk appetite, and the ethical requirements become a single certifiable practice, the shape ISO 27001 gave information security, applied to AI. Three companion pieces complete it.

  • 01

    ISO/IEC 42001 — the management system

    Policy, risk assessment, controls, roles, and continual improvement across the AI lifecycle. Certifiable by a third party, which is what makes it evidence rather than assertion.

  • 02

    ISO/IEC 42005 — impact assessment

    Published in 2025. The method for assessing what an AI system does to the people and groups it affects — the input that makes a risk register about consequences rather than components.

  • 03

    ISO/IEC 42006 — certification bodies

    Published in 2025. Requirements for the bodies that audit and certify an AIMS, building on ISO/IEC 17021-1. It matters because it determines what an auditor will actually look for.

  • 04

    Connected, not parallel

    The AI management system extends the risk and compliance practice already running. A second, separate governance function competing with the first is a common and expensive failure.

Get direction and delivery working together.

Tell us what governance exists today and where AI actually runs. We will respond with a read on how the direction and the management system line up, and a sensible first step — often an ISO/IEC 42001 gap analysis.

Start a conversation