AI Risk

AI risk is now the hard part, and it is wider than the model.

Deploying AI is no longer the challenge. The challenge is knowing where it creates exposure — against regulation, against your own risk appetite, and against the reality that these systems fail in ways traditional software does not. But exposure is only half of what risk means, and it is the more familiar half.

Written against
ISO 31000:2018 · ISO/IEC 23894:2023 · ISO/IEC 42001:2023 · ISO/IEC 38507:2022 · NIST AI RMF 1.0 · Regulation (EU) 2024/1689
Last updated
The principle underneath

Risk management exists to create and protect value

ISO 31000 states the purpose of risk management in a single line: the creation and protection of value. Not the avoidance of loss. That is a deliberate choice of words, and it changes what the work is for.

It follows from how the standards define risk itself. ISO 31000 puts it as the effect of uncertainty on objectives. ISO/IEC 42001 states it in the harmonised form — risk is the effect of uncertainty — and Note 1 to the entry keeps the decisive part intact: an effect is a deviation from the expected, in a positive or negative sense (clause 3.7). Its planning clause is titled accordingly: actions to address risks and opportunities (6.1). Upside is not an optimistic addition to a threat register. It sits inside the definition.

Used in full, the definition cuts both ways. A programme that counts only what could go wrong will under-invest in the AI that is working, and will hand the board a register of threats with no view of where return actually sits. Read properly, the same evidence answers both questions, which is why the risk work and the value work are one engagement seen from two directions.

AI enters an organization through many doors — sanctioned platforms, embedded vendor features, employee tools, and increasingly autonomous agents. Each one can shape a decision, touch personal data, or trigger an action. Very few were ever reviewed as AI, and fewer still were classified against the regulation that now governs them.

The first failure is rarely a bad model. It is not knowing where the risk sits. A system that looked harmless in a demo becomes a high-risk system the moment it informs an employment, credit, or safety decision. Exposure hides in tools that were never catalogued and in data flows never scoped as AI — until an auditor, a regulator, or an incident surfaces it.

AI risk is not one thing. It spans distinct domains, each with its own failure modes, owners, and controls, and each with a different way of becoming expensive. A credible risk picture covers all six, not just the accuracy of the model.

  • 01

    Technical & model risk

    Hallucination, drift, bias, and brittleness. Models fail probabilistically and degrade as the world moves away from their training data — quietly, because a degraded model answers with the same confidence as a healthy one. The failure mode is not an outage. It is a slow decline in decision quality nobody is watching for, found when a customer complains or an auditor samples output. By then the wrong decisions have been made, at whatever volume the system runs at.

  • 02

    Data risk

    Quality, provenance, privacy, poisoning, and the copyright exposure carried in training data. Data failures are invisible by construction: a model trained on flawed or improperly licensed data passes its benchmarks and fails where it matters. The consequence usually arrives from outside, a regulator asking where the data came from, a rights holder asking under what licence, a subject asking to be removed from something that cannot easily forget.

  • 03

    Operational risk

    Shadow AI, vendor concentration, and integration failure. The largest uncontrolled exposure is often the AI nobody approved: a feature switched on inside a tool the organization already pays for, running outside sanctioned workflows. It carries the same obligations as a deliberate deployment and none of the controls. Vendor concentration compounds it — when one provider sits behind six of your processes, their outage, price change, or deprecation becomes your continuity problem.

  • 04

    Governance risk

    No clear owner, no inventory, no audit trail. What cannot be seen cannot be governed, and naming every AI system in use — let alone who is accountable for each — is harder than it sounds once vendor features and employee tools are counted. Failures in the other five domains arrive unannounced when this one is weak: without ownership there is nobody whose job it is to notice, and without records there is no way to reconstruct what happened once somebody does.

  • 05

    Legal & regulatory risk

    Penalties, liability split across providers and deployers, and cross-jurisdictional conflict. The exposure is rarely the fine alone. It is the obligation to suspend a system mid-process, the discovery burden of proving how a decision was reached, and the liability that travels onward through your own contracts to customers who relied on the output.

  • 06

    Agentic risk

    Autonomous agents that plan, call tools, and pass context between one another. Traceability breaks where context is compressed, boundaries blur where one agent invokes another, and the decision path a control was written for is assembled at runtime rather than at design time. No framework governs it fully yet, so the controls have to be designed from scratch, and it is the fastest-growing part of most estates.

Likelihood on its own does not answer the question a board actually asks, which is what happens if this is real, and for AI the answer is rarely a single event.

The first cost is decisions. A system that has drifted does not stop — it keeps deciding, at volume, until somebody notices. The interval between degradation and detection is where the damage accumulates, and it is usually measured in months, because nothing was instrumented to shorten it.

The second is remediation. Re-running affected decisions, contacting affected people, and correcting downstream records is manual work that scales with exactly the volume the system was bought to handle. Automation that saved a year of effort can cost more than that to unwind.

The third cost leaves your organization. Your contracts promise outcomes to your own customers, and a model failure does not stay inside your organization. Add the copyright and IP exposure carried in generated output, and the retrofit cost of adding provenance, logging and oversight to something already in production — always dearer than designing them in.

What the incident record shows
  • Only around 2% of reported AI incidents occurred pre-deployment. Almost all harm materialises in operation, not in development.
  • Roughly a third of reported incidents involve systems that would fall in the EU AI Act’s high-risk tier; about 3% in the prohibited tier.
  • About half were assessed as intentionally caused, which means about half were not.
  • The largest domain is malicious use, dominated by fraud, scams and targeted manipulation. The next is system safety, failures and limitations.

MIT AI Risk Initiative incident tracker, classifying reports from the AI Incident Database. Treat the proportions as indicative rather than representative: reporting is voluntary, so the sample skews toward incidents somebody chose to publicise, and classification is automated, without manual review.

The 2% figure is the one that matters most here. If almost all harm materialises after deployment, then almost all of it lands on the deployer, the organization running the system, not the one that built it.

Every finding needs a yardstick. Without one an assessment produces observations, and an observation carries no obligation to act.

Risk appetite is the amount and type of risk an organization is willing to pursue or retain. For AI it has to be stated in terms the estate can actually be measured against: how much model uncertainty is acceptable in which decisions, how much autonomy a system may exercise unaided, how much opacity is tolerable where an outcome affects a person, and how much dependence on a single provider the business will carry.

ISO/IEC 38507 places that decision with the governing body, and the placement matters more than it sounds. A threshold set by the team that built the system is a preference. The same threshold approved by the board is a control, with an owner, and a consequence when it breaches.

Tolerance is the narrower question: how far a specific system may deviate before somebody must act. That is what the indicators further down are measured against, and it is why they are set before the assessment rather than after. Thresholds chosen once the findings are known tend to be chosen so that nothing has breached.

Finding exposure is not a one-off exercise. The work follows the risk management process ISO 31000 defines, the same discipline NIST's AI RMF re-cuts for AI, with governance lifted out as a continuous function rather than a first step. Four stages, one foundation, and a layer that never stops.

Govern — across all four

Communication, consultation, recording and reporting are not a stage. ISO 31000 runs them alongside the entire process; NIST lifts the same idea into a cross-cutting function that informs the other three. Governance that happens only at the end is the failure this page is about.

01MAP

Scope, context and criteria

What the assessment covers, which systems and processes fall inside it, and the criteria they will be judged against — including the risk appetite approved at board level under ISO/IEC 38507. Findings without criteria are observations. Findings against criteria are decisions.

02MAP · MEASURE

Risk assessment

Three moves in one stage. Identify what is running and what it touches; analyse how each system fails across the six domains; evaluate the result against the criteria set in stage one. Classification against the EU AI Act, NIST AI RMF and ISO/IEC 42001 happens here.

03MANAGE

Risk treatment

Decide what changes. Controls designed and built, exposure formally accepted, an activity stopped, or risk shared with a third party. This is the stage where findings become decisions, and where the architectural guardrails come from.

04MEASURE · MANAGE

Monitoring and review

Indicators measured against the thresholds set in stage one, so movement surfaces before an audit does. When the estate grows, the regulation moves, or appetite shifts, it returns to stage one.

Foundation — the AI inventory

Every use case and every AI system in one place: what is proposed as well as what already runs. ISO/IEC 42001 never uses the word inventory, but the objective behind its A.4 controls is that an organization accounts for the resources of its AI systems — components and assets included — precisely so that risks and impacts can be understood and addressed, and A.4.2 requires those resources to be identified and documented. NIST states it outright in GOVERN 1.6: mechanisms are in place to inventory AI systems, resourced according to risk priorities. Without it, the four stages above run on recollection.

The process is iterative, not sequential — ISO 31000 is explicit about that, and AI estates change faster than most risk registers were designed for.

Stage 03 in full — the four decisions
Avoid

Do not run the use case, or stop running it. The option most often missing from an AI programme, because stopping reads as failure rather than as a decision.

Reduce

Design and build controls until residual risk sits inside appetite. This is where the architectural guardrails come from, and where most of an engagement usually is.

Share

Move part of the exposure to a third party by contract or insurance — noting that regulatory duties on a deployer do not transfer, whatever the contract says.

Accept

Retain the risk deliberately: recorded, owned, and revisited. A documented accepted risk is a governance artifact, not a failure. Without a mechanism for accepting risk deliberately, it gets accepted by default instead — silently, and by whoever happened to be closest to it.

Regulation is one source of risk among the six, not the subject of the discipline. It earns disproportionate attention because it is the one source that arrives with a deadline and an enforcer attached.

Two things about it change your exposure more than any date. The first is role: the EU AI Act assigns duties by role, and the role most organizations occupy is deployer — the party using an AI system under its own authority. Buying software with AI inside makes you one, whether or not procurement noticed, and the duties that attach are not transferable by contract. The second is that the line moves. Under Article 25 of the Act — responsibilities along the AI value chain — rebranding a system, modifying it substantially, or repurposing it into a use it was never sold for attaches the full provider obligation set instead.

The obligations are live rather than pending. Transparency duties and enforcement against general-purpose AI providers have applied since August 2026, and the deferral of the high-risk tiers moved the date those obligations begin to apply, leaving the requirements themselves untouched. The full timetable, the role-by-role obligations, and the supervisory authorities for the DACH region and Canada are set out on the standards page.

The definition this page opened with cuts both ways, and a risk function that counts only downside is failing at half its job.

Under-investment is exposure. So is a control posture so cautious that value never ships, a governance function able only to say no is producing risk rather than managing it, and the deviation from the expected outcome is every bit as real as an incident.

The competitive form is sharper. If a comparable organization automates a process you left manual, the gap compounds while your register records nothing, because nothing went wrong. Opportunity risk is invisible to a threat-led programme by construction.

The same evidence answers it. The inventory that maps exposure also shows where a system is under-used, where a proven capability has not been extended to an adjacent decision, and where a manual process sits beside an automated one doing similar work. The same rows are a risk register when you read them for downside, and a pipeline when you read them for upside. That is the bridge to the value work, and the reason the two are one engagement.

These measure exposure, not whether the compliance work got done. Whether classifications are complete and assessments have been filed is control effectiveness, and it sits with the governance indicators. What follows is what moves when the risk itself moves — read against thresholds set in advance, which is why appetite comes before measurement rather than after it.

A · Exposure — how much is riding on it
  • Consequential decisions influenced by AI per period, and the direction of travel.
  • Share of those taken without human review before they take effect.
  • People subject to an AI-influenced decision per period.
  • Value flowing through AI-influenced decisions — credit extended, claims settled, candidates screened, spend approved.
  • Systems running in production that fall in the high-risk tier.
  • Share of business-critical processes dependent on a single model or provider.
B · Early warning — signals that move before harm does
  • Human override and correction rate on AI output. The strongest signal available without access to the model, and it rises before quality failures become visible.
  • Share of outputs falling below an agreed confidence threshold.
  • Drift against baseline, where the model is yours to instrument.
  • Share of inputs failing validation, and retrieval corpus past its freshness limit.
  • Shadow-AI discoveries per period, and time to bring each into the sanctioned estate.
  • Vendor AI features activated without review — the bought-software exposure, made measurable.
  • Systems whose intended purpose has changed since classification — the EU AI Act Article 25 trigger that turns a deployer into a provider, watched rather than discovered.
  • Agent actions outside the defined action space, and the depth of agent-to-agent delegation.
C · The other direction — opportunity, measured
  • Decisions still made manually where an approved capability already exists.
  • Use cases stalled between gate approval and production, and for how long.
  • Provisioned capacity paid for and unused, where no decision has been taken either to use it or release it.
  • Time from an opportunity being identified to a decision being taken on it.

These are illustrative. A usable set is derived from the appetite established earlier on this page: the threshold is what makes an indicator actionable, and the approved appetite is what makes the threshold defensible. Several require instrumentation that has to be built first, which is usually part of the work rather than a precondition for it.

What happens when one breaches

An escalation path is what separates an indicator from a dashboard tile. Each threshold needs three things attached before it is useful: who is told, within what period, and what decision that triggers — re-assessment, suspension, or acceptance at a level with the authority to accept. Each also needs a stated direction, because several of these are two-sided: a falling override rate can mean the model improved, or that the reviewers stopped looking.

Foundation

ISO 31000 and ISO/IEC 23894

AI risk management is a specialization of established practice, not a new discipline. ISO 31000 supplies the principles, framework and process — including the two-sided definition this page opens with. ISO/IEC 23894 layers the AI-specific guidance on top, mirroring the same clause structure and extending the principles where AI behaves differently, particularly around dynamism, the quality of available information, and human and cultural factors. It cannot be implemented on its own; it assumes the general baseline underneath. Where a mature risk function already runs, the machinery is largely in place and only the AI-specific extension is new. NIST’s AI RMF reaches the same place from another angle: Govern, Map, Measure, Manage is that process re-cut so governance runs continuously instead of as a precondition. Adopting it over an existing ISO 31000 practice is an extension, not a second programme.

Get a defensible view of your AI risk.

Tell us where AI operates in your organization. We will respond with a read on where the exposure sits, where the upside sits alongside it, and a sensible first step — usually a bounded AI risk assessment.

Start a conversation