Governed AI is not the same as AI that delivers.
AI Value is the return on AI: whether a use case is worth doing, whether the one you built is delivering, and where the next increment of return actually sits. Adoption is no longer the constraint. AI now runs across decisions, processes, and customer contact, and the harder question is what any of it returned. The spending is real and recurring; unless measurement was designed in, the evidence a business case can point to is a demo, a usage chart, and a strong opinion. This work closes that gap in both directions — proving what value exists, and finding where more of it is available.
- Written against
- ISO 31000:2018 · ISO/IEC 42001:2023 · ISO/IEC 38507:2022 · ISO 21502:2020 · ISO 9001 · PMBOK Guide 8th ed.
- Last updated
Value sits inside the definition of the work.
The project-management body of knowledge does not treat value as an aspiration bolted onto delivery. The current PMBOK Guide defines a project as a temporary initiative undertaken to create value, defines value as the worth, importance or usefulness of something — financial or not — and names value per unit of investment as the ultimate indicator of success.
The same text is blunt about the mechanics. Projects are investments, and the expected value has to clear a threshold before the work is justified at all — which is what the gate on this page enforces. And what a project hands over is not yet value: deliverables lead to outcomes, outcomes to benefits, benefits to value, so the judgement runs on the far end of that chain, not on whether something shipped.
The definition scales, and the scaling is the point. A portfolio is managed as a group to achieve strategic objectives — it is where an organization decides which value is worth pursuing, and measured project value is the data that decision runs on: select, continue, or stop. That is the shape of this page: value decided at the gate, read across the portfolio, protected in operation, and delivered under a discipline that can prove it.
The measurement problem is structural, not lazy. AI is bought as a capability and deployed into a process that was already running, so the counterfactual disappears on day one. What the process cost beforehand is seldom on record, because there was no reason to capture it. Activity metrics fill the vacuum — queries served, seats licensed, documents summarised — and none of them is a financial outcome.
The second failure is timing. A business case written after deployment gets reverse-engineered from whatever happened to be instrumented, which means it measures what was easy to capture rather than what the investment was meant to change. By then the honest question — would we fund this again — has no defensible answer.
Perception makes it worse. Felt speed and measured speed come apart: effort falls while elapsed time does not, so the reported experience runs ahead of the number. That gap is not dishonesty. It is also the argument for measuring rather than surveying.
The cheapest place to fix an AI investment is before it is committed to, and committing rarely means building. Far more often it means buying a licence, switching on a feature that shipped disabled, or extending a tool into a decision it was never bought for. Every proposal passes the same gate whichever of those it is, and a meaningful share should not clear it. What does clear it enters the inventory, long before anything is live.
- 01
Expected value
Which financial line moves, by how much, and who owns that number. Revenue, cost, cycle time, or margin — named before the commitment rather than reverse-engineered afterwards. And stated as a range with a confidence attached, because a forecast of thirty percent that could plausibly be eight is a different investment from one that could plausibly be forty.
- 02
Feasibility
Whether the capability is actually reachable with the data, the platform, and the integration surface available, and what would have to change if it is not.
- 03
Risk
Both directions, both priced. Downside: what the system could get wrong, at the volume it runs at, what remediation would cost, and the standing cost of the controls it obliges, which rises steeply where it falls in a high-risk tier or puts you in the provider role. Upside: how much confidence the benefit estimate actually deserves. A proposal can clear on headline value and fail here. That is the gate working.
- 04
Data readiness
Whether the inputs exist at the quality, provenance, and permission the case requires — and, where the system is bought rather than built, whether you control enough of the input to be accountable for it. Data failures often arrive disguised as model failures.
- 05
Kill criterion
What would have to be true for this to be stopped, written down before it starts. Not a risk register entry — a threshold, named at the point where nobody is yet invested in the answer. A use case with no stated stopping condition does not get killed when it underperforms. It gets defended, and then renewed.
Risk belongs in the ROI, in both directions
A business case that states one number for benefit and files risk separately is not a case. It is a forecast with the uncertainty taken out of it. Under the definition this practice works to, the uncertainty in the benefit is itself risk — deviation from the expected, in the favourable direction as much as the adverse one.
So the forecast improvement carries a range and a confidence rather than a point. A thirty percent reduction in handling time that assumes full adoption, clean inputs, and no exception path is a different proposition from the same thirty percent with those assumptions tested. An untested forecast clears approval faster, which is a large part of why realised value lands short of the paper.
The downside enters the same arithmetic rather than a parallel register: the cost of decisions the system gets wrong at the volume it runs at, remediation and redress once they surface, and the standing operating cost of the controls the system obliges — oversight staffing, logging, evaluation, assessment. That last item is the easiest to leave out. A high-risk-tier system is more expensive to run than a minimal-risk one doing similar work, and that difference belongs in the case at gate time, before it surfaces in year two.
Netted, this changes the ranking. Two use cases with the same headline return are not equivalent if one has a narrow distribution and the other a wide one, and a portfolio ordered on expected value alone systematically over-selects the volatile. Risk-adjusting the comparison is the difference between a list of projects and an investment decision.
AI creates value when it improves what an organization achieves without creating consequences that outweigh the benefit. The financial return matters and it stays the measure this page is built on — but decision quality, resilience, human capability, customer trust and the effect on the people a system reaches decide whether that return is real.
Automation changes more than cost. A decision to automate has to account for its effect on people, skills, accountability and organizational resilience, not only the hours it removes. Fifty people at ten hours a week is a number, not yet a case. Whether the expertise that verified the work still exists once nobody does it by hand, whether the process still runs if the model or the vendor behind it becomes unavailable, and who absorbs the consequence when the output is wrong at the volume the system was bought to handle — all three belong in the same decision as the saving.
Where they carry a cost, they belong in the ledger, discounted like any other exposure. Where they do not, they belong in the criteria: the conditions a use case has to meet to proceed, and the ones that stop it. Neither belongs in a benefits case as a sentiment.
Value cannot be managed across AI nobody can see. Before anything can be ranked there has to be one place holding every use case and every system, and that inventory is the same asset the risk and governance work reads, from a different angle. Built once, it answers three different questions.
Value — the portfolio
What each system costs to run, what it returns against its case, and where the next increment of investment earns the most. Also where to stop: a healthy portfolio retires systems that no longer justify their cost.
Risk — the exposure map
What each system touches, what it would cost if it went wrong at the volume it runs at, and how much confidence the original value estimate still deserves. The same rows, read for what could deviate — in either direction.
Governance — the register
Who owns each system, which controls apply, and what evidence exists that they operate. This is the view ISO/IEC 38507 places with the governing body.
Every use case and every AI system in one place: proposals that cleared the gate alongside systems already in production. That dual scope is what makes it a portfolio rather than an asset register: the next increment of investment can only be placed well if candidates and incumbents sit on the same page. Both standards point at it from the risk side: ISO/IEC 42001 asks for the resources behind each AI system to be identified and documented (A.4.2), and NIST AI RMF asks for the inventory outright (GOVERN 1.6). The value case is simpler still. Nothing downstream can be ranked without it.
Not every AI investment is a project. Treating the portfolio as single projects at one end and one aggregate at the other loses the level in between, which is where most AI value is argued about. Several use cases sharing a data foundation, a platform, and the literacy to operate them form a single benefit case, and it behaves differently from the sum of its parts: the shared components produce no benefit on their own, and the use cases cannot produce theirs without them.
Managed as a group, the benefit case runs a life cycle rather than ending at a delivery date — benefits identified from the business case, planned and mapped, monitored through delivery, transitioned to the areas that will actually use them, and sustained after the delivery work closes. The last two are where the benefit either materialises or does not. Model monitoring, retraining and adoption support are obligations planned before handover, not discoveries made afterwards. That life cycle and its register are programme management’s own. Nothing in them is invented for AI.
The artefact that makes this governable is a benefits register, and one property of it settles the argument over what a shared data platform is worth: the register is benefit-first, and every benefit is mapped to the components that produce it. A shared data foundation therefore gets no line of its own. Its contribution appears only through the benefits mapped to the components it enables — which is both the honest answer to "what did the platform return" and the reason platform investment is so hard to defend when it is assessed on its own.
Benefit viability is then a standing question rather than a decision taken once. If life-cycle cost overtakes the benefit, or the window the benefit depended on closes, the case is reassessed — and a use case can be stopped or reshaped mid-flight without the wider benefit case having failed. The benefits management plan carries a process for removing a benefit that is no longer needed, which is what makes the kill criterion set at the gate enforceable later, once stopping has become socially expensive.
Risk and return on one ledger
Exposure and upside reach a decision by different routes: different people, different times, different scales. Nothing forces them onto one. Valued together means valued on one scale: expected return discounted for the confidence it deserves, less the expected cost of the ways it can go wrong, less the standing cost of controlling it. Ranked that way the inventory stops being a collection of projects and starts behaving like a portfolio — investment concentrates where the risk-adjusted balance is strongest, and attention goes where exposure has quietly outgrown the benefit. That comparison is what makes the decision defensible to a board.
Value has to be assured, not just ranked
Ranking the AI portfolio decides where investment goes. It does not establish that what was funded turned into anything, and the chain from one to the other is longer than it looks: deliverables produce outcomes, outcomes produce benefits, benefits produce value. Each link can hold while the next one fails, and value reporting that jumps from the first to the last is describing an intention.
Assuring that chain across the portfolio is ordinary discipline — tracing requirements through to what was actually built, reviewing at a defined number of points placed where decisions genuinely sit, and separating acceptance criteria from completion criteria. That distinction is the one the quality plan draws a level down: work can be finished and still not acceptable, and a portfolio tracking only completion will not register the difference.
The empty benefit record is a separate failure and a worse one. A deliverable accepted exactly as specified can still produce no outcome, and no amount of acceptance discipline will detect that — only measuring the next link does. A portfolio can therefore hold a complete record of accepted deliverables and an empty record of realised benefits without anything in its reporting registering a contradiction.
Measured value also needs an agreed measurement framework, decided before the numbers arrive rather than after. Value reporting is contested by nature — it decides where budget goes next — and a measurement basis agreed in advance is the only version of the conversation that is about the AI rather than about the arithmetic.
One method, three levels
The same discipline reads at each level, which is what stops the three from becoming three separate reporting regimes. A use case is a project, measured against its own baseline. A benefit case spanning use cases is a programme, measured on benefits realised against benefits planned. The whole is a portfolio, measured on value added against the strategy it was funded to serve. Planned-versus-realised is the constant, and only the unit changes.
Who answers for the value is not a matter of taste — the standards settle it. The sponsors of the portfolio own its value. The organization that receives a benefit owns keeping it alive. So the value number cannot be outsourced, and anyone offering to own it is selling something that belongs to the client. What can be handed over is the discipline that produces the number.
Opportunity means nothing if the system degrades on the way to it. Models drift as the world moves away from their training data. They hallucinate under ambiguity. They buckle under context load, where accuracy falls quietly as input grows rather than failing in a way anyone would notice.
None of that announces itself. A degraded model returns answers with the same confidence as a healthy one, which is why degradation is discovered by a customer, an auditor, or a headline rather than by the team that owns it. Evaluation, monitoring, and traceability are what turn a silent failure into a visible one.
This is where value and risk stop being separate disciplines. A model that has drifted is a business failure and a compliance exposure at the same moment: the value case erodes and the obligation to demonstrate control is breached, from a single root cause.
What clears the gate still has to be delivered
Permalink to “What clears the gate still has to be delivered”Between the gate that admits a use case and the indicators that report what it returned sits the work that decides whether the second bears any relation to the first. A gate that admits into an unmanaged pipeline has decided nothing. The proposal has been approved and released, which is a different thing, and the measurement at the far end will faithfully record the consequences.
Delivery is a managed activity or it is an assumption
The disciplines that turn an approved case into a delivered one are not new and not specific to AI. Work is planned in phases with a decision at the end of each — continue, change, or stop — against criteria set before the phase began. Decisions that belong to the client stay with the client and are named in advance rather than absorbed by whoever is closest to the build. Progress is reported against a baseline rather than against the last report.
None of this is novel, which is precisely the point. The case for running without it is that the technology is exploratory — and exploratory work needs more decision discipline than routine work, not less. The phase-gate model is decades old and the project-management standards have carried it since.
Three things AI changes, and only three
Delivery method is adapted for AI, not imported unchanged and not reinvented. Some assumptions break, and a usable method names them rather than absorbing them into contingency.
Data is a design input, and usually the dominant one.
Classic design control assumes the specification governs the output. In a learned system the training data shapes behaviour more than the specification does, and a plan that treats data as a supply problem rather than a design decision has misplaced the thing most likely to determine the result.
The output is uncertain by construction.
A conventional deliverable either meets its specification or does not. A model performs across a distribution, and the accurate statement of what it does is statistical. Plans that promise a capability rather than a measured performance level against a defined test set are promising something the technology cannot deliver.
Validation is iterative and does not end at release.
The evidence that a system works is produced by evaluating it, learning that the target was wrong or the data insufficient, and evaluating again. This is a planned loop with a defined exit, not a symptom of poor estimation — and treating it as slippage is how AI projects acquire the reputation for never finishing.
What "delivered" is allowed to mean
Delivery ends in a judgement: this is done. Left informal, that judgement is made by whoever is most tired of the question, and the record afterwards shows an artefact that exists rather than an artefact that was verified.
A quality plan makes it explicit. The acceptance criteria for each deliverable — what makes it acceptable, not merely finished — are written before work on it starts, and are not edited once work begins. A criterion later found to be wrong, or overtaken by a requirement that has legitimately changed, is superseded with a dated note rather than rewritten to match what was produced. Each criterion is answerable yes or no against evidence somebody can locate. The verification method is named in advance. The result is recorded as passed, failed and corrected, or failed and open, with the name of whoever checked it and the date they did. An artefact that fails and stays open is not released, unless someone with the authority to carry the consequence accepts it in writing and is recorded as having done so.
This is ordinary quality practice, which is the argument for it. What is not ordinary is where it meets machine learning.
Where classic quality assumptions break on AI
The adaptations above change how the work is planned. What follows changes how the result is judged, and it is the harder half. Quality management arrived at its methods through manufacturing and engineering, and four of its load-bearing assumptions do not survive contact with a learned system. Naming them is not a caveat. It is the difference between a quality plan that works and a quality plan that is decoration.
Conformity stops being per-unit.
Established practice assumes a nonconforming output can be identified and separated from the good ones. Every output of a generative system is new, and a wrong one — a fabricated citation, an incorrect tolerance — may be indistinguishable from a correct one without expert review. The discipline has met this before in another form: where a process produces output that cannot be fully verified by inspecting it, as in welding or sterilisation, the answer is to validate the process and the conditions it runs under rather than the unit, and to qualify the people operating it. Inference is that kind of process. What follows is the same move — conformity defined statistically across a distribution, with sampling, review gates and output monitoring designed in deliberately.
The specification does not determine the behaviour.
Design control verifies a design against inputs stated up front, and the assumption beneath it is that those inputs govern what gets built. In a learned system the training data governs more of the result than the specification does. Data quality, provenance and preparation are therefore design inputs and have to enter design control as such — otherwise the control is a formality performed over the least influential half of the system.
Validation does not stay valid.
Validation assumes conditions of intended use can be characterised once. A model valid at deployment degrades as the world moves, and nothing changes inside the system to trigger a review — the change happens outside it. Periodic revalidation and performance monitoring are already ordinary practice. What has to be re-keyed is the trigger: from the engineering change request to the monitored measurement.
There is no calibration chain for model evaluation.
Measurement discipline grounds instruments in reference standards, and that route is unavailable here — benchmarks are not reference standards, and results move with the test set. The obligation does not disappear with the route. The instrument still has to be fit for what is being measured. Its suitability simply has to be argued for each system rather than inherited from a calibration regime.
All four have quality-engineering answers once they are named. What no quality system supplies at any point is an account of what the system does to the people it touches — which is what an AI management system is for, and why the two run together rather than one instead of the other.
The shortfall is visible long before the end
Delivery reporting usually answers whether work is proceeding. The sharper question is whether the work the value case depended on is being accomplished at the rate the case assumed, and that one can be answered in flight rather than after the fact.
It takes three numbers and no new instrumentation. What the plan said would be finished by now. What has been finished. What has been spent getting there. The first two compared give the shortfall against schedule, the second and third the shortfall against cost, and both extrapolated give a defensible estimate of where the work lands if nothing changes. That turns "are we on track" from a judgement into arithmetic, and it surfaces the gap while there is still budget left to respond to it.
This is earned value management, and two properties make it fit AI rather than fight it. The unit does not have to be money — the technique is defined to work in whatever unit measures the work, which matters where an AI benefit is real and not naturally monetary. And it is defined for iterative delivery, not only for fixed-scope plans. Neither is a concession made for AI. Both were in the method already.
What the arithmetic tracks is output against plan, and output is not value. A project can earn every unit of its baseline and return nothing. Applied above the project the principle extends and the unit changes with it, but here it answers a narrower question: whether the thing the case was betting on is being built at the rate the bet assumed. That is worth knowing early. The principle underneath is smaller than the technique and survives without it — something was planned, something happened, and the gap between them is reportable long before the end.
Performance indicators answer one question: is the AI delivering what it was funded to deliver? Four groups — what has actually landed, what is still owed, what it took to get, and whether the system for choosing between use cases is working. They are read alongside the exposure indicators on the risk page and the control indicators on the governance page, because a strong performance number can mask weakening controls beneath it.
- —Realised value against the range the case was approved on, per use case, not against a point estimate nobody committed to.
- —Cumulative value realised across the portfolio, against the cumulative forecast it was funded on.
- —Share of use cases with a defined and measured financial outcome, rather than an activity metric standing in for one.
- —Value still being realised twelve months after go-live, against value at first measurement. AI benefit decays; adoption drifts and processes move around it.
- —Adoption and utilisation of deployed AI against provisioned capacity.
- —Value approved and not yet realised: the committed pipeline, and how long each tranche has been waiting for it.
- —Value by stage — proposed, cleared the gate, in delivery, live but not yet measured, measured. Where value accumulates is where the portfolio is stuck.
- —Realisation rate: measured value as a share of approved value. A persistent gap is either delivery lag or forecasts that never survived contact with production.
- —Value written off — approved cases abandoned or quietly stalled, and what they were worth on paper. Unrecorded, this is the portfolio’s largest blind spot.
- —Total cost of ownership per system, including the standing cost of oversight, logging, evaluation, and assessment, not licence and build alone.
- —Cost per decision or per inference, and its direction as volume grows. Unit economics that work at pilot scale do not always survive production.
- —Share of total AI spend attributable to systems with no measured outcome.
- —Share of proposals that do not clear the gate. A gate that rejects nothing is not a gate.
- —Forecast accuracy across the portfolio: realised value against the approved range, and whether the misses run consistently in one direction.
- —Time from commitment to first measured value, not time to production, which is an engineering milestone rather than a business one.
- —Decommission rate: systems retired for want of value, which a healthy portfolio does deliberately rather than by neglect.
Forecasts only become evidence once they are compared against outcomes. The in-flight numbers are the simplest of the four groups to produce: a business case records its promised value once, at approval, and nothing downstream requires anyone to revisit it, which is how a portfolio comes to owe more than anyone has noticed.
ISO 31000 — risk management creates and protects value
The idea that value and risk belong in one conversation is not a positioning device. It is the stated purpose of the general risk management standard: risk management exists for the creation and protection of value, not merely the avoidance of loss. The definition of risk carries the same logic: an effect of uncertainty is a deviation from the expected in either direction, so opportunity sits inside the term rather than beside it. Read seriously, that makes finding where AI could deliver more part of risk management itself, which is exactly how it is treated here.
Make the return on your AI provable.
Tell us where AI is deployed and what it was meant to change. We will respond with a read on what can be measured today, what it would take to make the case provable, and a sensible first step.