Insights · AI Governance

AI controls in production: where they live, whether they still work, what to measure

A control written into a policy may never reach the running system. This piece is for enterprise architects and assurance leads: where AI controls have to be enforced, how to test that they still work, and which indicators show it.

By Reinhardt Mühlhäusser · Published

Written against
ISO/IEC 42001:2023 · Regulation (EU) 2024/1689
Last updated

Writing a control down is the easier half. The harder half is production, where a policy asks for human oversight and the workflow offers no point to intervene. Agentic systems widen that gap, because they assemble their decision path at runtime. And a control that worked at launch can drift as permissions widen and exceptions pile up. This piece follows a control from design into operation: where it has to live in the architecture, what changes when you bought the system rather than built it, how to test design and operating effectiveness, and which indicators give early warning.

Specifying a control and enforcing one are different activities, and the gap between them is where governance is exposed. Written governance describes intent; production enforces behaviour. In between sits a policy requiring human oversight with no point in the workflow where a human can intervene, an approval requirement with no record that approval occurred, and a retention rule the pipeline never implemented.

The gap widens with autonomy. Agentic systems plan, call tools, and pass context between one another, so the decision path a control was written to constrain is assembled at runtime rather than at design time. Controls written for request-and-response do not bind a system that composes its own steps.

This is also where the data-side techniques earn their place, under the requirement they serve rather than as a framework of their own. Provenance and lineage are what turn a claim about where training and retrieval data came from into a lookup instead of an assertion. Event logging and traceability are what turn a decision path into evidence instead of an intention. A defined intervention point, with authority attached and a record that it was available, is what stops human oversight being an entry on an org chart. Bounding the action space a system may operate within — which tools it may call, which data it may reach, which decisions it may take unaided — is what keeps an agentic workload inside something anyone can reason about afterwards.

Which of those you are legally obliged to produce depends on where you operate and what the system does. For high-risk systems within the EU AI Act’s scope the Act names them directly: data governance in Article 10, record-keeping in Article 12, human oversight in Article 14, accuracy, robustness and cybersecurity in Article 15. Outside that scope none of those articles binds, and the same capabilities are still what ISO/IEC 42001’s data and lifecycle controls ask for, and still what any usable answer to a customer, an insurer, a regulator in another jurisdiction or a court will require. The technique does not change. Only the reason you can be compelled to produce it does.

What changes when you bought it rather than built it

Deployers control less of the stack and remain accountable for the outcome. The model is not yours to inspect, the training data is not yours to govern, and the vendor’s documentation is what you have to work with. That narrows the work. It does not remove it.

What stays in your hands is substantial: which decisions the system is allowed to touch, what data you feed it, where a human can intervene and whether that person has real authority to stop it, what you log and how long you keep it, and how you detect that behaviour has shifted. Those are configuration, integration, and process controls, and they are precisely what a deployer is asked to demonstrate.

02

Testing that the controls still work

Permalink to “Testing that the controls still work”

A control that was designed correctly and is never tested is an assumption. Assurance is what turns it back into a control.

There are two questions, not one. Design effectiveness asks whether this is the right control for the risk — whether a human review step placed after the decision has already taken effect can prevent anything at all. Operating effectiveness asks whether it is working now, at the volume it currently runs at, with the people currently doing it. A control can pass one and fail the other, and that gap is where an audit finding gets written.

Cadence comes from the risk tier rather than the calendar. A system informing credit decisions and a system drafting internal summaries do not warrant the same test frequency, and treating them alike is how an assurance budget gets spent on the wrong things.

Evidence is the third piece, and it decides whether an audit is routine or a project. If demonstrating that oversight happened means reconstructing it from memory and email, the control may well have operated, and it cannot be shown to have.

Controls drift too, and more quietly than models do. Permissions widen. An exception path gets added for a deadline and never removed. A threshold is relaxed after a noisy quarter and never restored. None of that surfaces in a policy review, because the policy still says exactly what it always said.

Assurance also has an allocation, and it is finer than the system-level one. Four roles per control, and the second is the one the independence of the whole rests on.

Per control — who does what
Operates

The team running the control day to day: the reviewer exercising oversight, the engineer maintaining the logging, the analyst watching a drift threshold. Named per control, not per department.

Tests

Someone other than the operator. A control tested by the people who run it is a self-assessment, which cannot supply the independence the test exists for. That independence is what makes the result worth having.

Reviews

Who receives the findings, decides what they mean, and owns a remediation date. A named owner and a date are what close a finding; without them it ages, and then it gets reported as a trend.

Escalates

Who is told when a control fails, within what period, and who holds the authority to suspend the system rather than log the failure and carry on. Decided in advance, because the moment it is needed nobody is reading a policy.

Control indicators answer whether each safeguard is operating as designed. They are the earliest signal in the chain: controls degrade first, exposure rises next, value is lost last. Three groups — whether the controls exist, whether they work, and whether they are keeping up with a portfolio that changes faster than the review cycle.

A · Are the controls in place?
  • —Share of production AI with complete logging and traceability, against the record-keeping obligations that apply.
  • —Human-oversight coverage: high-impact decisions with a defined and exercised override point.
  • —Share of systems carrying a named accountable individual, not a function, not a committee.
  • —Share of AI systems classified against the frameworks that apply to them.
  • —High-risk systems with a completed impact or conformity assessment.
B · Are they working?
  • —Control-test pass rate by control type, rather than in aggregate — an average hides the one control that never passes.
  • —Open findings against AI controls, and their age.
  • —Incidents where a control existed and did not operate. The most informative number in the set, and the one that requires somebody to have been looking.
  • —Exception volume, and how long exceptions persist past their stated expiry.
C · Are they keeping up?
  • —Time from a system entering the portfolio to its classification.
  • —Time from a purpose change to re-assessment — the EU AI Act Article 25 trigger that can turn a deployer into a provider, controlled rather than discovered.
  • —Share of controls tested within their required cadence.
  • —Share of use cases that cleared the pre-implementation gate before development began.

Group B takes the most work of the three, and tells you the most once it exists. Presence is straightforward to evidence and operation is not, which is why a control inventory and a control assurance programme are different things, and why an audit that samples operation reliably finds what a self-assessment did not.

A good place to start is one high-impact system. For each control around it, ask three questions. Where in the landscape is it enforced? When was it last tested, and by someone other than the operator? What evidence would show an auditor that it operated? Wherever an answer is missing, that's where to begin. None of this calls for a separate structure. Testing, findings and corrective action already run in a management system such as ISO 9001 or ISO/IEC 27001. Valment integrates the AI controls into that machinery, so one audit programme covers them.

Get direction and delivery working together.

Tell us what governance exists today and where AI actually runs. We will respond with a read on how the direction and the management system line up, and a sensible first step — often an ISO/IEC 42001 gap analysis.