Skip to main content

ai governance

the accountability half of the job

regulation and case law are closing the gap between running a model and owning what it produces. the enablement budget usually covered only the first half.

most orgs we begin to work with have learned to operate ai in some capacity. almost none of them have decided who is accountable for what the ai produces. that gap is most of what happened in enterprise ai over the last 2 years, and optional hygiene is no longer a safe description of it.

mit's nanda initiative1, in a 2025 report called "the genai divide," found that about 95% of enterprise generative-ai pilots produced no measurable impact on the bottom line. the report took criticism over its method, so hold the figure loosely. the direction is harder to wave away: most of this spending returned nothing anyone could point to on a p&l.

the comfortable explanation is that the tools were not ready, or the teams were not skilled enough. in the regulated builds we have run, the model was rarely the point of failure. a capability went live and no single person was accountable for the decisions it started making. operating the system had an owner; standing behind its output did not.

you can watch the same gap reach a courtroom. in 2024, air canada's website chatbot2 gave a passenger wrong advice about bereavement fares. when the case landed at a small-claims tribunal, the airline argued the chatbot was responsible for its own actions. the tribunal rejected that and made the airline pay. that defense can only come out of a company where no person had been named as the owner of what the tool says.

part of the reason it stays unowned is simple mechanics. a model spreads responsibility thin by default. when a person writes the first draft, they own it. when a model writes it and a person edits it, ownership goes fuzzy in exactly the way that lets everyone assume someone else has it. that fuzziness holds until something goes wrong, and then nobody can answer a simple question about who decided.

the other reason is less comfortable. owning a decision a model made means being able to tell whether the model got it right. people were taught to operate the tool. few programs checked whether they could judge what came out of it. that takes knowing the process, the standards, and what good work looks like. it takes being able to catch a confident answer that happens to be wrong. a name against a decision, with no ability to evaluate it, is just a name on whatever the model produced.

for a while you could treat that as internal hygiene and get away with it. that window is closing, in writing. the eu ai act3 entered into force in august 2024 and has been phasing in since: bans on certain uses and ai-literacy duties in february 2025, obligations for general-purpose models that august. the heaviest requirements, for high-risk systems, were set for 2026 and are now sliding toward late 2027 under the commission's "digital omnibus" package. california moved too. sb 534, signed in september 2025, took effect in january 2026. it requires the largest model developers to publish a safety framework and report serious incidents.

underneath the differing details and the moving dates, the expectation is the same: a named party is responsible for what a model does, and can be asked to show their work.

that is the half of the job the enablement budget usually skipped. people were taught to prompt. accountability for the output was left unnamed.

we spent the summer on the companion problem in our lo & slo series: an ai stack that speeds the work while wearing the worker down. accountability is the other half of the same problem. if a person is still answering for the output, that person has to be rested enough to catch what is wrong with it.

if you run one exercise this quarter, run this one. write down the decisions your team already hands to a model. put a single person's name against each, a person and not a committee. then, for each, write down what would make you stop trusting it. it takes an afternoon, none of it is fun, and it is most of what separated the pilots that mattered from the 95% that did not.