The AI Decisions You’re Already Making
Part 3: What you'd actually find if someone asked you to prove it
In the first piece in this series I looked at why a business is actually pursuing AI and who's authorised to say yes before a new use case goes live. In the second, I looked at where the risk tends to hide once those conversations have started, usually somewhere other than where people are looking.
This last piece is about proof. Not what you intend to put in place, but what exists today, and whether you'd know if something had already gone wrong.
Could you show your working, not just state the outcome?
The decisions that get documented are usually the ones that feel significant at the time. A new AI use case going in front of a committee, a model being signed off through a formal review. These tend to leave a paper trail, because the process that generated them was designed to produce one.
The gaps are rarely there. They're in the decisions that didn't feel like decisions when they were made: someone connecting two systems together that hadn't been connected before, someone enabling an AI feature that arrived embedded in an existing tool, someone building something and making it accessible without anyone else being involved in the call that it was ready. None of these feel like significant choices at the time, and some of them turn out to be significant later. The problem isn't that the thinking didn't happen. It's that it often lives nowhere except in the memory of whoever was in the room, or in a message thread that nobody would know to look for.
If a regulator, a customer, or a new hire asked "why was this considered acceptable," the answer depends on that person still being around, still remembering it the same way, and still being willing to reconstruct it under pressure. None of those are safe assumptions six months later.
This is a smaller ask than it sounds. It doesn't mean a compliance framework for every tool anyone touches. It means the handful of decisions that actually matter, the ones with a real answer to "what could go wrong here," get written down at the time, in enough detail that someone else could read it cold and understand why you did what you did. Not a policy document nobody reads, but a short, dated record of the actual reasoning behind the choices that mattered.
If your honest answer is "we'd have to piece that together from memory," you're describing the default state, not a failure. The question worth asking is what happens the first time someone external asks for it and the memory isn't there, or isn't consistent.
Would you know if something had already gone wrong?
This is the question that gets skipped most often, and it's harder than the previous one in a specific way: the answer depends on who built the thing, not on how consequential it actually is.
AI that goes through a formal engineering process (models developed by a data science team, agents built as part of a technical platform) tends to come with automated tests, benchmarks, and some form of continuous evaluation. That's not because anyone made a particular effort on governance; it's because it goes through a development process that includes those things by default.
Operational AI is different. A tool helping someone draft a recommendation, an assistant supporting a decision about a customer: these typically rely on a human in the loop to check the output and remain accountable for the end result. That sounds like governance, and in a sense it is, but it rests on an assumption that deserves more scrutiny than it usually gets: that the person checking maintains the right level of scepticism, consistently, under normal working conditions.
That assumption is fragile in a way that doesn't announce itself. It doesn't fail visibly; it degrades gradually, as familiarity makes outputs that look normal feel normal. The failure mode isn't a system going down or an error surfacing in a log. It's a decision that seemed fine at the time, accepted by someone who had stopped scrutinising as carefully as they once did, discovered later by someone asking why it doesn't add up.
You don't need to monitor everything. You do need an honest view of which AI-assisted decisions would cause real harm if they were wrong for a sustained period without anyone noticing, and whether anything in your current setup would catch that before someone outside the business does.
Across these three pieces, the pattern is the same. Not one big gap, but several reasonable decisions, each one sensible when it was made, that nobody's gone back to check still add up. Most businesses aren't behind on this. They're roughly where everyone else is: further along than they'd guess on some of it, and further behind than they'd like on the rest.
If you're fairly sure which category your own business falls into, that's a good sign. If these pieces have left you less certain than you'd like to be, that's usually the more useful place to start from. If you'd like to find out where you actually stand, rather than where you assume you do, get in touch - an AI Sense Check is designed for exactly this.
Ed Ball is Head of Data and Security at a regulated UK lender, and the founder of Sapien Solutions.
Part 1: Why you're doing this, and who's actually deciding
Part 2: The risk isn't where you're looking