“When was AI last audited?”
It is being asked in audit committees across the country, and the honest answer at most firms is that it hasn’t been — not because internal audit is avoiding it, but because nobody is quite sure what the audit is. There’s no widely adopted program, no standard work paper set, and a real risk of producing something that either restates the technology audit you already do or drifts into a data science review the function can’t defend.
Having built control frameworks that auditors tested, directed design-effectiveness testing across an enterprise control library, and spent years on both sides of the finding, I’d offer this: an AI audit is far more familiar than it sounds. It is a governance and controls audit. The subject matter is unusual; the discipline is not.
Here’s what it actually contains.
First, what it is not
It is not an assessment of whether the models are any good. Internal audit is not the right function to opine on whether a gradient boosting model was the correct architectural choice, and a function that tries will be out-argued by the team it’s auditing and rightly ignored by management.
It is also not a repeat of your IT general controls audit. Access, change and operations controls over the platforms AI runs on matter, and you likely already cover them. If your AI audit report reads like an ITGC report with “AI” substituted in, you have audited the infrastructure and missed the risk.
The risk you are auditing is this: an output that materially influences a decision, produced by a system whose behavior nobody has established, monitored or owned. Everything you scope should trace back to that sentence.
The scoping problem, which is the whole first phase
Almost every AI audit stalls in the same place: nobody can tell you what’s in scope, because nobody knows what AI the firm is actually using.
That’s not evasion. It’s genuinely hard. AI arrives in three ways and only one of them is visible. There’s what the firm built or configured deliberately. There’s what arrived inside software the firm already licensed — the vendor enabled a feature, and it now summarizes, ranks, or drafts something. And there’s what staff adopted independently.
So the first phase is discovery, and it is not a request to IT for a list. In practice it means: reviewing the software inventory and asking which products have added AI capability in the last two years; reading vendor release notes and contract amendments; interviewing process owners about anything automated that produces a recommendation, score, ranking or draft; checking expense claims and corporate card statements for tool subscriptions; and asking staff directly, in a way that doesn’t sound like an investigation.
Two things worth knowing before you start. First, if the firm has no AI inventory, that is itself your first finding, and often the most important one — because everything else in the framework is unenforceable without it. Second, the discovery work you do will effectively become the firm’s first inventory. Hand it over deliberately, with a recommendation that management own and maintain it, rather than letting it live only in your work papers.
The four layers you test
Once scope exists, the audit has a natural shape. Four layers, each with its own testing approach.
Layer one: governance. Is there a policy? Was it approved by someone with the authority to approve it, and when? Is there a named owner for AI governance overall, with the standing to say no? Is there a route by which a new use case gets assessed, and does the evidence show it was used — or was it bypassed for the most significant deployment in the firm?
Testing here is straightforward document and interview work. The revealing test is always the same: take the firm’s most consequential AI use and trace it backwards through the governance process. Policies look excellent until you follow one real case through them.
Layer two: inventory and tiering. Does the inventory exist, is it complete, and does it record what a governance process needs — purpose, data consumed, decisions influenced, owner, risk tier? Is the tiering criteria written down and applied consistently? Test completeness against your own discovery work, and test tiering by re-performing it on a sample: pick five entries, apply the firm’s own criteria independently, and see whether you land where they did. Where you don’t, the disagreement is the finding.
Layer three: lifecycle controls, proportionate to tier. For higher-tier uses: was there review before deployment by someone independent of the build or the buying decision? Is there documentation of intended use, data, limitations and known failure modes? Are changes — including a vendor retraining a model underneath you — controlled and communicated? Is there a defined point at which the system is retired or re-approved?
For vendor-supplied AI, this layer connects directly to third-party risk. What did due diligence establish about the vendor’s own model governance? Was the assessment refreshed when the vendor materially changed the product? If the answer is that a procurement questionnaire from three years ago is the only artifact, that’s a finding that belongs to the third-party risk process as much as to AI.
Layer four: monitoring and use in practice. Are outputs monitored for performance and drift, with thresholds that trigger something? Is there evidence anyone reviewed the monitoring? Where the policy requires human review before an output is used, does that review actually happen — and is it substantive?
That last test is the one that most often produces the finding that matters. Sample real outputs and trace them to the decision they informed. A “human in the loop” who approves forty recommendations in nine minutes is not a control, and the evidence for that is in the timestamps.
What evidence actually looks like
The most common failure in a first AI audit is accepting assurance in place of evidence. Concretely, what you should be collecting:
Approved policy documents with dates and approvers. The inventory as a maintained artifact, not a spreadsheet assembled the week you asked. Assessment records for individual use cases. Pre-deployment review documentation with an identifiable reviewer independent of the build. Model or system documentation stating intended use and limitations. Change records, including vendor change notifications. Monitoring output plus evidence of review — minutes, sign-offs, tickets. Timestamped samples of human review where the policy requires it. Third-party assessment records for vendor-supplied AI. Incident records, if any.
That final item deserves attention. Ask specifically whether anything has gone wrong: an output that was wrong in a way that reached a client, a tool used with data it shouldn’t have been. A firm that reports no incidents at all in two years of AI adoption has either exceptional controls or no reporting route, and it is worth establishing which.
Writing findings people act on
Findings in this space fail in two directions. Too technical and management can’t act. Too abstract and they can’t disagree, which sounds like agreement but isn’t.
The pattern that works is the one that works everywhere: state the condition specifically, name the risk in business terms, identify the root cause rather than the symptom, and rate by consequence rather than by how easy the gap was to spot.
“No AI governance framework exists” is accurate and useless — it’s unrateable, and the remediation is a year of work nobody will commit to. Compare: “Twelve AI-enabled capabilities were identified across four business processes, of which three influence client-facing decisions. None have documented pre-deployment review, and no inventory exists, so management cannot currently identify which processes would be affected if a vendor changed a model.” That’s specific, traceable, rated by consequence, and the remediation is obvious.
One more thing about rating. Resist inflating findings because the subject is topical. AI attracts attention, and a function that reports everything as high on a fashionable topic loses credibility on the ones that genuinely are. Rate by consequence, exactly as you would anywhere else.
Who can actually do this work
Honestly: an experienced auditor who understands governance, controls and risk can perform layers one, two and four with a modest amount of preparation. The frameworks translate — this is COBIT and COSO thinking applied to a new object.
Layer three, for the highest-tier uses, is where specialist input matters. Assessing whether a validation was meaningful, whether the documented limitations are plausible, and whether monitoring metrics are the right ones requires someone who understands how models fail. That’s the piece small internal audit functions typically co-source, and it’s a defensible use of external specialists — your function retains ownership, methodology and sign-off, and buys depth for one layer of one audit.
If you’re building capability internally, ISACA’s AAIA credential now sits on top of CISA for exactly this work, and the professional standards material is maturing quickly.
What the audit committee should ask for
If you sit on an audit committee and want to know whether the AI coverage you asked for was real, four questions distinguish substantive work from a paper exercise:
How many AI-enabled systems did the audit identify, and how many did management know about before it started? The gap between those numbers is the most informative statistic in the report.
Which of them influence decisions affecting clients? If nobody can answer, the inventory isn’t fit for governance.
Did we test whether the controls operate, or only whether they exist? Design-only testing on a first audit is a legitimate choice. It should be stated, not implied.
What did we find when we traced a real output to a real decision? If the audit never did this, it stayed in the documentation and didn’t reach the risk.
Where to start
If AI coverage is on your plan this year and you don’t know where to begin: start with discovery and scope the first engagement narrowly. A well-executed inventory-and-governance audit — layers one and two, done properly — is more valuable than an ambitious full-scope review that runs out of time in fieldwork and reports on documentation alone.
The second-year audit, with an inventory to work from, is where the depth becomes possible. Sequence it that way deliberately and say so in the plan.
If your audit plan has AI on it and your function doesn’t yet have the specialist depth for layer three, that’s precisely the work we co-source — under your methodology, with your function retaining ownership and sign-off.