AccuroAI
Products
What We Do
Solutions
Company
Resources
Book a demo
← Blog·AI Compliance10 min read

AI Governance Audit Preparation: What an Auditor Actually Asks For

An AI governance auditor does not tick your environment off against a control list. They take your own statement of applicability and test whether you did what it says, which means the highest-value week of preparation is spent on a document most teams write last.

S
Sofia Reyes
Head of Compliance
Sep 18, 2026

An AI governance auditor does not arrive with a list of controls and tick them off against your environment. They arrive with your own document and test whether you did what it says. Teams that understand this prepare in a completely different order from teams that do not, and the difference usually shows up in the first morning.

This piece is about what actually gets requested, in the sequence it gets requested, across ISO/IEC 42001, SOC 2 and the internal audits that increasingly carry AI scope. It is deliberately not a single-framework walkthrough. If you want the framework-by-framework mapping, we keep a crosswalk between NIST, ISO and the EU AI Act, and the ISO/IEC 42001 page covers that standard on its own terms.

What does an AI governance auditor actually ask for?

First, your scope statement. Second, the document that says which controls you decided apply. Third, evidence that those controls ran.

Everything else is a variation on the third item. The order matters because each step constrains the next: your scope decides what is in the audit, your control decisions decide what evidence exists, and the evidence decides whether you pass. Teams that start by collecting evidence are working backwards and usually discover in week five that half of what they gathered sits outside the scope they wrote.

The Statement of Applicability is the document the audit runs on

ISO/IEC 42001 defines the statement of applicability as:

"documentation of all necessary controls and justification for inclusion or exclusion of controls"

and adds a note that matters more than the definition:

"Organizations may not require all controls listed in Annex A or may even exceed the list in Annex A with additional controls established by the organization itself."

Read that carefully. The auditor is not measuring you against the standard's control list. They are measuring you against your own statement of applicability, including your exclusions and the reasons you gave for them.

This reframes preparation entirely. A well-written exclusion, justified and dated, is a pass. An unwritten exclusion is a finding. And an included control with no evidence behind it is a worse finding than an excluded one, because you declared it applicable and then did not run it.

The practical consequence: the highest-value week of preparation is spent on the statement of applicability, not on evidence collection. Most teams spend it the other way around.

What happens in Stage 1 versus Stage 2?

Certification against a management system standard runs as a two-stage initial audit followed by surveillance, on a cycle that is normally three years. Stage 1 is a readiness and documentation review. Stage 2 is the implementation audit, where the auditor looks for proof that the system described in Stage 1 is the one actually operating.

Stage 1 is where scope gets challenged. It is a documentation review, so the questions are about boundaries: what is in your AI management system, what is out, and can you defend the line. ISO/IEC 42001 requires that the scope be available as documented information, and requires the organization to determine its roles with respect to the AI systems it develops, provides or uses.

That last requirement catches people. An organization that both builds an internal model and buys three vendor tools holds different roles for each, and the scope statement has to say so.

Stage 2 is where sampling starts.

How do auditors sample, and what fails sampling?

Sampling is the part nobody prepares for, because a policy document reads as complete regardless of whether anyone followed it.

An auditor picks instances. Not "do you have an AI risk assessment process" but "show me the risk assessment for this system, and the three before it." The failure mode is consistent across every audit we have seen described: the process exists, the first two instances are immaculate because they were done during implementation, and instances four through nine were never done at all.

Three sampling targets come up repeatedly:

What gets sampledWhat the auditor is testingCommon failure
Individual AI system entries against the inventoryThat the inventory is current, not a snapshot from implementationSystems adopted after the inventory was built
Impact assessments for specific systemsThat the process ran for every system that triggered itThe trigger criteria were never written, so nothing formally triggered
Change records against deployed configurationThat changes went through the processEmergency changes with retrospective approval that was never filed

ISO/IEC 42001 defines the AI system impact assessment as a formal, documented process by which impacts on individuals, groups and societies are identified, evaluated and addressed. The word doing the work there is formal. An assessment that happened in a meeting and lives in someone's notes is not one.

Which artifacts should already exist before the evidence request arrives?

Six, and you can build all of them before anyone books an auditor.

A current AI system inventory, with an owner per system and your role for each. A statement of applicability with justified inclusions and exclusions. Impact assessments with written trigger criteria. Records showing human oversight happened. Change and incident records for AI systems. And evidence of competence and awareness for the people running all of it.

The inventory is the one that decays fastest, and it is also the one the audit opens with. An inventory maintained by quarterly survey is stale by the time it is read. This is the argument for deriving it from telemetry rather than from asking people, which is the same argument that drives the twelve fields an agent log needs.

If you want to know where you stand before committing to a date, scoring yourself against a control library first is cheaper than discovering it in Stage 1.

What evidence proves human oversight was real?

This is the hardest evidence category in AI governance, and the one auditors have become most pointed about.

A policy saying a human reviews high-impact outputs proves nothing. A screenshot of a review interface proves the interface exists. What satisfies an auditor is a record, per decision, showing that a named person saw the output, had the ability to change or reject it, and that some non-trivial fraction of the time they actually did.

The last clause is the test. An oversight log where the human approved one hundred percent of ten thousand decisions is evidence of a rubber stamp, and a competent auditor reads it that way. Override rate is the number that makes oversight credible, and almost nobody instruments for it in advance.

Ownership is the companion question. Auditors ask who signs, and the answer has to be a role that exists on an org chart, which is the practical reason governance roles need to be written down before the audit rather than assembled during it.

What do they ask about systems outside your scope?

Scope is a boundary you draw. It is not a claim that nothing exists on the other side of it, and auditors know the difference.

The question that lands hardest is some version of: how do you know what is in scope. Not what is in scope, but how you know. An organization that answers with a spreadsheet maintained by asking department heads has just told the auditor that its boundary is built on self-reporting, and self-reporting is exactly what shadow AI defeats.

Excluding a system is legitimate and common. Excluding a category of systems you have no way to detect is different, because you cannot justify excluding something you never knew about. This is where an AI management system audit diverges from an information security one: the population itself is uncertain, and the auditor is entitled to test how you bounded it.

What satisfies the question is a discovery method with coverage you can describe. If a tool was adopted by a team with a corporate card and a browser, can your process see it, and within how long. A first discovery run against a mid-sized estate commonly surfaces dozens of AI tools nobody had recorded, and finding them before Stage 1 is considerably cheaper than an auditor finding one after.

The practical form of the answer is two sentences in the scope statement: how the population is determined, and how often that determination is refreshed. Most scope statements have neither.

What nobody can tell you about common findings

Here is a thing worth saying plainly, because the internet is full of the opposite.

There is no published, finding-level data on ISO/IEC 42001 non-conformities. Accredited certification only began recently, the standard is not yet counted in the ISO Survey, and no accreditation body or standards body has published an aggregate. Every "most common ISO 42001 findings" listicle in circulation is a vendor or certification-body blog drawing on a handful of engagements, and several are visibly copying each other.

The nearest thing to a primary observation comes from the Standards Council of Canada, which wrote up its pilot accreditation of an AI management system certification body in September 2025. Its reported gaps were AI inventories, data lineage tracking, third-party risk management, and sector-specific overlays. That is one accreditation body describing one pilot, not a finding tally, and it should be read as exactly that.

Treat any article that gives you a ranked list of common findings with percentages as fiction until someone publishes the data.

Which edition was your certification body accredited against?

This is the question almost nobody asks, and it has become the sharpest one available.

The requirements for bodies that audit and certify AI management systems are set out in ISO/IEC 42006, published July 2025. It is additive to ISO/IEC 17021-1 rather than a replacement. It requires, among other things, that the audit team collectively hold knowledge of all the standard's Annex A controls and their implementation, and it bars a certification body from providing AI, security, data protection or risk management consulting to its own certification clients, including a clause preventing that restriction from being evaded by renaming the activity.

Two facts make this worth raising with any certification body you engage.

The first is timing. Certification against ISO/IEC 42001 was being issued before ISO/IEC 42006 published, which means some certification bodies were accredited against a draft. The Standards Council of Canada says as much about its own pilot. European co-operation for Accreditation made ISO/IEC 42006 mandatory for its members at its November 2025 General Assembly, which tells you the floor moved and when.

The second is structural. The International Accreditation Forum ceased operations on 1 January 2026, merging with ILAC into a single body. The mandatory documents the certification industry runs on now sit on a legacy archive, and there has never been a mandatory document covering AI management systems at all, despite one being issued for essentially every other new certification-body criteria standard.

None of that invalidates a certificate. It does mean that a certificate is a claim about a certification body as much as about you, and a customer doing real diligence will ask which edition and which accreditor. You should be able to answer before they do.

The six-week sequence

If you have six weeks and no date booked yet, run it in this order. Each week's output is the next week's input.

WeekDo thisOutput
1Write the scope statement and your role per AI systemDocumented scope, defensible boundary
2Build or refresh the inventory from telemetry, not a surveyCurrent inventory with named owners
3Write the statement of applicability, including exclusions and their justificationThe document the audit will run on
4Write impact assessment trigger criteria, then run the backlog they exposeAssessments that a sample cannot break
5Instrument oversight so override rate is measurableEvidence that review is real
6Sample yourself: pick five systems at random and pull every artifactYour own findings, before theirs

Week six is the one people skip and the one that pays. Pull the evidence for five randomly chosen systems exactly as an auditor would, with no warning to the owners, and count how many produce a complete set. If the answer is five, you are ready. If the answer is two, you have found your gaps while they are still cheap.

Where continuous evidence helps is in weeks two and five, because both are asking for a record that has to exist before the audit rather than be assembled for it, which is the case for reporting built on live telemetry rather than on a quarterly collection exercise.

One closing note on cost, since it shapes what you can verify yourself. The standard that defines the controls, the standard that defines what your auditor must be competent in, and the standard that defines the certification cycle are all paywalled, at roughly CHF 580 for the three. The free previews cover scope, terms and the early clauses, which is why the statement of applicability definition quoted above is one of the few load-bearing sentences anyone can check without paying. It is also, conveniently, the most important one.

See AccuroAI in action.
30-minute demo tailored to your top AI risk.
Book a demo
More from the blog
See AccuroAI in action.

Book a 30-minute demo and see how security teams use AccuroAI to discover, govern, and protect every AI asset across their organization.

Book a demoRun the free assessment

15 enterprises secured · under 38ms p99 · live on your own estate in 72 hours