Annex A of ISO/IEC 42001 contains 38 controls across nine objectives, numbered A.2 through A.10. The standard tells you what each control is for. It does not tell you which document, log or screenshot proves you did it, and that gap is where most first-time certifications lose six weeks.
This piece is the control-by-control answer. For each of the 38, the artifact an auditor will accept, and whether you write that artifact by hand or pull it from a system you already run. If you want the shape of the audit itself, the stages, the sampling, and what gets you a nonconformity, that is a separate piece on audit preparation. Read this one when you are filling in the Statement of Applicability and need to know what goes in the evidence column.
One caveat before the tables. Control titles below are references, not the standard's text. ISO/IEC 42001 is copyrighted and you need a licensed copy to implement it properly. Buy the standard.
Why the evidence column is the hard part
An AI management system certification works the way an ISO 27001 certification works. You declare a scope, you produce a Statement of Applicability listing every Annex A control with an applicable or excluded decision and a justification, and then the auditor tests whether the applicable ones actually operate. Stage 1 reviews the documents. Stage 2 asks for proof.
The failure mode is predictable. Teams write excellent policy for A.2 and A.3, because policy is a writing exercise and security teams are good at it. Then they reach A.6 and A.7, where the evidence is operational, and discover nobody logged anything. A policy saying event logs are retained is not evidence. The logs are evidence.
So the useful way to read Annex A is by evidence type, not by control number. Three types exist.
Written artifacts are documents a human authors once and reviews on a schedule: policies, role definitions, impact assessment templates, supplier terms. Roughly half of Annex A is this.
Operational records are generated continuously by systems in use: event logs, monitoring output, access records, incident tickets. These cannot be back-filled. If you are not capturing them today, you cannot produce twelve months of them in March.
Decision trails sit between the two. An impact assessment is a written artifact, but the record of who approved it and when is operational. Auditors sample both halves, and the second half is where enterprises get surprised.
A.2 to A.5: the governance layer
Fourteen controls, and almost all of them are written artifacts. This is the part you can finish in a quarter with no new tooling.
| Control | Reference title | Evidence that satisfies it | Type |
|---|---|---|---|
| A.2.2 | AI policy | Approved AI policy, with the approval record and version history | Written |
| A.2.3 | Alignment with other organizational policies | A mapping showing where the AI policy touches security, privacy and HR policy, and how conflicts resolve | Written |
| A.2.4 | Review of the AI policy | Dated review minutes, plus the next scheduled review | Decision trail |
| A.3.2 | AI roles and responsibilities | A role matrix naming individuals, not job titles, against each AI governance duty | Written |
| A.3.3 | Reporting of concerns | The channel, and at least one worked example of a concern raised and closed | Decision trail |
| A.4.2 | Resource documentation | An inventory of the resources your AI systems depend on | Written |
| A.4.3 | Data resources | Which datasets are in scope, where they live, who owns them | Written |
| A.4.4 | Tooling resources | The tool inventory, including the AI tools your staff use but procurement never bought | Operational |
| A.4.5 | System and computing resources | Compute and hosting documentation for in-scope systems | Written |
| A.4.6 | Human resources | Competence records: who is trained, on what, when | Decision trail |
| A.5.2 | AI system impact assessment process | The documented process itself | Written |
| A.5.3 | Documentation of impact assessments | Completed assessments for in-scope systems, not a blank template | Decision trail |
| A.5.4 | Assessing impacts on individuals or groups | The section of each assessment covering affected persons, with findings | Decision trail |
| A.5.5 | Assessing societal impacts | The societal section, and evidence it was considered rather than deleted | Decision trail |
Two of these are quietly harder than they look.
A.4.4 asks for tooling resources. If your scope statement covers employee use of AI systems, then the AI tools employees actually use are in scope, and an inventory built from procurement records will be incomplete. Every enterprise that has run a discovery exercise has found tools nobody bought. That is the one control in this block where a written artifact is not enough, because the honest answer changes weekly.
A.3.2 fails for a different reason. Role matrices that name job titles pass Stage 1 and fail Stage 2, because the auditor asks the named person what their AI duties are, and the named person has not heard of the matrix.
A.6: the AI system life cycle, and the nine controls that need logs
This is the largest objective and the one that decides your timeline. Nine controls, and the back half of them cannot be written retrospectively.
| Control | Reference title | Evidence that satisfies it | Type |
|---|---|---|---|
| A.6.1.2 | Objectives for responsible development | Stated development objectives, approved | Written |
| A.6.1.3 | Processes for responsible design and development | The documented development process, with gates | Written |
| A.6.2.2 | AI system requirements and specification | Requirements documents per in-scope system | Written |
| A.6.2.3 | Documentation of design and development | Design records, including decisions rejected | Written |
| A.6.2.4 | Verification and validation | Test plans and results, including tests that failed and what changed | Decision trail |
| A.6.2.5 | AI system deployment | Deployment records and approvals, per release | Operational |
| A.6.2.6 | Operation and monitoring | Monitoring output over a period, with the alerts it raised | Operational |
| A.6.2.7 | Technical documentation | The technical file per system, current as of audit | Written |
| A.6.2.8 | Recording of event logs | Actual event logs, retained, tamper-evident, covering the audit period | Operational |
A.6.2.8 is the control that sets your certification date. You need event logs covering a period the auditor considers representative, and the clock starts when logging starts, not when the project starts. A team that begins its ISO 42001 program in January and turns on logging in May is not ready in June, no matter how good the documents are.
If you buy rather than build AI systems, read A.6 with care. Several of these controls assume you develop the system, and for a deployer they are candidates for exclusion in the Statement of Applicability. Exclusion is allowed. Unjustified exclusion is a finding. Write the justification in the same sentence as the exclusion, and make it about your scope rather than about your effort.
A.6.2.6 and A.6.2.8 are rarely excludable even for pure deployers, because you still operate the system and someone still has to watch it. This pair is the operational core of the whole Annex, and it is also the pair that maps most directly onto tooling you may already own. Log pipelines, monitoring and alert routing are not new disciplines. What is new is that the events are prompts, tool calls and model outputs rather than network flows, and most existing pipelines were never pointed at them. We covered the plumbing in AI audit trails and SIEM integration.
A.7 to A.10: data, disclosure, use and third parties
| Control | Reference title | Evidence that satisfies it | Type |
|---|---|---|---|
| A.7.2 | Data for development and enhancement | Which data is used for development, and on what basis | Written |
| A.7.3 | Acquisition of data | Source records and acquisition terms | Written |
| A.7.4 | Quality of data | Quality criteria, and the results of applying them | Decision trail |
| A.7.5 | Data provenance | Provenance records traceable per dataset | Written |
| A.7.6 | Data preparation | Preparation steps documented and repeatable | Written |
| A.8.2 | System documentation and information for users | What users are told, as shipped | Written |
| A.8.3 | External reporting | The reporting channel and anything reported through it | Decision trail |
| A.8.4 | Communication of incidents | Incident records with notification timestamps | Operational |
| A.8.5 | Information for interested parties | A record of what was disclosed, to whom, when | Decision trail |
| A.9.2 | Processes for responsible use | The acceptable-use process as operated | Written |
| A.9.3 | Objectives for responsible use | Stated use objectives, approved | Written |
| A.9.4 | Intended use of the AI system | Intended use statements, and evidence of use staying inside them | Operational |
| A.10.2 | Allocation of responsibilities | Who owns what across you and your vendors | Written |
| A.10.3 | Suppliers | Supplier assessments and contract terms covering AI | Decision trail |
| A.10.4 | Customers | What you commit to customers, and how you honor it | Written |
A.9.4 deserves attention. It asks for intended use, which sounds like a written statement, and most teams supply one. The auditor then asks how you know use stayed inside it. That second question is operational, and answering it means you can see what people actually did with the system. A policy forbidding sensitive data in a public model is a written artifact. Evidence that nobody sent any is a different thing entirely, and it comes from the point of use.
A.10.3 is the control most likely to stall. Supplier assessments for AI vendors need to cover training data, retention, subprocessors and model changes, and standard security questionnaires do not ask about any of those. If you need a starting point, our AI vendor security questionnaire covers the AI-specific questions, and the reasoning behind each one is on the blog.
What you can produce automatically, and what you cannot
Counting by evidence type across all 38 controls: roughly 20 are written artifacts, 11 are decision trails, and 7 are operational records. The written half is a documentation project with a known end date. The operational third is a capability question, and the answer is either yes or it is a six-month delay.
The seven operational controls are A.4.4, A.6.2.5, A.6.2.6, A.6.2.8, A.8.4 and A.9.4, plus whichever decision trails your auditor decides to sample from live systems rather than from your binder. If you can answer those from telemetry you already collect, your certification timeline is set by the writing. If you cannot, it is set by how long you need to run before you have a representative period of logs.
This is the argument for starting the logging before starting the paperwork, which is the reverse of how most programs sequence it.
Using this with the other frameworks you already map
Almost nobody implements ISO 42001 alone. The same control evidence usually has to serve NIST AI RMF and, for anyone with EU exposure, the AI Act. The mapping is imperfect and the overlaps are real, and we have worked through where they line up and where they genuinely diverge in the unified compliance crosswalk. The short version is that the Annex A operational controls do most of the double duty, because evidence of what actually happened is framework-agnostic. Policy language is not.
AccuroAI maps enforcement decisions to 8 frameworks and customers report 11 times faster audit preparation, and the reason is narrow rather than magical: the operational third of Annex A is the expensive part to assemble by hand, and it is the part that comes out of a system of record rather than a document library.
The sequence that works
Define scope first, and define it narrowly enough that A.6 exclusions are defensible. Then turn on logging for every in-scope system, because that clock is the long one. Then write, in this order: policy and roles, the impact assessment process, supplier terms, and technical files. Run an internal audit before Stage 1 and find your own nonconformities, because the ones you find are free.
If you want a view of where your program stands before committing to a certification date, our AI governance maturity assessment walks the same ground in about ten minutes and tells you which of the three evidence types you are weakest on. Teams that already know the answer is logs can skip it and go turn on logging.