Compliance2026-09-099 min read

The AI Act Audit Asks for a Technical File, Not a Policy

The transition period for Annex III high risk AI systems ended on 2 August 2026, and market surveillance authorities across the EU are reported to be running their first coordinated document requests. What gets asked for is the Annex IV technical file: architecture, training data provenance, accuracy metrics, human oversight design and post market monitoring, written per system. Most teams who believe they did their AI Act work built a policy set instead, and the two are not interchangeable.

The AI Act stopped being a planning exercise on 2 August 2026. That was the applicability date for the Annex III high risk categories, the ones covering employment and worker management, access to essential private and public services including credit scoring, education, biometrics, critical infrastructure and law enforcement. Since then the AI governance trackers have described a first coordinated wave of document requests, with the European AI Office and national market surveillance authorities including CNIL in France, the BfDI in Germany and AESIA in Spain concentrating on three obvious populations: automated resume screening in recruitment, algorithmic credit assessment in retail banking, and triage tools in private healthcare. Treat that sector list as reporting rather than as something to plan around, because it may or may not describe where the second wave goes. Treat the underlying change as settled, because it is written into the regulation rather than into a news cycle: the transition period is over, and Article 21 entitles a competent authority to make a reasoned request for the documentation that demonstrates conformity, and to be given it.

The thing they ask for has a name and a defined shape, and this is where a lot of programmes discover a gap. Article 11 requires a provider of a high risk system to draw up technical documentation before the system goes on the market and to keep it current afterwards. Annex IV sets out what has to be in it, in nine parts: a general description of the system, its intended purpose and its versions; a detailed account of how it was developed, covering architecture, computational resources and the training data; information on capabilities, limitations, accuracy and foreseeable risks to health, safety and fundamental rights; a justification for why the chosen performance metrics are the appropriate ones; the risk management system required by Article 9; a record of changes made across the lifecycle; the harmonised standards applied, or an explanation of how conformity was reached without them; a copy of the EU declaration of conformity; and the post market monitoring plan required by Article 72. Article 18 then requires that file be kept for ten years after the system is placed on the market.

Read that list next to what most organisations actually built during 2025 and 2026 and the mismatch is obvious. The work that got done was governance work: an AI policy, an acceptable use standard, a model inventory spreadsheet, a risk assessment template, a steering committee with a charter, perhaps an ISO 42001 gap analysis. All of that is genuinely useful and none of it answers a request for Annex IV. A policy says how decisions ought to be made. A technical file is the evidence that one particular system was built, tested and monitored a particular way, and it is written per system rather than per organisation. If you have twelve AI features in production and one AI policy, you have roughly one twelfth of a filing cabinet. The distinction matters most at the moment of the request, because an authority that asks for the technical documentation of a named system and receives a policy set has learned something about you that you did not intend to tell it.

Three sections cause more trouble than the rest, and they are the ones worth checking first. The training data description in part two requires provenance, scope, characteristics, how the data was obtained and selected, labelling procedures and cleaning methods, which is difficult for anyone who fine tuned on a scraped corpus and impossible for anyone who cannot say where that corpus came from. The accuracy material in part three requires your metrics, and part four then requires a justification for why those metrics are appropriate for this system and the people it affects, which is a question most teams have never been asked and cannot answer out of a dashboard. The human oversight material requires a description of the measures built in so a person can understand, override and interrupt the system, and the reporting suggests auditors are asking for the schematic rather than the sentence. Any of the three can be produced retrospectively at cost. All three are far cheaper to capture while the system is being built, which is the argument for starting on new work now instead of waiting for a request.

The file is also not a document you finish. Part six asks for changes made across the lifecycle and part nine asks for the post market monitoring plan and what it found, which together mean the technical file has to move whenever the system moves. That is a hard requirement to meet when any part of your system is a third party model, because the model changes on a schedule you do not control. We have covered the churn repeatedly this year: an assistants API retired, model rosters rotated underneath coding tools, introductory pricing withdrawn, weights swapped for a newer generation with measurably different behaviour. Every one of those is a lifecycle change to any high risk system that depends on it, and every one needs a dated entry, a retest against your stated accuracy figures, and an updated declaration if the behaviour moved. A provider who cannot say which model version was serving production traffic in March cannot honestly complete part six, and that is a records problem rather than a compliance opinion.

The next trap catches companies who assumed none of this reaches them. The obligations above land on the provider, and many teams file themselves under deployer on the grounds that they bought the model from somebody else. The Act does not work that way. Put your own name or trademark on a high risk system, make a substantial modification to one, or change the intended purpose of a general purpose system so that it becomes high risk, and you become the provider with the full Article 11 obligation attached. That is precisely the shape of a great deal of current building. A recruitment screening feature assembled on top of ChatGPT or Claude, a credit pre assessment flow scaffolded in Cursor or Devin, an internal HR tool built in Bolt, v0 or Lovable, a candidate interview product running ElevenLabs voices: each of those is somebody putting their own name on a system that does something Annex III lists. The model vendor will publish documentation for its own obligations, and that documentation is an input to your technical file rather than a replacement for it. It is the same argument we made about platform certifications: a vendor report covers the vendor.

None of this makes your existing certification work wasted, and it helps to be precise about what carries over. ISO 42001 gives you the management system, the roles and the impact assessment discipline, and it maps closely onto Articles 9, 10, 12, 14 and 17, which is a large part of why it is worth having. What it does not do is produce the per system evidence, and as we noted when EN 18286 stalled, it does not currently hand you a presumption of conformity either. ISO 27001 covers the security half of the picture, and SOC 2 covers whether your controls operated over a period rather than on the day you looked, which feeds parts five and nine usefully. The compliance automation platforms we cover, Vanta, Drata, Secureframe, Sprinto, Thoropass and Hyperproof, have all shipped AI governance modules, and they will hold the inventory, drive the risk assessments and collect evidence on a schedule. What none of them will do is write your architecture description, source your training data provenance or justify your accuracy metrics, because those are engineering artefacts only the team that built the system can produce. Buying a platform and expecting a technical file to fall out of it is the specific mistake to avoid this quarter.

The work is bounded and it is much cheaper before somebody asks. List every AI system you operate and mark each one provider or deployer, being honest about the ones where you put your own name on the output. For anything landing in an Annex III category, open a file per system with the nine Annex IV headings as empty sections and fill in what you already have, because the gaps become visible in an afternoon and that is worth more than another gap analysis. Take the three hard sections first, training data provenance, accuracy metrics with a written justification, and the human oversight design, and assign each to the engineer who knows the answer rather than to the compliance function. Add a rule that any change to an underlying model version is logged as a lifecycle change with a date and a retest result, and wire it into the release process you already run rather than into a new one. Confirm you could actually retrieve all of it on a deadline set by somebody else, rather than on the timeline it takes when the material is scattered across six different teams. And if you genuinely have no Annex III system, write that determination down with the reasoning behind it, because being able to show why you concluded you are out of scope is the cheapest piece of evidence you will ever produce.

EU AI ActAnnex IVtechnical documentationhigh-risk AImarket surveillanceISO 42001ISO 27001SOC 2AI governanceVantaDrata

Editorial note: AES Tech reviews are independent. Some outbound links are affiliate links and are marked sponsored; they never change our rankings. See our disclosure.

// Signal, not noise

Get the next post by email

One short email when something worth knowing ships. No spam, unsubscribe anytime.

Loading comments...

Add a comment

Corrections and first-hand experience are the most useful things you can leave. Comments are screened automatically and reviewed by a human; see the moderation policy.

0/4000 · plain text · links are held for review

More from the blog