EU AI Act Annex III: Your Workflow Decides the Risk
The model name does not decide high-risk status. Intended use, workflow influence and procurement changes…
The short version
- EU AI Act Annex III classifies complete systems by intended use and influence over consequential decisions.
- Recruitment audits can test vocabulary, candidate personas, demographic-label effects, and adverse impact before deployment.
- Companies have until 2 December 2027, but model updates and workflow changes require repeatable evidence.
At noon, the same Mistral model drafts a rejection email. Five minutes later, it helps decide who gets rejected. The second workflow should make your lawyer spill their espresso.
EU AI Act Annex III classification follows the complete system’s intended use and influence over consequential decisions.
Model logos, parameter counts and national origin do not decide it. Annex III lists eight stand-alone high-risk categories: the complete route, separate from the product-safety route in Annex I.
Europe got the principle right. Automation that changes somebody’s life should have to prove itself.
Now comes the engineering. Dio mio, spare me another PDF pilgrimage.
start with the decision
My first question is simple: what does this system decide or influence? “Our company uses AI” tells me nothing. The bakery near my apartment uses AI and still cannot update its opening hours online. “This system recommends which applicants receive an interview” immediately raises an Annex III recruitment question. Regulation AI’s analysis draws the same boundary: the categories cover specific purposes, not every company in a named industry. A law firm’s research assistant stays ordinary until its intended purpose enters covered work, such as helping a judicial authority apply law to concrete facts. SafeLegalAI similarly notes that workplace emotion recognition generally cannot be rescued by human oversight and a tasteful dashboard.
Procurement needs to stop worshipping the model name. I identify the input and output, then follow that output to its user. Does it rank someone, recommend an outcome, exclude a person or determine access to something consequential? I compare that function with Annex III and check whether it profiles a natural person. Then I document who controls the intended purpose and any post-purchase configuration that changed it. Finally, I classify the full workflow, including the human decision consuming the output. That separates software drafting a job advertisement from software filtering applicants, even if both call the same API.
Article 6 offers a limited exception when a listed system creates no significant risk. It can cover a narrow preparatory task, pattern detection that does not improperly replace human review, or improvement of work already completed by a person. The provider must document this conclusion. Regulation AI notes that profiling a natural person closes the route, so calling candidate scoring “administrative assistance” requires more than a creative contract title.
Procurement can also shift provider roles. Under Article 25, a customer can move from deployer to provider by rebranding a high-risk system, substantially modifying it or changing a lower-risk product’s intended purpose until it becomes high-risk. Connecting an assistant to applicant records does not settle the issue. Configuring it to recommend interview candidates changes the workflow, risk and potentially the customer’s legal role.
I once wasted an indecent evening on a vendor security questionnaire that answered every encryption question beautifully but never explained what the product decided. Now I ask for the intended-purpose sentence before discussing contract length. Beige branding has fooled stronger men than me.
recruitment bias hides inside the workflow
The model supplies general capabilities. Prompts and retrieval sources make it a system; deployment connects that system to organisational data and a decision about a person. Annex III classification must follow every layer. A Mistral AI vs Claude comparison can help product selection, but the research available here contains no direct benchmark crowning either model for Annex III recruitment. I will not invent one from benchmark vibes. We already have fantasy football for LLMs.
Kitahara and Yamaguchi audited six open-weight models in controlled recruitment experiments: Mistral, Llama, Gemma, Qwen, Phi and DeepSeek. Their protocol first scores job-posting vocabulary, especially agentic or communal wording. The researchers then give candidate personas to the models while holding the recruiter task and surrounding prompt conditions steady, comparing recommendation scores and expressed interest across personas. When differences appear, label ablation removes the explicit demographic label to test whether it caused them. Agentic wording was associated with lower recruiter recommendation scores for female personas; communal wording partly reversed the pattern. Ablation identified explicit demographic labels as the primary driver of the reported bias effects.
That causal chain matters. Posting language changed the context, the persona label added a demographic signal, and the task converted both into a recruitment recommendation. A broad model leaderboard misses this interaction because it tests general capability outside the hiring workflow. Kitahara and Yamaguchi therefore propose pre-deployment evidence based on vocabulary scoring and persona-conditioned probing. Adverse-impact flagging would capture resulting differences for Annex III review, showing how the configured system behaved before anyone trusted it with applicants. Better than a vendor slide announcing that the foundation model aced a benchmark last spring.
Alt text: EU AI Act Annex III recruitment audit with vocabulary scoring, persona testing, label ablation and adverse-impact flagging.
My honest concession: these were controlled simulations. Nobody knows whether the protocol detects or reduces discrimination in live hiring outcomes. Still, configuration-specific evidence gives me something concrete to challenge, reproduce and rerun after updates.
Searches for Mistral AI stock or Mistral AI valuation answer another procurement question. The supplied evidence establishes neither a current valuation nor retail stock availability, so I will not fish answers from the internet’s financial-information minestrone. Investors care about ownership and financing. My Annex III file needs the declared purpose, system limitations, test evidence, change policy and documentation handover.
I want European AI champions, including Mistral, to win. Europe needs shared compute and an integrated market where companies can compete with American and Chinese platforms. European origin deserves strategic consideration. Recruitment bias still requires evidence, whether the model came from Paris or Palo Alto.
GDPR evidence has company
GDPR-compliant AI still falls under the AI Act when used within Annex III. Davis Wright Tremaine explains that the AI Act borrows GDPR concepts including profiling and special-category data, while GDPR continues governing personal-data processing. A data protection impact assessment can supply evidence for an AI Act fundamental-rights impact assessment where they overlap. Their scopes differ, so a DPIA cannot complete every AI Act duty. Reuse the evidence. Renaming the file and hoping nobody opens it is the compliance equivalent of parsley on yesterday’s pasta.
Paul and Nandy’s Governance-as-Code proposal puts evidence inside software delivery. A team converts declared compliance requirements into system-checkable acceptance criteria, then machine-executable modules in the CI/CD pipeline. Every run tests the current deployment, including changes since the last legal review, and emits audit evidence indexed to the relevant requirements. Reviewers can trace failed checks to underlying obligations instead of doing SharePoint archaeology. Across the evaluated enterprise deployments, the framework reproduced a manual expert audit’s findings while cutting audit labour by about 75% against that baseline. Nobody knows whether the saving generalises beyond those deployments, but production evidence exposes what binders forget.
Paul and Nandy also identify a major weakness in applying high-risk rules to generative AI. The requirements were largely drafted for predictive systems, leaving gaps around provenance, emergent behaviour, fairness and human oversight. Ronald Schnitzer and his co-authors reached a related conclusion: a minority of high-risk requirements directly address AI-specific risk sources; most cover organisational processes and documentation. Those duties matter. Complete paperwork still cannot prove an open-ended model behaves safely.
Enforcement needs the same honesty. An institutional-design analysis of Germany’s KI-MIG supervisory chamber argues that structural independence is weak because the same agency leadership would staff the chamber and run the broader hierarchical authority. Its enforcement performance cannot yet be assessed. Nobody knows how consistently national authorities will apply the Annex III classification exception once the regime begins.
The European Data Protection Board is trying to align national practice on fines. Its guidelines replace the previous case-specific approach to corrective measures with a five-step methodology. Authorities must confirm their legal power to fine and determine liability, then examine intent or negligence. They also weigh aggravating or mitigating factors and test whether the result is effective, proportionate and dissuasive. Fourteen worked examples now illustrate decisions previously left to case-specific assessment. The EDPB says minor infringements will generally receive no fine and may bring a reprimand, while serious cases carry a strong presumption favouring a fine. Public consultation remains open until 13 November 2026, following adoption at the Board’s latest plenary.
EDPB Deputy Chair Jelena Virant Burnik said in the Board’s announcement:
The new EDPB guidelines are a major step in further aligning how Data Protection Authorities decide whether an administrative fine should be imposed, either on its own or alongside other corrective measures. The GDPR significantly increased the corrective powers of DPAs, with fines serving as an important instrument for effective enforcement. The guidelines reaffirm our commitment to providing greater clarity and ensuring the consistent application of the GDPR across Europe.
This is why I am a European federalist. Twenty-seven incompatible enforcement habits would turn the single market into a legal escape room whose prize is another outside-counsel invoice.
spend the delay on evidence
Stand-alone Annex III obligations now apply from 2 December 2027, replacing the previous deadline of 2 August 2026. The calendar moved. Every classification and evidence problem remains.
Regulation (EU) 2026/1744 is the amending law that replaced the earlier timetable. Clausebench captured the practical consequence perfectly in its September timeline analysis:
A later date is more time to do the same amount of work, not less work.
The sensible order: first, inventory systems and embedded AI features, recording each intended purpose beside the decision its output influences. Then determine whether the organisation is provider or deployer, including modifications that could change its role. For every Annex III system, map existing evidence to applicable duties and expose gaps. Run recruitment probes or other use-specific tests before deployment, then repeat them after a model update, rewritten system prompt, new data pipeline or expanded task. Add the deadline after classification. A calendar built on an unclassified spreadsheet is fake precision with conditional formatting.
Other timelines are already running. At its ninth meeting, the AI Board discussed European Commission implementation support, including guidelines and a Code of Practice. That work supports transparency rules applicable since 2 August 2026, unlike the period before they took effect. Waiting for the Annex III deadline can leave companies late on obligations already here.
By the end of 2027, I expect serious European providers to push machine-readable classification files and test evidence directly into procurement systems. Laggards will arrive with a PDF, a panicked lawyer and an espresso priced like airport Champagne. Europe’s first durable AI champions will turn compliance into plumbing before anyone forces them to.
Frequently asked questions
What systems are high-risk under EU AI Act Annex III?
EU AI Act Annex III covers stand-alone systems used for listed purposes in eight categories. Classification depends on the complete system’s intended use, its influence over consequential decisions, and whether it profiles a natural person. A model’s brand, parameter count, or national origin does not determine high-risk status.
When do EU AI Act Annex III obligations apply?
Stand-alone Annex III obligations apply from 2 December 2027, replacing the previous 2 August 2026 deadline. The delay provides more preparation time but does not reduce the work: organisations still need classification, role mapping, use-specific testing, evidence mapping, and repeat testing after material system changes.
Can GDPR-compliant AI still be high-risk under the AI Act?
GDPR-compliant AI can still be high-risk under the AI Act when its intended use falls within Annex III. GDPR governs personal-data processing, while the AI Act adds system-specific obligations. A data protection impact assessment can support a fundamental-rights impact assessment where evidence overlaps, but it cannot satisfy every AI Act duty.
Sources
- Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment
- Governing in the Gap: The EU AI Act Deferral, the Non-Life Blind Spot, and the Sixteen-Month Window for Insurance Regulators in Emerging Markets
- Changes to Regulation (EU) 2024/1689: the AI Act as it now stands
- Safety components and machinery under the AI Omnibus: What’s changing for manufacturers
- EU AI Act obligations for law firms and legal-AI vendors, by date, article and enforcer, after the Omnibus
- What the Digital Omnibus moved, and which notes to re-date