The EU AI Act is the first major Western regulation that maps specific obligations onto AI system providers and deployers, and a meaningful share of those obligations are evaluation requirements. The Act entered into force on August 1, 2024, and applies in phases. The cheatsheet below is the version you pin above your desk 1.
CAUTION
The timeline moved this month. The Digital Omnibus on AI, Regulation (EU) 2026/1744, was published in the Official Journal on July 24, 2026 and entered into force on July 27, 2026, deferring both high-risk tranches 2. Any cheatsheet you saved before that date, including several still-circulating trackers, has the high-risk dates wrong. The table below reflects the post-Omnibus schedule.
Application timeline
| Date | What goes live |
|---|
| Feb 2, 2025 | Prohibited practices ban (Article 5) and AI literacy obligations |
| Aug 2, 2025 | General-Purpose AI (GPAI) model obligations under Chapter V, plus governance and most penalty provisions |
| Aug 2, 2026 | Article 50 transparency rules apply, and enforcement begins at national and EU level for GPAI models, prohibitions, transparency, and AI literacy 3 |
| Dec 2, 2026 | New prohibition on non-consensual intimate imagery and CSAM; grace period ends for Article 50(2) machine-readable marking on systems placed on the market before Aug 2, 2026 |
| Aug 2, 2027 | Deadline for GPAI models placed on the market before Aug 2, 2025 to reach compliance (Article 111(3)) |
| Dec 2, 2027 | High-risk obligations for standalone Annex III systems (deferred by the Omnibus from Aug 2, 2026) |
| Aug 2, 2028 | High-risk obligations for Annex I systems embedded in regulated products (deferred from Aug 2, 2027) |
The bolded rows are what changed. Two things follow for planning. The deferral bought high-risk providers roughly sixteen months, and the new dates are fixed calendar dates rather than the standards-readiness trigger originally floated, so there is no further conditional slip to wait for. But nothing about August 2, 2026 moved: that is when the Commission can start fining GPAI providers, and it is days away.
Read the dates as the latest a given obligation can come into force. Treating them as soft has not been a winning bet on prior EU technology regulation 1.
Figure: EU AI Act application timeline after the 2026 Digital Omnibus, Regulation (EU) 2026/1744. GPAI enforcement and fines of up to 3% of worldwide turnover or EUR 15 million start August 2, 2026, unchanged, while high-risk obligations shift to December 2, 2027 for standalone Annex III systems and August 2, 2028 for Annex I embedded products.
What August 2, 2026 actually turns on
Enforcement. The GPAI obligations have applied since August 2, 2025, but Article 101, the Commission's power to fine GPAI model providers directly, was expressly carved out of that tranche by Article 113(b) and falls to the general date. From August 2, 2026 the Commission can impose fines of up to 3% of annual total worldwide turnover or EUR 15,000,000, whichever is higher, on a GPAI provider that breaches the Regulation, fails to supply requested documents, supplies inaccurate information, refuses a Commission-ordered measure, or denies the Commission access to evaluate the model.
Keep that separate from the Article 99 tiers, which national authorities administer and which are the numbers usually quoted in headlines: up to EUR 35,000,000 or 7% of worldwide turnover for prohibited practices, EUR 15,000,000 or 3% for most other infringements including Article 50 transparency, and EUR 7,500,000 or 1% for supplying incorrect or misleading information. For SMEs each cap is the lower of the two figures rather than the higher.
The practical reading for an evals owner: the denial-of-access limb is the one that touches your work. If the Commission asks to evaluate a model and the provider cannot produce the evaluation record, that is the exposure.
The four risk tiers
| Tier | What it covers | Eval implication |
|---|
| Prohibited (Art. 5) | Social scoring, manipulative systems, real-time biometric ID in public (narrow exceptions) | If you are in this tier, the obligation is not eval; it is exit |
| High-risk (Annex III) | Critical infra, education, employment, essential services, law enforcement, migration, justice, democratic process | Quality management system, conformity assessment, post-market monitoring, automated logging, accuracy and robustness testing, human oversight |
| Limited-risk | Chatbots, generative content, emotion recognition (where used) | Transparency: users informed they are interacting with AI, content marked as AI-generated |
| Minimal-risk | Everything else | No specific obligations |
A separate axis cuts across the tiers for GPAI models (the regime targets foundation-model providers specifically) and for GPAI models with systemic risk (the largest models, currently identified by training-compute thresholds and Commission designation).
GPAI obligations, in plain text
If you provide a general-purpose AI model placed on the EU market on or after August 2, 2025, three obligations apply regardless of tier:
- Technical documentation. A summary of training content, training methodology, energy consumption, and capability and limitation testing. The Commission publishes a template under the GPAI code of practice.
- Copyright policy. A documented policy for respecting EU copyright law, including a policy on text and data mining opt-outs.
- Summary of training data. A public-facing summary "sufficiently detailed" to let copyright holders assess whether their works were used.
If your model is classified as having systemic risk, four further obligations apply:
- Model evaluation under standardized protocols, including adversarial testing.
- Systemic-risk assessment and mitigation, with documentation maintained.
- Serious-incident reporting to the AI Office and national authorities.
- Cybersecurity protections appropriate to the model's capabilities.
Obligation 4 is where the evals chapter of your governance program lives. The Act does not name specific benchmarks; the GPAI Code of Practice is the soft-law instrument that points at concrete protocols.
What the Code of Practice actually asks for
The Code was published on July 10, 2025 and confirmed as an adequate voluntary compliance tool by the Commission and the AI Board on August 1, 2025. Some 26 organisations signed at the August 2025 launch, including OpenAI, Google, Microsoft, Anthropic, Amazon, IBM, Mistral, and Cohere; the launch count included xAI, which signed only the Safety and Security chapter, and Meta declined to sign. The Commission maintains the current signatories list, which stood at 23 full signatories plus xAI as of its April 2026 update 4.
That Safety and Security chapter carries ten commitments and applies only to models with systemic risk. Commitment 3, systemic risk analysis, is the one your eval program answers to. Measure 3.2 requires "at least state-of-the-art model evaluations" using "methods that are appropriate for the model and the systemic risk," and requires that evaluations "include open-ended testing of the model to improve the understanding of the systemic risk." The methods it enumerates are worth reading as a checklist, because they map directly onto sections of this site: question-and-answer sets, task-based evaluations, benchmarks, red-teaming and other adversarial testing, human uplift studies, model organisms, simulations, and proxy evaluations for classified materials 4.
Two readings follow. First, "state-of-the-art" is a moving bar by construction, so a frozen eval suite drifts out of compliance without anyone changing a line of it; treat the Code like the OWASP Top 10, a document you re-baseline against. Second, "open-ended testing" is explicitly not satisfiable by benchmark scores alone, which is the same argument the red-team program design chapter makes on engineering grounds: the standing corpus regresses what you know, and human sessions are where new categories come from.
High-risk obligations, in plain text
If you provide a high-risk AI system (Annex III categories, eight families covering critical infrastructure, education, employment, essential public and private services, law enforcement, migration, administration of justice, and democratic processes), the obligations include a quality management system, technical documentation maintained throughout the lifecycle, automated event logging, human oversight provisions, accuracy and cybersecurity testing, and conformity assessment before placement on the market. These now bite on December 2, 2027 for standalone Annex III systems and August 2, 2028 for Annex I systems embedded in regulated products 2.
The deferral is a scheduling change, not a scope change. Conformity assessment still requires evidence that accumulates over quarters, not weeks: a documented accuracy and robustness testing regime, a risk-management file with revision history, and post-market monitoring that has actually run. Teams that read the extra sixteen months as permission to start later tend to discover that the evidence they need is retrospective.
The eval-specific obligations are accuracy testing (Art. 15), risk management (Art. 9), and post-market monitoring (Art. 72). Accuracy and robustness testing must be documented and proportionate to intended use. The accepted approach is to point at a recognized framework (the NIST AI RMF maps cleanly, see the next chapter) and demonstrate that your eval activities cover the required dimensions 5.
What to document, regardless of tier
The minimum evidence pack for an Act audit conversation:
| Document | What it contains |
|---|
| Eval methodology summary | What benchmarks, what cadence, what release gates |
| Risk register (see AI risk register) | Identified risks, severity, mitigations, residual risk |
| Model and system cards (see Customer trust artifacts) | Capabilities, limitations, eval results with dates |
| Post-market monitoring report | Production incidents, response, model updates triggered |
| Human oversight description | Which humans, when, with what authority to override |
Microsoft and other major vendors publish reference implementations of this evidence pack 6. The shape is becoming standardized; you do not need to invent it from scratch.
A common misreading, named
The Act applies to systems "placed on the market or put into service" in the EU. Some teams read this as "we are a US company, this does not apply." It does. The Act has explicit extraterritorial scope when (a) the output of the system is used in the EU, (b) the system's deployer is established in the EU, or (c) the system is made available on the EU market. For most AI SaaS products, at least one of these is true.
The second common misreading is treating compliance as a one-time audit. The Act's post-market monitoring and risk-management provisions are continuous; the eval program is what feeds those obligations. A spike in production failure rate that you do not investigate and respond to is a separate compliance issue from the rate itself.
What to do this quarter
- Re-date your compliance plan against the post-Omnibus schedule. If your plan still says Annex III lands August 2026, it is describing a law that was amended on July 27.
- Identify whether your system is in scope (most are; check Annex III against your product surfaces).
- Map your existing eval activities to the four eval-adjacent obligations: accuracy and robustness, risk management, post-market monitoring, and adversarial testing for GPAI systemic risk.
- Stand up the evidence pack above. The risk register, model card, and post-market monitoring report are the three artifacts that take real work; the rest are summaries of work you are already doing.
The next chapter, NIST AI RMF mapped, gives you the cross-walk that makes the eval-to-obligation mapping defensible in writing.
NOTE
Verified against primary sources on July 29, 2026, two days after the Omnibus entered into force. Several widely-linked timeline trackers had not yet been updated at that point and still showed the pre-Omnibus high-risk dates. When a date on this page matters to a decision, read it against the Official Journal text rather than any tracker, including this one.