Operational guidance, not legal advice. This page distills named public sources (regulator guidance and industry practice). It is not a legal determination, not a notification decision, and not a substitute for your counsel, insurer, or a retained DFIR firm. Verify applicability and current deadlines for your facts and jurisdiction.

What model evaluation and red teaming does the EU AI Act require?

Updated

Regulation (EU) 2024/1689 splits evaluation in two: Article 15 is high-risk accuracy, robustness and cybersecurity, declared in the instructions for use; Article 55(1)(a) is GPAI systemic-risk model evaluation including documented adversarial testing. Not legal advice. It does not start a clock.

AI model evaluation and adversarial testing, last verified 10 September 2026 against Articles 9, 13, 15, 51, 55, 56 and 113 of Regulation (EU) 2024/1689 (OJ L 2024/1689, 12.7.2024). Article 15 is not Article 55(1)(a), and Article 9(6) and 9(8) testing is a third thing. The GPAI Code of Practice safety and security chapter is Commission and AI Office methodology — guidance, not the regulation, and not a legal substitute for Article 55(1)(a). NIST AI RMF 1.0, including its Measure function, is guidance, not law. ISO/IEC 42001:2023 is a management-system standard, not the regulation. This page is not legal advice, not a filing, and does not start a clock. It does not determine that the Act applies to YOU, that YOU are a provider or a deployer, that YOUR system is high-risk, or that YOUR model has systemic risk. It does not run YOUR tests and does not evaluate YOUR model.

This is Article 15 and Article 55(1)(a), not YOUR test plan

Audience: an ML lead, CISO, or counsel deciding what evaluation and red-team evidence Regulation (EU) 2024/1689 actually expects, and from whom. This page is not legal advice. It does not start a clock. Reading it does not start a clock. Mapping a row is not a determination that the Act applies, that YOU are a provider or a deployer, that YOUR system is high-risk, or that YOUR model has systemic risk. This page does not run YOUR tests, does not evaluate YOUR model, does not adversarially test anything, does not issue certifications, and does not file with the AI Office.

The AI Act is Regulation (EU) 2024/1689 of 13 June 2024, OJ L 2024/1689, 12.7.2024. ELI: http://data.europa.eu/eli/reg/2024/1689/oj. Three different texts are routinely collapsed into one phrase called model evaluation. Article 15 sets accuracy, robustness and cybersecurity requirements for high-risk AI systems, with the accuracy metrics declared in the instructions for use under Articles 15(3) and 13(3)(b)(ii). Articles 9(6) and 9(8) require testing to identify risk-management measures, against prior defined metrics and probabilistic thresholds. Article 55(1)(a) requires providers of general-purpose AI models with systemic risk to perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing. They are different duties on different subjects with different application dates. Last verified 10 September 2026. Not legal advice.

  • Statute versus methodology versus standard versus recommendation: Articles 9, 13, 15, 51, 55, 56 and 113 of 2024/1689 are legal requirements only if they apply. The GPAI Code of Practice safety and security chapter is Commission and AI Office methodology — guidance, not the regulation. NIST AI RMF 1.0 and its Measure function are guidance, not law. ISO/IEC 42001:2023 is a management-system standard, not the regulation. The test-plan illustration below is a ShipReadyMetrics recommendation. This page labels which kind of text each claim rests on.
  • The AI-risk-management-requirements guide on this site is the Articles 9 and 55 page. The GPAI-systemic-risk guide on this site is the Article 51–55 page. The GPAI-requirements guide on this site is the Article 53 baseline page. The technical-documentation guide on this site is the Articles 11 and 53 Annex IV/XI/XII page. The AI-audit-evidence guide on this site is the evidence-request-by-role page. A dedicated AI-cybersecurity-requirements guide is not on this site yet. Naming it is not a link.
  • This page does not invent a 2 August 2026 date for Annex I. Articles 9 and 15 sit in Chapter III Section 2. The original Article 113 second paragraph applies the rest of the Regulation from 2 August 2026, and Article 113(c) keeps Article 6(1) and the corresponding obligations — Annex I product-embedded high-risk — on 2 August 2027. Article 55 sits in Chapter V, which Article 113(b) applies from 2 August 2025 with the exception of Article 101. This page does not invent a 2 August 2026 start date for GPAI. Those dates are not one number.

Article 15 is not Article 55(1)(a)

Do not conflate them. Article 15 is a requirement on a high-risk AI system, gated by Article 6. Article 55(1)(a) is a duty on the provider of a general-purpose AI model with systemic risk, gated by Article 51. Articles 9(6) and 9(8) are high-risk testing in service of the Article 9 risk-management system — related to Article 15, and not the same text as Article 55(1)(a). A high-risk AI system is not automatically a systemic-risk GPAI model, and a systemic-risk GPAI model is not automatically a high-risk AI system. Last verified 10 September 2026. Not legal advice.

Article 15 versus Articles 9(6) and 9(8) versus Article 55(1)(a) (not YOUR test plan; not a determination that any of them binds YOU; not legal advice)
DutyWhat the cited text isKind of textLast verified
Article 15 — high-risk accuracy, robustness, cybersecurityHigh-risk AI systems shall achieve an appropriate level of accuracy, robustness and cybersecurity and perform consistently in those respects throughout their lifecycle. The accuracy levels and relevant accuracy metrics are declared in the accompanying instructions of use. It is a property requirement on the system, not a prescribed evaluation protocol, and it does not mention adversarial testing by that name.Article 15 of 2024/1689. Legal requirement, only if Article 6 high-risk applies. Distinct from Article 55(1)(a). This page does not find that YOUR system is high-risk.10 September 2026
Articles 9(6) and 9(8) — testing for the risk-management systemHigh-risk AI systems shall be tested for the purpose of identifying the most appropriate and targeted risk management measures, at any time throughout development and in any event before being placed on the market or put into service, against prior defined metrics and probabilistic thresholds. This is testing in service of Article 9, related to Article 15 but not the same paragraph, and not Article 55(1)(a).Articles 9(6) and 9(8) of 2024/1689. Legal requirement, only if it applies. The AI-risk-management-requirements guide on this site is the Articles 9 and 55 page. This page does not run YOUR tests.10 September 2026
Article 55(1)(a) — GPAI systemic-risk model evaluationProviders of general-purpose AI models with systemic risk shall perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model with a view to identifying and mitigating systemic risks. This is the only place the Regulation names adversarial testing as a duty, and it is a model-level duty, not a high-risk-system requirement.Article 55(1)(a) of 2024/1689. Legal requirement, only if Article 51 systemic-risk applies. In force 2 August 2025 via Article 113(b). This page does not designate YOUR model and does not evaluate YOUR model.10 September 2026
Article 6 — the high-risk gate, not the systemic-risk gateArticle 6 classifies high-risk AI systems (Annex I product-embedded under Article 6(1); Annex III use-cases under Article 6(2)). That classification is the gate for Articles 9 and 15. It is not the Article 51 systemic-risk gate.Article 6 of 2024/1689. Legal requirement, only if it applies. This page does not classify YOUR system.10 September 2026
Article 51 — the systemic-risk gate, not the high-risk gateArticle 51 classifies general-purpose AI models with systemic risk. That classification is the gate for Article 55(1)(a). It is not the Article 6 high-risk gate, and it is not Article 15.Article 51 of 2024/1689. Legal requirement, only if it applies. The GPAI-systemic-risk guide on this site is the Article 51–55 page.10 September 2026

What original Article 15 actually says

Last verified 10 September 2026 against Article 15 of Regulation (EU) 2024/1689 on EUR-Lex (OJ L 2024/1689, 12.7.2024). These are legal requirements of the original regulation, only if they apply. This page does not apply them to YOU and does not test YOUR system against them. Not legal advice.

Article 15 as the original regulation states it (not YOUR evaluation; not a determination that Article 15 binds YOU; not legal advice)
PointWhat the cited text saysKind of textLast verified
Article 15(1)Authentic Article 15(1): High-risk AI systems shall be designed and developed in such a way that they achieve an appropriate level of accuracy, robustness, and cybersecurity, and that they perform consistently in those respects throughout their lifecycle.Article 15(1) of 2024/1689. Legal requirement, only if it applies. Note the words appropriate and throughout their lifecycle — no figure is stated. This page does not decide what is appropriate for YOUR system.10 September 2026
Article 15(2)Authentic Article 15(2): To address the technical aspects of how to measure the appropriate levels of accuracy and robustness set out in paragraph 1 and any other relevant performance metrics, the Commission shall, in cooperation with relevant stakeholders and organisations such as metrology and benchmarking authorities, encourage, as appropriate, the development of benchmarks and measurement methodologies.Article 15(2) of 2024/1689. Legal requirement addressed to the Commission, not to you. Benchmarks encouraged under this paragraph are measurement methodology — guidance, not the regulation.10 September 2026
Article 15(3)Authentic Article 15(3): The levels of accuracy and the relevant accuracy metrics of high-risk AI systems shall be declared in the accompanying instructions of use.Article 15(3) of 2024/1689. Legal requirement, only if it applies. Read with Article 13(3)(b)(ii). This page does not draft YOUR instructions for use.10 September 2026
Article 15(4), first subparagraphAuthentic Article 15(4): High-risk AI systems shall be as resilient as possible regarding errors, faults or inconsistencies that may occur within the system or the environment in which the system operates, in particular due to their interaction with natural persons or other systems. Technical and organisational measures shall be taken in this regard.Article 15(4) of 2024/1689. Legal requirement, only if it applies. This is the robustness limb. This page does not judge YOUR resilience.10 September 2026
Article 15(4), second and third subparagraphsAuthentic Article 15(4): The robustness of high-risk AI systems may be achieved through technical redundancy solutions, which may include backup or fail-safe plans. High-risk AI systems that continue to learn after being placed on the market or put into service shall be developed in such a way as to eliminate or reduce as far as possible the risk of possibly biased outputs influencing input for future operations (feedback loops), and as to ensure that any such feedback loops are duly addressed with appropriate mitigation measures.Article 15(4) of 2024/1689. Legal requirement, only if it applies. The feedback-loop subparagraph is why a one-off pre-deployment result does not evidence a continuously learning system.10 September 2026
Article 15(5), first and second subparagraphsAuthentic Article 15(5): High-risk AI systems shall be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting system vulnerabilities. The technical solutions aiming to ensure the cybersecurity of high-risk AI systems shall be appropriate to the relevant circumstances and the risks.Article 15(5) of 2024/1689. Legal requirement, only if it applies. A dedicated AI-cybersecurity-requirements guide is not on this site yet. Naming it is not a link.10 September 2026
Article 15(5), third subparagraphAuthentic Article 15(5): The technical solutions to address AI specific vulnerabilities shall include, where appropriate, measures to prevent, detect, respond to, resolve and control for attacks trying to manipulate the training data set (data poisoning), or pre-trained components used in training (model poisoning), inputs designed to cause the AI model to make a mistake (adversarial examples or model evasion), confidentiality attacks or model flaws.Article 15(5) of 2024/1689. Legal requirement, only if it applies. This is the closest the high-risk track comes to naming adversarial attacks, and it is a design-measure duty, not the Article 55(1)(a) adversarial-testing duty.10 September 2026
Article 13(3)(b)(ii) — the declared metricsAuthentic Article 13(3)(b)(ii): the level of accuracy, including its metrics, robustness and cybersecurity referred to in Article 15 against which the high-risk AI system has been tested and validated and which can be expected, and any known and foreseeable circumstances that may have an impact on that expected level of accuracy, robustness and cybersecurity;Article 13(3)(b)(ii) of 2024/1689. Legal requirement, only if it applies. This is where a number ends up in front of a deployer. This page does not state a figure for YOUR system.10 September 2026

What original Articles 9(6) and 9(8) actually say — the third text

Last verified 10 September 2026 against Article 9 of Regulation (EU) 2024/1689 on EUR-Lex (OJ L 2024/1689, 12.7.2024). Article 9 testing is not Article 15, and it is not Article 55(1)(a). It is testing in service of the risk-management system. The AI-risk-management-requirements guide on this site is the Articles 9 and 55 page. Not legal advice.

Articles 9(6), 9(7) and 9(8) as the original regulation states them (not YOUR testing; not Article 55(1)(a); not legal advice)
PointWhat the cited text saysKind of textLast verified
Article 9(6)Authentic Article 9(6): High-risk AI systems shall be tested for the purpose of identifying the most appropriate and targeted risk management measures. Testing shall ensure that high-risk AI systems perform consistently for their intended purpose and that they are in compliance with the requirements set out in this Section.Article 9(6) of 2024/1689. Legal requirement, only if it applies. This Section is Chapter III Section 2, which includes Article 15. This page does not run YOUR tests.10 September 2026
Article 9(7)Authentic Article 9(7): Testing procedures may include testing in real-world conditions in accordance with Article 60.Article 9(7) of 2024/1689. Legal requirement of the option, only if it applies. This page does not authorise YOUR real-world test.10 September 2026
Article 9(8)Authentic Article 9(8): The testing of high-risk AI systems shall be performed, as appropriate, at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service. Testing shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system.Article 9(8) of 2024/1689. Legal requirement, only if it applies. Prior defined means the metric and threshold are chosen before the test, not after the result. This page does not pick YOUR thresholds.10 September 2026

What original Article 55(1)(a) actually says — the GPAI evaluation track

Last verified 10 September 2026 against Article 55 of Regulation (EU) 2024/1689 on EUR-Lex (OJ L 2024/1689, 12.7.2024). Article 55(1)(a) is the adversarial-testing duty. It applies to providers of general-purpose AI models with systemic risk under Article 51, not to every model and not to every high-risk system. Mapping a row is not a designation that YOUR model has systemic risk. Not legal advice.

Article 55(1)(a) and its Article 56 route as the original regulation states them (not YOUR Article 55 file; not a designation; not legal advice)
PointWhat the cited text saysKind of textLast verified
Article 55(1) chapeauAuthentic Article 55(1): In addition to the obligations listed in Articles 53 and 54, providers of general-purpose AI models with systemic risk shall:Article 55(1) of 2024/1689. Legal requirement, only if Article 51 systemic-risk applies. In addition means the Article 53 baseline duties stay.10 September 2026
Article 55(1)(a)Authentic Article 55(1)(a): perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model with a view to identifying and mitigating systemic risks;Article 55(1)(a) of 2024/1689. Legal requirement, only if it applies. Note documenting: an undocumented red-team exercise is not the evidence this point describes. This page does not evaluate YOUR model and does not adversarially test it.10 September 2026
Article 55(1)(b)Authentic Article 55(1)(b): assess and mitigate possible systemic risks at Union level, including their sources, that may stem from the development, the placing on the market, or the use of general-purpose AI models with systemic risk;Article 55(1)(b) of 2024/1689. Legal requirement, only if it applies. The evaluation in point (a) feeds this assessment; it does not replace it.10 September 2026
Article 55(2) — the code-of-practice routeAuthentic Article 55(2): Providers of general-purpose AI models with systemic risk may rely on codes of practice within the meaning of Article 56 to demonstrate compliance with the obligations set out in paragraph 1 of this Article, until a harmonised standard is published. Compliance with European harmonised standards grants providers the presumption of conformity to the extent that those standards cover those obligations. Providers who do not adhere to an approved code of practice or do not comply with a European harmonised standard shall demonstrate alternative adequate means of compliance for assessment by the Commission.Article 55(2) of 2024/1689. Legal requirement of the option, only if it applies. A code of practice is voluntary. It is a route to demonstrating compliance, not a replacement for Article 55(1)(a).10 September 2026
Article 9(6) and 9(8) are not this dutyHigh-risk testing under Articles 9(6) and 9(8) is a different subject with a different gate. A provider running Article 9 testing has not thereby performed Article 55(1)(a) model evaluation, and a provider documenting adversarial testing under Article 55(1)(a) has not thereby met Article 15.Articles 9, 15 and 55 of 2024/1689. Legal requirements, only if they apply. Do not conflate them.10 September 2026

Article 113: the high-risk dates are not the GPAI dates

Last verified 10 September 2026 against Article 113 of Regulation (EU) 2024/1689 on EUR-Lex (OJ L 2024/1689, 12.7.2024). Articles 9 and 15 sit in Chapter III Section 2. The original Article 113 second paragraph applies the rest of the Regulation from 2 August 2026. Article 113(c) keeps Article 6(1) and the corresponding obligations — Annex I product-embedded high-risk — on 2 August 2027, not 2 August 2026. This page does not invent a 2 August 2026 date for Annex I.

Article 55 sits in Chapter V. Article 113(b): Chapter III Section 4, Chapter V, Chapter VII and Chapter XII and Article 78 shall apply from 2 August 2025, with the exception of Article 101. GPAI systemic-risk model evaluation and adversarial-testing duties therefore apply from 2 August 2025 under Article 113(b), except Article 101. They did not start on 2 August 2026. This page does not invent a 2 August 2026 start date for GPAI. Those dates are not one number.

Regulation (EU) 2026/1744 is an amending regulation. Counsel reads the authentic operative article of any amendment. This page does not apply 2026/1744 to YOU, does not treat a recital as rewriting Article 113(b) for Chapter V, and does not move original Article 113(c) Annex I off 2 August 2027. Not legal advice.

The GPAI Code of Practice is methodology, not the article

The General-Purpose AI Code of Practice, published 10 July 2025, has three chapters: transparency, copyright, and safety and security. The safety and security chapter is where evaluation and red-team methodology lives — how a signatory says it will run model evaluations, what it treats as state of the art, and how it documents adversarial testing. It is Commission and AI Office material. It is guidance, not the regulation. Signing it, or copying its method, is not a determination that Article 55(1)(a) is met, and declining to sign it does not remove Article 55(1)(a) — Article 55(2) leaves a provider who does not adhere to an approved code to demonstrate alternative adequate means of compliance for assessment by the Commission. Last verified 10 September 2026. Not legal advice.

NIST AI RMF 1.0 (NIST AI 100-1, January 2023) organises work under Govern, Map, Measure and Manage. The Measure function is where evaluation, testing and red-teaming practice sits. It is voluntary US agency guidance, not law, and mapping a Measure category is not a finding that Article 15 or Article 55(1)(a) is met. ISO/IEC 42001:2023 is a management-system standard: it asks an organisation to maintain AI system impact assessments and life-cycle records, not to run a specified evaluation protocol. None of the three is Regulation (EU) 2024/1689.

Legal requirement versus Code of Practice methodology versus guidance versus standard versus recommendation (not a ranking; not legal advice; last verified 10 September 2026)
TextWhat it isWhat this page does not do
Regulation (EU) 2024/1689 Article 15Legal requirement — appropriate accuracy, robustness and cybersecurity for a high-risk AI system, consistent throughout its lifecycle, with accuracy metrics declared in the instructions of use, only if Article 6 high-risk applies. Chapter III Section 2, with Annex I corresponding obligations on 2 August 2027 under Article 113(c).Does not determine that Article 15 binds YOU, does not test YOUR system, and does not invent a 2 August 2026 date for Annex I.
Regulation (EU) 2024/1689 Articles 9(6) and 9(8)Legal requirement — testing to identify risk-management measures, throughout development and before placing on the market, against prior defined metrics and probabilistic thresholds, only if it applies.Does not run YOUR tests, does not pick YOUR metrics or thresholds, and does not treat this testing as Article 55(1)(a) model evaluation.
Regulation (EU) 2024/1689 Article 55(1)(a)Legal requirement — model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing, only if Article 51 systemic-risk applies. In force 2 August 2025 under Article 113(b), except Article 101.Does not designate YOUR model, does not evaluate YOUR model, does not conflate Article 55(1)(a) with Article 15, and does not invent a 2 August 2026 start date for GPAI.
GPAI Code of Practice — safety and security chapter (published 10 July 2025)Commission and AI Office methodology. Guidance, not the regulation. Voluntary. Article 55(2) makes an Article 56 code a route to demonstrating compliance until a harmonised standard is published.Does not treat the Code of Practice as a legal substitute for Article 55(1)(a), and does not treat signature or non-signature as a compliance verdict.
NIST AI RMF 1.0 (NIST AI 100-1, January 2023), Measure functionGuidance, not law. A voluntary US agency framework whose Measure function covers evaluation, testing and red-teaming practice.Does not treat a Measure category as discharging Article 15 or Article 55(1)(a).
ISO/IEC 42001:2023Best practice / standard. An AI management-system standard, certifiable by an accredited certification body. Not a legal substitute for the Act.Does not treat an ISO 42001 certificate as an evaluation result, as CE marking, or as discharging Article 15.
Published benchmarks and open red-team toolingIndustry best practice. Useful, moving, and unaudited. Article 15(2) has the Commission encourage the development of benchmarks and measurement methodologies; a benchmark score is not a legal conclusion.Does not rank tools, does not name a benchmark as required, and does not treat a published ranking as evidence that a duty is met.
This product's AI model-evaluation registerShipReadyMetrics recommendation: a register of evaluations the organisation recorded, graded for coverage and currency. Not a legal determination and not an evaluation.Does not run evaluations, does not adversarially test, does not certify that a result was adequate, and does not file. A named human still owns the assessment.

Example test plan — an illustration, not YOUR Article 55 file

The table below is an illustration of how a pre- and post-deployment evaluation plan can be laid out so that each row points at the text it is meant to evidence. It is a ShipReadyMetrics recommendation. It is not YOUR Article 55 file, not YOUR Article 9 testing record, and not a template the Regulation prescribes — no article of 2024/1689 prescribes a test-plan format. Walking the rows is not a determination that any of these articles binds YOU. Last verified 10 September 2026. Not legal advice.

Example pre- and post-deployment test plan (an illustration; a ShipReadyMetrics recommendation; not YOUR Article 55 file; not legal advice)
Example stageWhat is tested in the exampleWhich text it points atKind of text
Define the metric and the threshold firstWrite down the accuracy metric, the robustness measure, and the pass threshold for the intended purpose before running anything, with the rationale for the number.Articles 9(8) and 15(1) and (3): prior defined metrics and probabilistic thresholds; accuracy levels and metrics declared in the instructions of use.Legal requirements of 2024/1689, only if they apply, laid out as a ShipReadyMetrics recommendation. This page does not pick YOUR threshold.
Pre-deployment functional evaluationMeasure accuracy against the defined metric on held-out data representative of the intended purpose, and record the environment, dataset version, and date.Articles 9(6) and 9(8): testing before being placed on the market or put into service, ensuring consistent performance for the intended purpose.Legal requirements of 2024/1689, only if they apply. NIST AI RMF Measure is guidance on how, not law.
Robustness and failure-mode evaluationPerturbed inputs, out-of-distribution inputs, degraded dependencies, and the fallback path; record what the system did rather than what it was supposed to do.Article 15(4): resilience to errors, faults or inconsistencies, including technical redundancy such as backup or fail-safe plans.Legal requirement of 2024/1689, only if it applies. The specific perturbation set is industry best practice, not a prescribed list.
Bias and non-discrimination evaluationDisaggregated performance across the groups the use case affects, plus a check for feedback loops if the system continues to learn after deployment.Article 15(4) third subparagraph on feedback loops, read with the Article 10 data-governance duties.Legal requirements of 2024/1689, only if they apply. This page does not decide which groups YOUR use case affects.
Security and adversarial evaluation of the deployed systemData-poisoning, model-poisoning, evasion, prompt-injection, extraction and confidentiality attempts against the deployed configuration, with the mitigation for each finding.Article 15(5): resilience against unauthorised third parties, with measures addressing data poisoning, model poisoning, adversarial examples or model evasion, confidentiality attacks and model flaws.Legal requirement of 2024/1689, only if it applies. A dedicated AI-cybersecurity-requirements guide is not on this site yet. Naming it is not a link.
Systemic-risk model evaluation and documented adversarial testingOnly for a provider of a GPAI model designated under Article 51: model evaluation under standardised protocols reflecting the state of the art, with adversarial testing conducted and documented, and the systemic risks identified and mitigated.Article 55(1)(a) and 55(1)(b): model evaluation including conducting and documenting adversarial testing; assess and mitigate possible systemic risks at Union level.Legal requirements of 2024/1689, only if Article 51 applies. The GPAI Code of Practice safety and security chapter is methodology — guidance, not the regulation.
Post-deployment re-evaluation on a stated cadenceRe-run the plan on a written cadence and after any substantial modification, because a result ages: the model, its inputs, and its environment move.Article 15(1) performing consistently throughout the lifecycle; Article 9(2) continuous iterative process requiring regular systematic review and updating; Article 72 post-market monitoring.Legal requirements of 2024/1689, only if they apply. The cadence figure is a ShipReadyMetrics recommendation, not a statutory interval.
Write the result where it will be readPut the declared accuracy metric and the expected level into the instructions for use, and keep the underlying results with the technical documentation.Articles 15(3) and 13(3)(b)(ii) for the declaration; Article 11 with Annex IV for the documentation.Legal requirements of 2024/1689, only if they apply. The technical-documentation guide on this site is the Articles 11 and 53 Annex IV/XI/XII page.

What to do now

As of last verification on 10 September 2026, Article 55 GPAI systemic-risk duties, including the Article 55(1)(a) model-evaluation and adversarial-testing duty, have applied since 2 August 2025 under Article 113(b), except Article 101. High-risk Chapter III Section 2 duties, including Articles 9 and 15, sit on the original 2 August 2026 residual, except Article 6(1) and the corresponding obligations on 2 August 2027 under Article 113(c). The list below is operational preparation. It is not a determination that Article 15 or Article 55(1)(a) binds YOU. Walk it with counsel.

  • Ask counsel which text you are actually under: Article 6 high-risk (Articles 9 and 15), Article 51 systemic-risk GPAI (Article 55(1)(a)), both, or neither. This page does not run either test, and the answer changes what evidence is worth producing.
  • Write the metric and the threshold down before the next evaluation run. Article 9(8), if it applies, speaks of prior defined metrics and probabilistic thresholds; a threshold chosen after seeing the result is not that.
  • If counsel finds a high-risk system, decide where the declared accuracy metric will live. Articles 15(3) and 13(3)(b)(ii) put it in the instructions for use, in front of the deployer, not only in an internal report.
  • If counsel finds a GPAI model with systemic risk, treat documentation as part of the duty. Article 55(1)(a) requires conducting and documenting adversarial testing; an undocumented exercise leaves nothing to show.
  • Use the GPAI Code of Practice safety and security chapter, NIST AI RMF Measure, and published benchmarks as methodology, and label them that way. They are guidance and best practice, not the regulation, and they do not discharge Article 15 or Article 55(1)(a).
  • Put a re-evaluation cadence in writing, and re-run after a substantial modification. Article 15(1) speaks of performing consistently throughout the lifecycle; a two-year-old result does not describe today's model.
  • The AI-risk-management-requirements guide on this site is the Articles 9 and 55 page. The GPAI-systemic-risk guide on this site is the Article 51–55 page. A dedicated AI-cybersecurity-requirements guide is not on this site yet. Naming it is not a link.

Evaluation checklist

This is a question list, not a filing, not an evaluation, and not a determination that Article 15 or Article 55(1)(a) binds YOU. Walk it with counsel. The EU AI Act overview on this site is the pillar page.

  • Does the Act apply to YOU at all? Articles 2 and 3. This page does not run that test.
  • Does Article 6 high-risk apply? Article 15 is a legal requirement only if it does. This page does not classify YOUR system.
  • Does Article 51 systemic-risk apply to a model YOU provide? Article 55(1)(a) is a legal requirement only if it does. This page does not designate YOUR model.
  • For each evaluation, was the metric and the probabilistic threshold defined before the run, and is the rationale written down?
  • Is the accuracy level, with its metrics, declared in the instructions for use under Articles 15(3) and 13(3)(b)(ii)?
  • Was robustness tested against errors, faults and inconsistencies, and is the fallback path evidenced rather than assumed?
  • Were the Article 15(5) attack classes considered — data poisoning, model poisoning, adversarial examples or model evasion, confidentiality attacks, model flaws — and is each mitigation recorded?
  • If the system continues to learn after deployment, is the feedback-loop risk addressed with mitigation measures, per Article 15(4)?
  • For a systemic-risk GPAI model, is the adversarial testing documented, not merely conducted, per Article 55(1)(a)?
  • Does the GPAI Code of Practice, NIST AI RMF, or an ISO 42001 certificate discharge Article 15 or Article 55(1)(a)? No. They are methodology, guidance and a standard, not the regulation.
  • Is there a written re-evaluation cadence, and does a substantial modification trigger a re-run?
  • Document the assessment, including a not-in-scope decision and a decision that no evaluation is owed. This page does not keep YOUR file.

Where this shows up in ShipReady Metrics

The bundled framework key eu_ai_act is customer-visible. Its version label is Regulation (EU) 2024/1689 high-risk obligations (starter subset). The control-set is a starter subset, illustrative, to be tailored by a compliance owner; not legal advice; not a conformity determination; not CE marking. Readiness is not compliance.

If you already have a session: signed-in app → Compliance → AI governance holds the AI inventory, the AI risk register, and a model-evaluation register. That register records evaluations a human entered against four dimensions — accuracy, robustness and cybersecurity for Article 15, and bias for the Article 10 data-governance duties — per registered high-risk system, and grades each dimension current, failing, stale, or missing. Only systems the organisation classified as high-risk are treated as Article 15 subjects; a limited or minimal system is not a subject, never a fabricated gap.

The grading is deliberately unflattering. A dimension with no recorded evaluation reads missing, never inferred as tested. An evaluation with no usable date, or one older than the currency window, reads stale rather than current, and an otherwise-recent evaluation recorded before the system's last substantial modification also reads stale, because it tested the previous model. A recorded fail reads failing and is never masked as current or green; a demonstrated failure also blocks the affected AI control from reading met on artifact presence alone. With no high-risk system registered there is nothing to assess, and the register says so rather than reporting a pass.

This product does not run evaluations, does not adversarially test or red-team anything, does not choose YOUR metrics or thresholds, does not certify that a recorded result was adequate, does not perform Article 55(1)(a) model evaluation, does not file with the AI Office, does not issue certifications, does not affix CE marks, and is not a notified body. It records and grades what a human entered. Missing data is a disclosed gap, not a pass. A named human still owns the assessment.

This page does not document a public demo URL. There is no public EU AI Act demo path. This product does not start a clock.

Primary sources (last verified 10 September 2026)

Every regulatory or guidance claim on this page is taken from one of these. If a later revision of a source changes the rule, the date above is how you can see we have not re-checked yet.

Regulation (EU) 2024/1689 of 13 June 2024 (Artificial Intelligence Act), Articles 6, 9, 10, 11, 13, 15, 51, 55, 56, 72 and 113, is a legal requirement only if it applies. Entry into force 1 August 2024. Article 113(a) 2 February 2025; Article 113(b) 2 August 2025 except Article 101; general application 2 August 2026; Article 113(c) Article 6(1) and the corresponding obligations from 2 August 2027. Regulation (EU) 2026/1744 is an amending regulation. The GPAI Code of Practice, including its safety and security chapter (published 10 July 2025), is Commission and AI Office material — guidance, not the regulation. NIST AI RMF 1.0 (NIST AI 100-1, January 2023) is guidance, not law. ISO/IEC 42001:2023 is a management-system standard, not the regulation. These are not a complete world list. Not legal advice.

The EU AI Act overview on this site is the pillar page. The requirements-in-force-2026 guide on this site is the Article 113 dates page. The AI-risk-management-requirements guide on this site is the Articles 9 and 55 page. The GPAI-systemic-risk guide on this site is the Article 51–55 page. The GPAI-requirements guide on this site is the Article 53 baseline page. The technical-documentation guide on this site is the Articles 11 and 53 Annex IV/XI/XII page. The AI-audit-evidence guide on this site is the evidence-request-by-role page. The how-shipreadymetrics-tracks-ai-risk guide on this site is the product page for the AI risk register. A dedicated AI-cybersecurity-requirements guide is not on this site yet. Naming it is not a link.

Frequently asked questions

Is this legal advice?

No. It is a dated map of the evaluation and adversarial-testing texts in Regulation (EU) 2024/1689 — Article 15, Articles 9(6) and 9(8), and Article 55(1)(a) — with the GPAI Code of Practice labelled as methodology, NIST AI RMF as guidance, and ISO/IEC 42001:2023 as a standard. Whether any of them applies to YOU is a legal question for counsel on your facts. This page does not start a clock and does not file with the AI Office.

Is Article 15 the same duty as Article 55(1)(a)?

No. Article 15 requires a high-risk AI system to achieve an appropriate level of accuracy, robustness and cybersecurity throughout its lifecycle, only if Article 6 high-risk applies. Article 55(1)(a) requires a provider of a general-purpose AI model with systemic risk to perform model evaluation including conducting and documenting adversarial testing, only if Article 51 applies. Different subject, different gate, different date. Do not conflate them.

Does the EU AI Act require red teaming?

Not in those words for high-risk systems. Article 55(1)(a) of Regulation (EU) 2024/1689 is the point that names adversarial testing, and it applies to providers of general-purpose AI models with systemic risk under Article 51. For high-risk AI systems, Article 15(5) requires resilience against unauthorised third parties and measures addressing data poisoning, model poisoning, adversarial examples or model evasion, confidentiality attacks and model flaws — a design-measure duty, not a named red-team exercise. Counsel applies both to your facts.

Does the GPAI Code of Practice replace Article 55(1)(a)?

No. The Code of Practice, including its safety and security chapter, is Commission and AI Office methodology and is voluntary. Article 55(2) permits relying on an Article 56 code to demonstrate compliance with Article 55(1) until a harmonised standard is published, and leaves a provider who does not adhere to an approved code to demonstrate alternative adequate means of compliance for assessment by the Commission. Guidance, not the regulation. Last verified 10 September 2026.

Does NIST AI RMF or ISO 42001 discharge Article 15?

No. NIST AI RMF 1.0 is voluntary US agency guidance, and its Measure function is a way of organising evaluation work, not law. ISO/IEC 42001:2023 is a management-system standard and an ISO 42001 certificate is not CE marking, not an Article 43 conformity assessment, and not an evaluation result. Mapping a Measure category or holding a certificate is not a determination that Article 15 or Article 55(1)(a) is met.

Did GPAI model-evaluation duties start on 2 August 2026?

No. Article 113(b) of Regulation (EU) 2024/1689 applies Chapter V, which contains Article 55, from 2 August 2025, with the exception of Article 101. High-risk Chapter III Section 2 duties, including Articles 9 and 15, sit on the original 2 August 2026 residual, and Article 113(c) keeps Article 6(1) and the corresponding obligations on 2 August 2027. Those dates are not one number, and this page does not invent a 2 August 2026 date for either track.

Does ShipReady run our model evaluations?

No. Signed-in app → Compliance → AI governance holds a model-evaluation register that records evaluations a human entered and grades each of accuracy, robustness, cybersecurity and bias as current, failing, stale, or missing for each registered high-risk system. It does not run evaluations, does not adversarially test, does not choose your metrics or thresholds, and does not certify that a result was adequate. A recorded fail is never shown as current, and a missing dimension is a disclosed gap rather than a pass. A named human still owns the assessment.

Published by ShipReady Metrics, an evidence-based technology and compliance intelligence platform. This guide is educational and vendor-neutral.