What would a validation strategy for probabilistic AI look like under GMP? It would need to show, proportionately to patient and product risk, that the intended use is defined, the system’s failure modes are understood, safeguards are demonstrably effective, accountable human oversight is workable, and performance remains controlled through updates, retraining and drift. This is not yet a final Annex 22 requirement: EMA’s June–July 2026 workshop was designed to gather expert evidence for guidance development. But its published questions provide a practical indication of the issues manufacturers should be ready to address.
For QA, Validation, Data Integrity, IT/OT and manufacturing leaders, the most useful response is not to wait for a final text. It is to separate low-risk experimentation from GMP use, establish an evidence-led AI governance model, and test whether proposed controls would stand up when the model is wrong, uncertain, changed or outside the manufacturer’s direct control.
What EMA has said about Annex 22
EMA describes Annex 22 as EU guidance on the use of artificial intelligence in medicines manufacturing. Its two-day multistakeholder workshop took place on 30 June and 1 July 2026. The first day was an open session for expert opinions and evidence; the second was a closed session in which EMA’s Annex 22 drafting group reviewed contributions. EMA stated that it expected the workshop to produce a report with expert input.
According to EMA, a 2025 stakeholder consultation on draft Annex 22 suggested support for potentially enabling technologies such as generative AI and large language models in medicines manufacturing. The draft had indicated that dynamic, adaptive and probabilistic models, including GenAI and LLMs, should not be used in critical GMP applications. EMA said it was still considering the consultation’s implications and was seeking expert input on possible risk controls and mitigations, including guardrails, within a proposed risk-based approach.
Regulatory fact: the published workshop page asks questions; it does not itself establish that any particular guardrail, validation method or human-review model is acceptable. GuideGxP interpretation: the questions make clear that a generic software-validation package or a vendor assurance statement alone is unlikely to be a sufficient rationale for using probabilistic AI in a high-risk GMP workflow.
Why probabilistic AI changes the validation conversation
Traditional GMP computerised-system validation generally concentrates on demonstrating that a specified system performs consistently for its intended use. Probabilistic AI adds a different operational challenge: the same prompt or input context may not always yield identical outputs, and a model can generate plausible but incorrect content. For adaptive systems, behaviour may also change after retraining, model updates or data drift.
That does not make all AI unusable. It means the validation claim must be precise. A manufacturer should not attempt to validate an AI model as universally “accurate” or “safe”. Instead, it should validate a controlled use case, its input boundaries, its permitted output, the decision rights retained by people, and the technical and procedural controls that prevent or detect unacceptable outcomes.
EMA’s questions connect this directly to GMP consequences: hallucinations, incorrect recommendations, fabricated data, incident escalation, data integrity, process control, auditability, traceability, cybersecurity, outsourced infrastructure and the feasibility limits of mitigation. This is a whole-system problem, not solely a model-performance problem.
A practical validation strategy for AI used in GMP
1. Define the intended use and the prohibited use
Start with a narrowly written intended-use statement. It should identify the GMP process, user population, inputs, outputs, interfaces, decision being supported and the maximum degree of automation. Equally important, specify what the AI must not do.
For example, an AI tool that drafts a deviation investigation summary for QA review is not the same use case as a tool that determines root cause, approves a batch disposition decision or changes a process-control parameter. The latter cases may have materially higher impact, lower detectability of deficiencies and much less tolerance for residual uncertainty.
2. Perform risk assessment across the full lifecycle
EMA explicitly asks how quality risk management should apply throughout the lifecycle of AI models in GMP high-risk areas, including situations with high patient impact and low detectability of deficiencies. The risk assessment should therefore go beyond a one-off implementation assessment.
| Lifecycle stage | Questions to answer | Evidence to retain |
|---|---|---|
| Design and selection | What task is being supported? Why is AI necessary? What errors are foreseeable? | Intended use, risk assessment, supplier assessment, architecture description |
| Implementation | Are inputs controlled? Are outputs constrained? Can the system access GMP records or execute actions? | Configuration specification, interface testing, access-control evidence |
| Validation | Can the system and its safeguards detect or prevent unacceptable outputs within the intended use? | Test protocol, challenge set, stress tests, failure analysis, acceptance rationale |
| Operation | Who reviews outputs? How are uncertainty, exceptions and incidents escalated? | SOPs, training records, audit trails, review records, deviation handling |
| Change and monitoring | How are updates, retraining, drift and supplier changes assessed and controlled? | Change control, monitoring trends, periodic review, revalidation rationale |
3. Treat guardrails as controls that require evidence
EMA asks to what extent guardrails can prevent, detect or contain hallucinations, incorrect recommendations and fabricated data, and what evidence would justify relying on them. This framing matters. A guardrail is not validated merely because it exists in a product demonstration or vendor documentation.
Test the guardrail against realistic failure conditions: ambiguous source material, incomplete records, conflicting information, prohibited requests, adversarial inputs, unexpected volume, unavailable dependencies and known model weaknesses. Establish what the guardrail does when it cannot provide a reliable response. A safe outcome may be refusal, routing to a qualified reviewer, or stopping an automated downstream action.
Acceptance criteria should be connected to the GMP risk. Where incorrect output could influence a critical decision, the system should not rely on a user noticing every error. The control design must make failure detectable and containable before it creates GMP impact.
4. Make human oversight specific and accountable
“Human in the loop” is not, by itself, a control strategy. EMA asks what level and form of oversight remains necessary after guardrails are implemented, and whether that oversight ensures accuracy, traceability and accountability.
A defensible design identifies the named role responsible for review, the information needed to make that review meaningful, the permitted response time, the circumstances requiring escalation and the record created by the decision. Reviewers need access to relevant source data and a way to distinguish AI-generated content from verified evidence. They also need training that covers the tool’s limitations, not only its operating steps.
5. Build change control around model evolution and supply chains
EMA’s workshop questions specifically raise updates, retraining and drift, as well as cloud-based AI supply chains where guardrail infrastructure sits outside the manufacturer’s direct control. These are not peripheral IT issues. They determine whether the validated state can be maintained.
Manufacturers should define which changes require assessment before implementation: base-model version changes, prompt or orchestration changes, retrieval-source changes, guardrail updates, infrastructure changes, changes to access permissions and new or altered interfaces. Where a supplier cannot provide adequate notification, change-control visibility or audit capability, that limitation should influence the intended use and risk decision.
Operational checklist for GMP teams
- Classify every proposed AI use case by GMP impact, detectability of error and degree of autonomy.
- Write an intended-use statement and explicit prohibited-use boundaries before configuration or testing.
- Map data flows, including sources, prompts, outputs, audit trails, retention and access permissions.
- Identify failure modes, including hallucination, fabricated data, incorrect recommendation, drift and unauthorised change.
- Define guardrails and test their effectiveness with normal, boundary and failure-condition scenarios.
- Document what happens when the system is uncertain, unavailable or outside validated operating conditions.
- Assign accountable human review roles, decision rights and escalation routes.
- Place model, configuration, knowledge-source and supplier changes under formal change control.
- Establish ongoing monitoring, periodic review and criteria for suspension or retirement.
- Ensure supplier qualification addresses security, tampering protection, auditability and outsourced-control limitations.
What not to conclude yet
It would be premature to state that EMA has authorised GenAI or LLMs for critical GMP applications, or that Annex 22 will definitely permit them if guardrails are used. The official workshop page says EMA is considering consultation implications and collecting expert input on possible safeguards and risk mitigations. Final expectations, scope and wording should not be assumed from consultation and workshop questions alone.
The sensible planning assumption is narrower: where a manufacturer proposes adaptive, dynamic or probabilistic AI for a GMP-relevant purpose, it should expect scrutiny of validation evidence, lifecycle management, human accountability, data integrity, traceability, supplier control and the point at which residual risk makes the use case unsuitable.
FAQ
Is Annex 22 final GMP guidance?
The EMA event page describes guidance development and an expert workshop. It does not present a final Annex 22 text or finalised requirements.
Can GenAI be used in GMP today?
The workshop page does not provide a blanket yes-or-no answer. Its questions focus on whether, and under what controls, dynamic, adaptive and probabilistic models could be accommodated in GMP applications.
Are guardrails enough for a critical GMP use case?
EMA explicitly asks whether guardrails sufficiently address high-risk issues and whether some critical decisions remain unsuitable despite guardrails and oversight. A manufacturer should therefore demonstrate effectiveness for its specific use case rather than assume that guardrails are sufficient.
Primary source
For practical GMP analysis delivered to your inbox, sign up to The Pragmatic GMP.