GxP Insights

EMA Annex 22: Preparing GMP Manufacturing for Probabilistic AI

EMA’s Annex 22 work places adaptive, probabilistic and generative AI at the centre of a GMP control-strategy discussion. The final guidance is not yet available; manufacturers can nevertheless prepare a defensible validation and governance approach now.

G GuideGxP 6 min read
✓ Official sources and references ✓ Practical approach ✓ For pharmaceutical professionals
GUIDEGXP · PRACTICAL GMP INSIGHTS
GuideGxP editorial illustration of quality, regulatory and operations specialists reviewing the transition to new EU veterinary GMP rules.

Artificial intelligence in GMP manufacturing is moving from a technology discussion to a quality-system question. EMA’s work on proposed Annex 22 is especially important because it directly addresses the difficult category: dynamic, adaptive and probabilistic models, including generative AI and large language models (LLMs), used in GMP applications.

The essential current position is straightforward: Annex 22 is still under development. It should not be treated as final EU GMP guidance or as a new binding set of requirements. EMA held a two-day multistakeholder workshop on 30 June and 1 July 2026 to collect expert opinion and evidence for the guidance, and stated that it expected the workshop to produce a report with expert input. The practical task for manufacturers is therefore to prepare for the questions now being tested, without claiming that the eventual Annex requires controls that have not yet been published.

What EMA has put on the table

EMA reports that its 2025 stakeholder consultation on the draft Annex 22 indicated support for potentially enabling technologies such as GenAI and LLMs in medicines manufacturing. The earlier draft had indicated that dynamic, adaptive and probabilistic models should not be used in critical GMP applications. EMA is now considering the implications of consultation feedback and has sought expert input on risk-based control and mitigation measures, including guardrails.

That is a significant regulatory development. It does not mean that such models are now accepted for critical GMP use. It means EMA is exploring whether, and under which control strategy, some of these technologies could be accommodated. The distinction matters for investment decisions, validation planning and internal communication.

EMA’s workshop questions concentrate on six connected areas:

  • regulatory pathways for adaptive and probabilistic AI in GMP;
  • technical reliability, including hallucinations, incorrect recommendations and fabricated data;
  • human oversight and accountability;
  • validation and lifecycle management, including updates, retraining and drift;
  • strategic risk and compliance limits for high-risk GMP uses; and
  • cybersecurity and outsourced activities, including cloud-based AI supply chains.

For QA, validation, data integrity and IT/OT leaders, the message is not to wait for a document title. The message is to make the intended use, failure modes, evidence and decision rights visible before an AI tool enters a GMP workflow.

The central challenge: validation where output is probabilistic

Traditional computerised-system validation is often built around predictable inputs, predefined processing and expected outputs. Probabilistic AI changes the nature of the assurance case. The same prompt or operational context may not reliably generate the same output; a model may change through supplier updates, retraining or interaction with changing data; and apparently plausible output may still be wrong.

EMA specifically asks what validation paradigm should apply to adaptive models and what a quality risk management approach should consider across the full lifecycle, particularly in GMP high-risk areas where deficiencies may have low detectability but high patient impact. It also asks what validation data, stress-testing results and failure analyses would justify reliance on guardrails as a mitigation measure.

Regulatory fact: these are questions submitted for expert consideration, not final prescribed validation requirements.

GuideGxP recommendation: frame validation as an evidence-based assurance case rather than a one-time test protocol. The assurance case should demonstrate that the system’s intended use is controlled, its limitations are understood, unacceptable outputs are either prevented or detected before GMP impact, and the system remains within its approved operating conditions over time.

A practical validation strategy

A proportionate strategy can be organised around the following sequence:

  1. Define the GMP decision boundary. State exactly whether the AI drafts information, recommends an action, prioritises review work, detects an anomaly, controls a process or makes a decision. Identify who can accept, reject or override the output.
  2. Classify impact and detectability. Assess the possible patient, product-quality, data-integrity and compliance consequences of failure. Give special weight to errors that are difficult for a reviewer to identify.
  3. Specify approved use conditions. Define permitted inputs, source data, users, locations, interfaces, output formats, prohibited use cases and escalation triggers. A system cannot be validated meaningfully if its use boundary is vague.
  4. Create a representative challenge set. Test normal cases, edge cases, incomplete records, conflicting evidence, malicious or anomalous inputs, ambiguous instructions and known failure patterns. Retain the rationale for coverage and acceptance criteria.
  5. Test the control strategy, not only the model. Evaluate guardrails, workflow restrictions, reviewer interfaces, audit trails, alarms, access controls and stop/escalation mechanisms as an integrated system.
  6. Establish lifecycle evidence. Define how performance, incidents, model changes, drift indicators and supplier changes will be assessed. Predefine when revalidation, reassessment or withdrawal from use is required.

Guardrails are controls, not a substitute for accountability

EMA asks whether available guardrails can reliably prevent, detect or contain hallucinations, incorrect recommendations and fabricated data in GMP workflows. It also asks what happens when guardrails fail or identify uncertainty, and what human-in-the-loop oversight remains necessary once such controls are implemented.

This is the right emphasis. A guardrail may reduce a defined risk, such as restricting a model to approved source content or blocking particular output types. It does not automatically establish that an output is accurate, complete, attributable or suitable for GMP decision-making. Nor does it remove the need to assign responsibility for the decision and for the AI system’s continuing state of control.

GuideGxP recommendation: every guardrail should have an owner, a documented purpose, defined failure behaviour and testable acceptance criteria. The quality system should answer four practical questions:

  • What risk does this guardrail mitigate?
  • What evidence shows it works under expected and adverse conditions?
  • How will its degradation or failure be detected?
  • Who acts, how quickly, and with what authority when it fails or output is uncertain?

For higher-risk use cases, human oversight should be designed rather than assumed. A nominal reviewer who lacks source access, sufficient time, relevant expertise or authority to reject the output is not an effective control. Review must be meaningful, traceable and commensurate with the risk of an undetected error.

Data integrity and traceability must be designed into the workflow

EMA identifies data governance, model evaluation, transparency, accountability and human oversight as key principles for the Annex 22 discussion. It also asks whether guardrail approaches sufficiently address data integrity, process control, auditability and traceability where generative AI may influence or automate decision-making.

In practice, AI governance must preserve the evidential chain around a GMP-relevant output. The organisation should be able to reconstruct what was submitted to the system, which approved configuration and model version were used, what information was available to the model, what output was generated, what controls were applied, who reviewed it and what final action was taken.

GuideGxP recommendation: do not treat an AI interface as an isolated productivity tool. Map it as part of the GMP process and data flow. This exposes where records are created, transformed, retained, transferred, reviewed and approved, and where the audit trail or accountability chain could break.

Change control must address the model supply chain

EMA’s questions explicitly include model evolution through updates, retraining and drift, as well as risks where guardrail infrastructure sits outside the manufacturer’s direct control. The Agency asks how supplier qualification, change-control visibility and independent audit capability can be preserved in cloud-based AI supply chains.

That focus is critical for externally hosted models. A supplier’s routine model update may change output quality, behaviour, security characteristics or the effectiveness of a previously validated guardrail. For GMP use, the relevant question is not simply whether a supplier is reputable. It is whether the manufacturer has adequate knowledge, contractual leverage, technical controls and evidence to keep the intended use in a validated state.

Control areaGuideGxP practical question
Supplier oversightCan the organisation understand and assess relevant model, infrastructure and guardrail changes before they affect GMP use?
Configuration controlCan it prove which model, version, parameters, knowledge sources and workflow rules generated a specific output?
Change managementAre changes triaged by GMP impact, with defined verification or revalidation before release?
Operational monitoringAre performance signals, incidents, uncertainty events and user overrides reviewed for emerging loss of control?
Exit and continuityCan GMP records, evidence and critical workflow capability be retained or transferred if the service changes or becomes unavailable?

What senior management should decide now

Senior management should avoid two unhelpful extremes: prohibiting all AI without examining the use case, or permitting broad experimentation that quietly becomes operational reliance. A portfolio approach is more defensible.

Start with an inventory of AI already used or proposed across manufacturing, laboratories, quality operations, engineering, supply chain and corporate functions that support GMP work. Separate low-impact drafting or administrative assistance from systems that influence GMP records, investigations, release decisions, process control or product-quality conclusions. Then require a clearly owned risk assessment and an approval pathway before any use case crosses into GMP-relevant activity.

EMA’s Annex 22 discussion may ultimately clarify what risk-based accommodation of adaptive, probabilistic and generative models could look like. Until final guidance is available, organisations that can demonstrate disciplined intended-use definition, transparent evidence, effective oversight and controlled lifecycle management will be best placed to assess and adopt it responsibly.

Sources

THE PRAGMATIC GMP · EVERY MONDAY

The GMP topics that matter, in 7 minutes.

One GMP topic, one real-world example and one practical action, based on official sources and inspection trends.
Discover The Pragmatic GMP