GxP Insights

PIC/S draft guidelines for new GMP Annex 22 on Artificial Intelligence

PIC/S has published draft guidelines for a new Annex 22: Artificial Intelligence. The proposal sets out a lifecycle-oriented framework for static, deterministic AI/ML models used in critical GMP applications, with implications for intended use, test-data governance, validation, explainability, change control and ongoing monitoring.

G GuideGxP 6 min read
✓ Official sources and references ✓ Practical approach ✓ For pharmaceutical professionals
GUIDEGXP · PRACTICAL GMP INSIGHTS
Clean colour GuideGxP 2D editorial comic illustrating the GMP topic: PIC/S draft guidelines for new GMP Annex 22 on Artificial Intelligence.

PIC/S has published “Draft guidelines: New annex 22 - Artificial intelligence” within the PIC/S GMP Guide section of its publications. The proposed Annex 22: Artificial Intelligence is a new annex, not a revision of an existing one. It is therefore an important regulatory development for organisations using, procuring or assessing AI/ML-enabled computerised systems in GMP operations.

The document remains a draft. It should not be treated as a final, independently binding GMP requirement or as evidence of a confirmed implementation timetable. Its practical importance lies in the level of specificity it brings to the management of AI models used in critical GMP applications, particularly where outputs can directly affect patient safety, product quality or data integrity.

The draft is explicitly positioned as additional guidance to Annex 11 for computerised systems in which AI models are embedded. For senior quality, validation, IT and manufacturing leaders, the central message is clear: AI assurance is proposed as more than conventional CSV applied to a novel algorithm. It requires demonstrable control of intended use, data, model performance, human roles and operational change throughout the model lifecycle.

What the PIC/S draft covers

The proposed scope covers computerised systems used in the manufacture of medicinal products and active substances where AI models are used in critical applications with direct impact on patient safety, product quality or data integrity. Examples given include the prediction or classification of data.

Its focus is machine learning models whose functionality is obtained through training with data rather than explicit programming. A model may comprise several individual models which automate specific process steps in GMP. The draft therefore has potential relevance to a broad range of use cases, including automated visual inspection, classification of manufacturing or laboratory data, and predictive tools embedded in GMP decision processes.

However, the scope is deliberately constrained. The draft applies to static models that do not adapt their performance during use by incorporating new data, and to models with deterministic output. Dynamic models that continuously and automatically learn during use, and probabilistic models that may provide different outputs for identical inputs, are not covered and are stated not to be suitable for critical GMP applications.

The document also states that it does not apply to Generative AI and Large Language Models (LLM), and that such models should not be used in critical GMP applications. Where they are used in non-critical GMP applications without direct impact on patient safety, product quality or data integrity, adequately qualified and trained personnel should remain responsible for deciding whether outputs are suitable for intended use. This is a human-in-the-loop (HITL) expectation, not a blanket approval of Generative AI in GMP work.

The proposed control model: intended use before technology

Annex 22 starts from the GMP principle that the intended use must be understood and defined. The proposed guidance expects a detailed description of the task that the model assists or automates, grounded in in-depth knowledge of the process in which it is integrated. It also expects comprehensive characterisation of the input data, including common and rare variations, limitations, and potentially erroneous or biased inputs.

A process subject matter expert (SME) should be responsible for the adequacy of this description, which should be documented and approved before acceptance testing begins. Where relevant, the input sample space should be divided into subgroups. These can reflect decision outcome, site or equipment baseline, material or product characteristics, or task-specific characteristics such as defect type and severity.

This is significant because it moves the discussion beyond generic statements that an algorithm is “accurate”. The proposed approach requires organisations to define for which process, population, conditions and edge cases the model is acceptable. A favourable overall accuracy metric may be inadequate if critical subgroups perform poorly.

Validation evidence: metrics, acceptance criteria and independent test data

The draft proposes case-dependent metrics aligned to intended use. For a classification model, examples include a confusion matrix, sensitivity, specificity, accuracy, precision and F1 score. Acceptance criteria should be established and approved before testing, with responsibility assigned to a process SME. Criteria may differ between relevant subgroups.

A particularly important proposed principle is that the model’s acceptance criteria should be at least as high as the performance of the process it replaces. This makes a robust baseline indispensable: an organisation should understand the existing manual or automated process performance before claiming that an AI model represents an acceptable replacement.

The proposed requirements for test data are detailed. Test data should represent, and expand, the full sample space of intended use; be stratified; include relevant subgroups; and reflect limitations, complexity, and common and rare variations. Dataset size should support calculation of test metrics with adequate statistical confidence. Labelling should be verified through a process that delivers a very high degree of correctness, potentially using independent experts, validated equipment or laboratory tests.

Test-data independence is also a prominent control. The draft proposes technical and/or procedural measures to ensure that data used for final testing were not used during model development, training or validation. When a test set is split before training, personnel involved in development and training should not have access to it. The test data should be protected with access control and audit trail functionality, with no copies outside the controlled repository. The guidance also addresses staff independence and identifies a 4-eyes principle as a possible mitigation where full separation cannot be maintained.

Explainability, confidence and operation

For models used in critical GMP applications, the draft proposes capture and recording of features that contributed to a classification or decision during testing. Where applicable, feature-attribution techniques such as SHAP values or LIME, and visual tools such as heat maps, should highlight key factors contributing to an outcome. Review of those features should form part of test-result approval, on a risk basis.

The proposal also addresses confidence scores and thresholds. Where applicable, prediction or classification models should log confidence scores. A model should be configured so that it makes an outcome only when confidence is suitable; for very low confidence, an “undecided” outcome may be more appropriate than an unreliable prediction or classification.

Once deployed, the model, the computerised system and the process being assisted or automated should be under change control. Changes to the model, system, process or relevant physical input objects should be evaluated for retesting. The draft also proposes configuration control, detection of unauthorised change, regular performance monitoring and monitoring that input data remain within the defined sample space and intended use. This last point is directly relevant to data drift and environmental or process changes, such as changed lighting conditions in an inspection application.

GuideGxP recommendation: prepare now, without overstating the draft

Organisations do not need to wait for a final Annex 22 to establish a defensible AI governance baseline. The appropriate action is not to retrospectively label every advanced analytical tool as AI, nor to introduce an unstructured “AI policy”. Instead, perform a focused, risk-based assessment of systems that use trained models and that may influence GMP decisions or critical data.

Readiness areaPractical GuideGxP recommendation
Use-case inventoryIdentify AI/ML models in GMP processes, including supplier-provided functionality, and classify whether their outputs have a direct impact on patient safety, product quality or data integrity.
Intended useEstablish approved intended-use statements, decision boundaries, input sample space, known limitations, failure modes and accountable process SMEs.
Data governanceMap training, validation and test datasets; preserve provenance, labelling rationale, access control, audit trails, independence and retention.
Validation strategyDefine risk-appropriate metrics and subgroup acceptance criteria before final testing; document the performance baseline for any process being replaced.
Operational controlBring the model, its configuration, interfaces and associated process under change control; define performance, drift and input-boundary monitoring.
Human oversightSpecify the operator’s decision responsibility and evidence of training and consistent performance wherever HITL is used.

QA should ensure that accountability is not fragmented between data science, IT/CSV, manufacturing and suppliers. The draft expects close co-operation among relevant parties during algorithm selection, training, validation, testing and operation, together with appropriate qualifications, defined responsibilities and suitable access. The regulated user is expected to have and review the relevant documentation even when training, validation or testing is performed by a supplier or service provider.

Key takeaway: proposed Annex 22 is a focused framework for demonstrating that a critical AI/ML model is fit for its defined GMP use, tested independently, explainable to an appropriate degree, controlled in operation and monitored against drift.

What to watch next

The PIC/S publications page currently presents “Draft guidelines: New annex 22 - Artificial intelligence” as a draft document in the PIC/S GMP Guide. Organisations should monitor PIC/S communications for consultation, revision, finalisation and any accompanying implementation information. Until then, internal readiness work should be managed through the existing PQS, QRM, supplier management, CSV, data integrity and change-control frameworks, while clearly distinguishing current GMP obligations from measures taken in anticipation of the draft.

Official sources

THE PRAGMATIC GMP · EVERY MONDAY

The GMP topics that matter, in 7 minutes.

One GMP topic, one real-world example and one practical action, based on official sources and inspection trends.
Discover The Pragmatic GMP