Saturday, July 11, 2026

EMA Annex 22 AI workshop: awaiting official conclusions, but the direction of discussion is important


Introduction EMA’s Annex 22 AI workshop highlights a possible shift from prohibiting GenAI and LLMs in GMP toward risk-based control using guardrails, human oversight, traceability, validation, lifecycle monitoring and pharmaceutical quality system governance.

Following the EMA GMP multistakeholder workshop organised on 30 June and 1 July 2026 on expert contributions to the development of EU GMP Annex 22 on Artificial Intelligence, the pharmaceutical industry is still awaiting official feedback, conclusions or a post-workshop report from EMA.

The workshop was organised to support the development of Annex 22 and to collect expert input on the possible use of artificial intelligence in medicines manufacturing. The first day was available as a live broadcast and included expert opinions on the future use of generative AI, large language models and other probabilistic or adaptive AI models in GMP environments.

One important message emerging from the expert discussion was that a simple, general prohibition of GenAI or LLMs in pharmaceutical GMP environments may not be the most effective or future-proof regulatory approach. Such a prohibition could quickly become outdated, especially as AI models, control mechanisms, guardrails, validation methods and monitoring tools continue to improve.

At the same time, this does not mean that GenAI or LLMs should be freely accepted in critical GMP applications. The more balanced and practical direction appears to be a risk-based approach: AI should be considered according to its intended use, GMP impact, patient risk, level of human oversight, data quality, traceability, model behaviour, validation evidence and lifecycle controls.

This is particularly important because the original draft Annex 22 indicated that dynamic, adaptive and probabilistic models, including GenAI and LLMs, should not be used in critical GMP applications. However, stakeholder feedback showed support for potentially enabling these technologies in medicines manufacturing if adequate control and mitigation measures can be demonstrated.

For pharmaceutical companies, the practical message is clear: the discussion is moving from “AI should be prohibited” toward “under what conditions can AI be controlled well enough for GMP use?”

The key control areas are likely to include:

  • clear definition of intended use;
  • GMP impact and patient-risk assessment;
  • approved and controlled source data;
  • model/version control;
  • guardrails and their verification;
  • human oversight and accountability;
  • traceability of AI outputs;
  • detection of hallucinations or incorrect recommendations;
  • incident escalation and prevention of GMP impact;
  • performance monitoring and drift detection;
  • supplier qualification and cloud-service oversight;
  • change control for model updates, retraining and configuration changes;
  • documented validation or qualification evidence.

One of the most important concepts discussed in the context of GenAI and LLMs in GMP is the use of “guardrails”. What could “guardrails” mean for AI in GMP?

In simple terms, guardrails are predefined controls built around an AI system to keep its use within safe, intended and acceptable boundaries.

In GMP language, guardrails can be understood as technical, procedural and human controls that prevent AI from being used in the wrong way, reduce the risk of incorrect or unsupported outputs, and stop AI-generated conclusions from becoming GMP decisions without appropriate human review.

Guardrails do not make AI perfect. They do not remove the need for validation, qualification, human oversight or quality risk management. Their purpose is to ensure that AI remains a controlled support tool, not an uncontrolled authority.
Guardrails are predefined technical, procedural and organizational controls that define what an AI system is allowed to do, what it must not do, how it should behave under uncertainty, and how its outputs are reviewed before they affect GMP activities, records or decisions.

Examples of AI guardrails in GMP

  • Intended-use limitation – the AI system may only be used for defined and approved purposes. For example, an AI tool may summarize SOP content, but may not approve a deviation, assign final root cause or make a batch disposition decision.
  • Source grounding – the AI must generate answers only from approved and controlled sources, such as current SOPs, specifications, batch records, validation reports, quality agreements or approved regulatory guidance.
  • Refusal rules – the AI must not guess when evidence is missing. For example, the system should state: “Insufficient information is available to conclude product impact” instead of generating an unsupported conclusion.
  • Output constraints – the AI output should be limited to predefined sections or fields, such as facts identified, source references, missing information, potential questions for review and recommended human follow-up. It should not create final GMP conclusions unless this is explicitly validated and approved for the intended use.
  • Human review gate – AI-generated outputs should be reviewed by qualified personnel before they are used in GMP records or decisions. For example, QA should review an AI-generated deviation summary before it is entered into the official investigation record.
  • Access control – only trained and authorized users should be able to use AI functions with potential GMP impact. For example, a batch review assistant should be available only to trained QA or manufacturing reviewers.
  • Audit trail – the system should retain evidence of AI use, including prompt, output, source documents, model or knowledge-base version, reviewer edits and final approval.
  • Change control – changes to the AI model, prompts, configuration, guardrails or knowledge base should be assessed for GMP impact. A new model version or major knowledge-base update may require impact assessment and requalification.
  • Performance monitoring – AI performance should be monitored in routine use. Examples include tracking false citations, unsupported claims, reviewer corrections, repeated failure modes and cases where AI output was rejected by users.

Practical example: AI support for deviation investigations

For an AI tool supporting deviation investigations, guardrails could be defined as follows:

The AI may identify relevant facts, summarize the event chronology, list missing information, identify potentially similar historical deviations and suggest questions for the investigator to consider. The AI must use only approved QMS records and controlled source documents. It must not assign the final root cause, conclude product impact, determine CAPA effectiveness or recommend batch disposition. If the available information is insufficient, the AI must clearly state that no conclusion can be made. Any AI-generated text must be reviewed, corrected where needed and approved by the investigation owner and QA before inclusion in the GMP record.

This example illustrates the practical meaning of guardrails: the AI can support the process, but it cannot replace GMP responsibility. The final decision remains with qualified personnel and must be justified by evidence.

The final Annex 22 position is not yet known. However, the workshop confirms that regulators are actively considering how to balance innovation with GMP control. For industry, this is a strong signal to start preparing practical AI governance frameworks now, rather than waiting until the final Annex 22 text is published.

AI in GMP should not be treated as an informal tool or uncontrolled black box. If AI supports or influences regulated manufacturing, quality or compliance decisions, it should be governed within the pharmaceutical quality system.

Previous note on the planned EMA workshop:
https://www.aiforpha … ment-of-ai-annex-22/

EMA workshop page:
https://www.ema.euro … development-annex-22