A new commentary published by Bio-IT World on 11 September 2026 addresses one of the most difficult questions facing pharmaceutical AI:
How can Generative AI be used in GMP when its outputs are inherently probabilistic rather than fully deterministic?
The article, written by Valentina Armiento and Pierre Bonzani of Genedata, argues that GenAI cannot simply be validated in the same way as conventional GMP software.
Traditional computerized systems are generally expected to behave predictably under defined conditions.
Generative AI is different.
Its output may vary depending on:
- the prompt;
- the surrounding context;
- model configuration;
- model version;
- sampling parameters;
- changes to the underlying model.
This makes repeatability, traceability and lifecycle control considerably more difficult.
The authors therefore argue that GenAI requires a different validation mindset built around:
bounded intended use → risk-based controls → traceability → human oversight → controlled model lifecycle → performance monitoring
Why GenAI Is Different from Conventional GMP Software
The fundamental issue is non-determinism.
Traditional GMP software typically executes predefined logic.
When given the same input under the same configuration, the expected result should normally be reproducible.
Large Language Models operate differently.
They generate outputs probabilistically.
Even when the underlying model has not changed, variations in prompts, context or generation parameters may influence the result.
This creates a validation challenge.
The relevant question is no longer only:
Does the system produce the expected output?
but increasingly:
Does the system consistently remain within an acceptable performance range for its intended use?
That distinction may be fundamental for future GenAI validation in GMP.
The Current Annex 22 Position Remains Restrictive
The current draft EU GMP Annex 22 takes a conservative position.
It applies to static AI/ML models used in critical GMP applications and states that dynamic models that continuously learn during use should not be used in critical GMP applications.
The draft also states that probabilistic models that may produce different outputs from identical inputs are outside its scope and should not be used in critical GMP applications.
It explicitly includes Generative AI and Large Language Models in this restriction.
For non-critical GMP applications, however, GenAI may be considered where qualified personnel remain responsible for ensuring that the output is suitable for its intended use.
In practical terms, the current draft creates a clear distinction between:
critical GMP use
and
non-critical GMP support under human oversight.
Official draft Annex 22:
https://health.ec.eu … ion_guideline_en.pdf
But the Regulatory Discussion Is Still Evolving
The position above should not be interpreted as the final regulatory endpoint.
EMA held a multistakeholder workshop on 30 June and 1 July 2026 specifically to explore whether adaptive and probabilistic AI models could potentially be accommodated within a future risk-based Annex 22 framework.
The workshop examined topics such as:
- guardrails;
- human oversight;
- model evaluation;
- data governance;
- accountability;
- transparency;
- risk mitigation;
- validation approaches for adaptive and probabilistic models.
This is important.
The regulatory question is gradually shifting from:
“Should probabilistic AI simply be excluded?”
towards:
“Can probabilistic AI be controlled sufficiently for a defined GMP use?”
EMA Annex 22 workshop:
https://www.ema.euro … development-annex-22
Six Practical Controls Highlighted in the Bio-IT World Commentary
The article identifies several elements that should form part of a GMP-oriented AI control strategy.
1. Clearly Defined Intended Use
The role of the AI should be defined precisely.
The organisation should know:
- what task the AI performs;
- which decisions it supports;
- which data it receives;
- under which conditions it may be used;
- where its boundaries are.
This is critical because validation can only be meaningful when performance is assessed against a clearly defined use case.
2. Risk-Based Human Oversight
The level of AI autonomy should depend on the risk of the activity.
For high-impact decisions, the AI should remain advisory and appropriately qualified personnel should retain responsibility for the final decision.
The important question is not simply whether a human is formally present in the workflow.
It is whether the human review is capable of detecting and correcting an unacceptable AI output.
3. Data and Decision Traceability
Inputs, outputs, model versions and relevant user interactions should be traceable.
For GenAI this can potentially include:
- source data;
- prompt or query;
- model version;
- configuration;
- generated output;
- reviewer;
- final accepted result.
This may become especially important when AI contributes to GMP documentation, investigations or other quality-system activities.
4. Explainability and Confidence
The authors also highlight explainability mechanisms.
For some AI systems this may involve techniques such as feature attribution or confidence scores.
For GenAI, explainability may be more difficult.
In many cases, a more practical control may be strong linkage between the generated answer and controlled source information so that the reviewer can verify the basis of the output.
5. Integration with Existing GMP Controls
AI should not operate outside the Pharmaceutical Quality System.
It should be incorporated, as applicable, into:
- change control;
- validation;
- configuration management;
- supplier management;
- periodic review;
- deviation management;
- performance monitoring.
This is an important point.
GenAI does not require abandoning established GMP controls.
It requires extending them to address new failure mechanisms.
6. Cross-Functional Governance
AI validation cannot be owned by one function.
Effective governance may require cooperation between:
- Quality Assurance;
- IT;
- Data Science;
- Validation;
- system and process owners;
- subject-matter experts.
This is consistent with the draft Annex 22 approach, which explicitly calls for cooperation between process SMEs, QA, data scientists, IT and other relevant parties.
The Real Problem: Model Updates Can Break the Validated State
One of the most important points in the Bio-IT World article concerns model updates.
For traditional systems, software changes are generally identifiable and controlled.
With GenAI, behaviour may change because:
- the underlying model is replaced;
- the provider updates the model;
- system prompts change;
- retrieval data changes;
- configuration changes;
- context handling changes.
Even where the visible application remains unchanged, the output behaviour may be different.
This creates a difficult lifecycle problem:
How can innovation continue without repeatedly destroying the validated state?
The article proposes several mechanisms:
- version-locked models;
- formal change control;
- predetermined update strategies;
- risk-based revalidation;
- continuous performance monitoring.
From Exact Repeatability to Acceptable Performance
Perhaps the most important conceptual idea in the article is that validation may need to move away from expecting identical text outputs.
For GenAI, the relevant acceptance criterion may instead be whether different acceptable outputs remain within predefined quality boundaries.
For example, performance could potentially be evaluated using:
- correctness;
- completeness;
- critical-error rate;
- hallucination rate;
- source-grounding;
- human correction rate;
- failure-to-escalate rate;
- performance on predefined challenge cases.
The objective would therefore not necessarily be:
same input → identical wording
but rather:
defined intended use → controlled variability → acceptable and measurable performance.
This represents a major conceptual shift from classical software testing.
Continuous Monitoring Becomes Part of Validation
The article also argues that AI performance should continue to be monitored after deployment.
This is closely aligned with the current draft Annex 22, which requires regular monitoring of model performance and input-data drift for AI models within its scope.
For future GenAI applications, similar principles may become essential.
A lifecycle approach could therefore look like:
Initial Evaluation
↓
Acceptance Testing
↓
Controlled Deployment
↓
Performance Monitoring
↓
Drift / Failure Detection
↓
Change Assessment
↓
Revalidation Where Required
This begins to resemble a maintained AI state of control rather than a one-time validation event.
A Practical Place for GenAI Today
The commentary identifies documentation and knowledge workflows as particularly promising areas.
Examples could include:
- summarising controlled information;
- knowledge search;
- drafting documentation;
- supporting investigation review;
- querying validated datasets;
- assisting with scientific or technical interpretation.
These use cases are more suitable when the AI does not make the final critical GMP decision and qualified personnel remain responsible for reviewing the result.
This remains very different from allowing an LLM to autonomously:
- release a batch;
- make a critical product-quality decision;
- change manufacturing parameters;
- approve a deviation;
- close a CAPA.
Why This Article Matters
The significance of this commentary is not that it creates a new regulatory framework.
It does not.
Its value is that it clearly describes the fundamental problem that the industry now needs to solve:
GMP validation was historically designed around deterministic systems, while Generative AI is inherently probabilistic.
Trying to force GenAI into an unchanged deterministic validation model may therefore be unrealistic.
But abandoning validation principles is equally unacceptable.
The likely direction is a middle ground based on:
bounded use → measurable performance → controlled variability → traceability → human oversight → change control → continuous monitoring.
This may ultimately be one of the key concepts determining whether Generative AI can move from GxP-adjacent experimentation into genuinely regulated pharmaceutical workflows.
Regulatory Status Note
The Bio-IT World article is contributed commentary written by representatives of Genedata. It is not regulatory guidance and does not represent the official position of FDA, EMA or another competent authority.
The discussion of Annex 22 should also be understood in the context of an evolving regulatory process. The current draft excludes GenAI and probabilistic models from critical GMP applications, but EMA is actively considering stakeholder input on whether future risk-based controls and guardrails could support broader use.
Sources
Bio-IT World – What It Takes to Make Generative AI Fit for GMP, published 11 September 2026:
https://www.bio-itwo … ative-ai-fit-for-gmp
European Commission – Draft EU GMP Annex 22: Artificial Intelligence:
https://health.ec.eu … ion_guideline_en.pdf
EMA – Multistakeholder workshop on AI guidance development for Annex 22: