Friday, September 11, 2026

Making Generative AI Fit for GMP: Validation Must Move Beyond Determinism


Introduction Generative AI challenges traditional GMP validation because probabilistic models cannot always provide identical outputs for identical inputs. A new Bio-IT World commentary explores how intended use, traceability, human oversight, controlled model updates and continuous performance monitoring could make GenAI more suitable for regulated pharmaceutical environments.

A new commentary published by Bio-IT World on 11 September 2026 addresses one of the most difficult questions facing pharmaceutical AI:

How can Generative AI be used in GMP when its outputs are inherently probabilistic rather than fully deterministic?

The article, written by Valentina Armiento and Pierre Bonzani of Genedata, argues that GenAI cannot simply be validated in the same way as conventional GMP software.

Traditional computerized systems are generally expected to behave predictably under defined conditions.

Generative AI is different.

Its output may vary depending on:

  • the prompt;
  • the surrounding context;
  • model configuration;
  • model version;
  • sampling parameters;
  • changes to the underlying model.

This makes repeatability, traceability and lifecycle control considerably more difficult.

The authors therefore argue that GenAI requires a different validation mindset built around:

bounded intended use → risk-based controls → traceability → human oversight → controlled model lifecycle → performance monitoring

Why GenAI Is Different from Conventional GMP Software

The fundamental issue is non-determinism.

Traditional GMP software typically executes predefined logic.

When given the same input under the same configuration, the expected result should normally be reproducible.

Large Language Models operate differently.

They generate outputs probabilistically.

Even when the underlying model has not changed, variations in prompts, context or generation parameters may influence the result.

This creates a validation challenge.

The relevant question is no longer only:

Does the system produce the expected output?

but increasingly:

Does the system consistently remain within an acceptable performance range for its intended use?

That distinction may be fundamental for future GenAI validation in GMP.

The Current Annex 22 Position Remains Restrictive

The current draft EU GMP Annex 22 takes a conservative position.

It applies to static AI/ML models used in critical GMP applications and states that dynamic models that continuously learn during use should not be used in critical GMP applications.

The draft also states that probabilistic models that may produce different outputs from identical inputs are outside its scope and should not be used in critical GMP applications.

It explicitly includes Generative AI and Large Language Models in this restriction.

For non-critical GMP applications, however, GenAI may be considered where qualified personnel remain responsible for ensuring that the output is suitable for its intended use.

In practical terms, the current draft creates a clear distinction between:

critical GMP use

and

non-critical GMP support under human oversight.

Official draft Annex 22:

https://health.ec.eu … ion_guideline_en.pdf

But the Regulatory Discussion Is Still Evolving

The position above should not be interpreted as the final regulatory endpoint.

EMA held a multistakeholder workshop on 30 June and 1 July 2026 specifically to explore whether adaptive and probabilistic AI models could potentially be accommodated within a future risk-based Annex 22 framework.

The workshop examined topics such as:

  • guardrails;
  • human oversight;
  • model evaluation;
  • data governance;
  • accountability;
  • transparency;
  • risk mitigation;
  • validation approaches for adaptive and probabilistic models.

This is important.

The regulatory question is gradually shifting from:

“Should probabilistic AI simply be excluded?”

towards:

“Can probabilistic AI be controlled sufficiently for a defined GMP use?”

EMA Annex 22 workshop:

https://www.ema.euro … development-annex-22

Six Practical Controls Highlighted in the Bio-IT World Commentary

The article identifies several elements that should form part of a GMP-oriented AI control strategy.

1. Clearly Defined Intended Use

The role of the AI should be defined precisely.

The organisation should know:

  • what task the AI performs;
  • which decisions it supports;
  • which data it receives;
  • under which conditions it may be used;
  • where its boundaries are.

This is critical because validation can only be meaningful when performance is assessed against a clearly defined use case.

2. Risk-Based Human Oversight

The level of AI autonomy should depend on the risk of the activity.

For high-impact decisions, the AI should remain advisory and appropriately qualified personnel should retain responsibility for the final decision.

The important question is not simply whether a human is formally present in the workflow.

It is whether the human review is capable of detecting and correcting an unacceptable AI output.

3. Data and Decision Traceability

Inputs, outputs, model versions and relevant user interactions should be traceable.

For GenAI this can potentially include:

  • source data;
  • prompt or query;
  • model version;
  • configuration;
  • generated output;
  • reviewer;
  • final accepted result.

This may become especially important when AI contributes to GMP documentation, investigations or other quality-system activities.

4. Explainability and Confidence

The authors also highlight explainability mechanisms.

For some AI systems this may involve techniques such as feature attribution or confidence scores.

For GenAI, explainability may be more difficult.

In many cases, a more practical control may be strong linkage between the generated answer and controlled source information so that the reviewer can verify the basis of the output.

5. Integration with Existing GMP Controls

AI should not operate outside the Pharmaceutical Quality System.

It should be incorporated, as applicable, into:

  • change control;
  • validation;
  • configuration management;
  • supplier management;
  • periodic review;
  • deviation management;
  • performance monitoring.

This is an important point.

GenAI does not require abandoning established GMP controls.

It requires extending them to address new failure mechanisms.

6. Cross-Functional Governance

AI validation cannot be owned by one function.

Effective governance may require cooperation between:

  • Quality Assurance;
  • IT;
  • Data Science;
  • Validation;
  • system and process owners;
  • subject-matter experts.

This is consistent with the draft Annex 22 approach, which explicitly calls for cooperation between process SMEs, QA, data scientists, IT and other relevant parties.

The Real Problem: Model Updates Can Break the Validated State

One of the most important points in the Bio-IT World article concerns model updates.

For traditional systems, software changes are generally identifiable and controlled.

With GenAI, behaviour may change because:

  • the underlying model is replaced;
  • the provider updates the model;
  • system prompts change;
  • retrieval data changes;
  • configuration changes;
  • context handling changes.

Even where the visible application remains unchanged, the output behaviour may be different.

This creates a difficult lifecycle problem:

How can innovation continue without repeatedly destroying the validated state?

The article proposes several mechanisms:

  • version-locked models;
  • formal change control;
  • predetermined update strategies;
  • risk-based revalidation;
  • continuous performance monitoring.

From Exact Repeatability to Acceptable Performance

Perhaps the most important conceptual idea in the article is that validation may need to move away from expecting identical text outputs.

For GenAI, the relevant acceptance criterion may instead be whether different acceptable outputs remain within predefined quality boundaries.

For example, performance could potentially be evaluated using:

  • correctness;
  • completeness;
  • critical-error rate;
  • hallucination rate;
  • source-grounding;
  • human correction rate;
  • failure-to-escalate rate;
  • performance on predefined challenge cases.

The objective would therefore not necessarily be:

same input → identical wording

but rather:

defined intended use → controlled variability → acceptable and measurable performance.

This represents a major conceptual shift from classical software testing.

Continuous Monitoring Becomes Part of Validation

The article also argues that AI performance should continue to be monitored after deployment.

This is closely aligned with the current draft Annex 22, which requires regular monitoring of model performance and input-data drift for AI models within its scope.

For future GenAI applications, similar principles may become essential.

A lifecycle approach could therefore look like:

Initial Evaluation

Acceptance Testing

Controlled Deployment

Performance Monitoring

Drift / Failure Detection

Change Assessment

Revalidation Where Required

This begins to resemble a maintained AI state of control rather than a one-time validation event.

A Practical Place for GenAI Today

The commentary identifies documentation and knowledge workflows as particularly promising areas.

Examples could include:

  • summarising controlled information;
  • knowledge search;
  • drafting documentation;
  • supporting investigation review;
  • querying validated datasets;
  • assisting with scientific or technical interpretation.

These use cases are more suitable when the AI does not make the final critical GMP decision and qualified personnel remain responsible for reviewing the result.

This remains very different from allowing an LLM to autonomously:

  • release a batch;
  • make a critical product-quality decision;
  • change manufacturing parameters;
  • approve a deviation;
  • close a CAPA.

Why This Article Matters

The significance of this commentary is not that it creates a new regulatory framework.

It does not.

Its value is that it clearly describes the fundamental problem that the industry now needs to solve:

GMP validation was historically designed around deterministic systems, while Generative AI is inherently probabilistic.

Trying to force GenAI into an unchanged deterministic validation model may therefore be unrealistic.

But abandoning validation principles is equally unacceptable.

The likely direction is a middle ground based on:

bounded use → measurable performance → controlled variability → traceability → human oversight → change control → continuous monitoring.

This may ultimately be one of the key concepts determining whether Generative AI can move from GxP-adjacent experimentation into genuinely regulated pharmaceutical workflows.

Regulatory Status Note

The Bio-IT World article is contributed commentary written by representatives of Genedata. It is not regulatory guidance and does not represent the official position of FDA, EMA or another competent authority.

The discussion of Annex 22 should also be understood in the context of an evolving regulatory process. The current draft excludes GenAI and probabilistic models from critical GMP applications, but EMA is actively considering stakeholder input on whether future risk-based controls and guardrails could support broader use.

Sources

Bio-IT World – What It Takes to Make Generative AI Fit for GMP, published 11 September 2026:

https://www.bio-itwo … ative-ai-fit-for-gmp

European Commission – Draft EU GMP Annex 22: Artificial Intelligence:

https://health.ec.eu … ion_guideline_en.pdf

EMA – Multistakeholder workshop on AI guidance development for Annex 22:

https://www.ema.euro … development-annex-22