Friday, September 11, 2026

Making Generative AI Fit for GMP: Validation Must Move Beyond Determinism


Introduction Generative AI challenges traditional GMP validation because probabilistic models cannot always provide identical outputs for identical inputs. A new Bio-IT World commentary explores how intended use, traceability, human oversight, controlled model updates and continuous performance monitoring could make GenAI more suitable for regulated pharmaceutical environments.

A new commentary published by Bio-IT World on 11 September 2026 addresses one of the most difficult questions facing pharmaceutical AI:

How can Generative AI be used in GMP when its outputs are inherently probabilistic rather than fully deterministic?

The article, written by Valentina Armiento and Pierre Bonzani of Genedata, argues that GenAI cannot simply be validated in the same way as conventional GMP software.

Traditional computerized systems are generally expected to behave predictably under defined conditions.

Generative AI is different.

Its output may vary depending on:

  • the prompt;
  • the surrounding context;
  • model configuration;
  • model version;
  • sampling parameters;
  • changes to the underlying model.

This makes repeatability, traceability and lifecycle control considerably more difficult.

The authors therefore argue that GenAI requires a different validation mindset built around:

bounded intended use → risk-based controls → traceability → human oversight → controlled model lifecycle → performance monitoring

Why GenAI Is Different from Conventional GMP Software

The fundamental issue is non-determinism.

Traditional GMP software typically executes predefined logic.

When given the same input under the same configuration, the expected result should normally be reproducible.

Large Language Models operate differently.

They generate outputs probabilistically.

Even when the underlying model has not changed, variations in prompts, context or generation parameters may influence the result.

This creates a validation challenge.

The relevant question is no longer only:

Does the system produce the expected output?

but increasingly:

Does the system consistently remain within an acceptable performance range for its intended use?

That distinction may be fundamental for future GenAI validation in GMP.

The Current Annex 22 Position Remains Restrictive

The current draft EU GMP Annex 22 takes a conservative position.

It applies to static AI/ML models used in critical GMP applications and states that dynamic models that continuously learn during use should not be used in critical GMP applications.

The draft also states that probabilistic models that may produce different outputs from identical inputs are outside its scope and should not be used in critical GMP applications.

It explicitly includes Generative AI and Large Language Models in this restriction.

For non-critical GMP applications, however, GenAI may be considered where qualified personnel remain responsible for ensuring that the output is suitable for its intended use.

In practical terms, the current draft creates a clear distinction between:

critical GMP use

and

non-critical GMP support under human oversight.

Official draft Annex 22:

https://health.ec.eu … ion_guideline_en.pdf

But the Regulatory Discussion Is Still Evolving

The position above should not be interpreted as the final regulatory endpoint.

EMA held a multistakeholder workshop on 30 June and 1 July 2026 specifically to explore whether adaptive and probabilistic AI models could potentially be accommodated within a future risk-based Annex 22 framework.

The workshop examined topics such as:

  • guardrails;
  • human oversight;
  • model evaluation;
  • data governance;
  • accountability;
  • transparency;
  • risk mitigation;
  • validation approaches for adaptive and probabilistic models.

This is important.

The regulatory question is gradually shifting from:

“Should probabilistic AI simply be excluded?”

towards:

“Can probabilistic AI be controlled sufficiently for a defined GMP use?”

EMA Annex 22 workshop:

https://www.ema.euro … development-annex-22

Six Practical Controls Highlighted in the Bio-IT World Commentary

The article identifies several elements that should form part of a GMP-oriented AI control strategy.

1. Clearly Defined Intended Use

The role of the AI should be defined precisely.

The organisation should know:

  • what task the AI performs;
  • which decisions it supports;
  • which data it receives;
  • under which conditions it may be used;
  • where its boundaries are.

This is critical because validation can only be meaningful when performance is assessed against a clearly defined use case.

2. Risk-Based Human Oversight

The level of AI autonomy should depend on the risk of the activity.

For high-impact decisions, the AI should remain advisory and appropriately qualified personnel should retain responsibility for the final decision.

The important question is not simply whether a human is formally present in the workflow.

It is whether the human review is capable of detecting and correcting an unacceptable AI output.

3. Data and Decision Traceability

Inputs, outputs, model versions and relevant user interactions should be traceable.

For GenAI this can potentially include:

  • source data;
  • prompt or query;
  • model version;
  • configuration;
  • generated output;
  • reviewer;
  • final accepted result.

This may become especially important when AI contributes to GMP documentation, investigations or other quality-system activities.

4. Explainability and Confidence

The authors also highlight explainability mechanisms.

For some AI systems this may involve techniques such as feature attribution or confidence scores.

For GenAI, explainability may be more difficult.

In many cases, a more practical control may be strong linkage between the generated answer and controlled source information so that the reviewer can verify the basis of the output.

5. Integration with Existing GMP Controls

AI should not operate outside the Pharmaceutical Quality System.

It should be incorporated, as applicable, into:

  • change control;
  • validation;
  • configuration management;
  • supplier management;
  • periodic review;
  • deviation management;
  • performance monitoring.

This is an important point.

GenAI does not require abandoning established GMP controls.

It requires extending them to address new failure mechanisms.

6. Cross-Functional Governance

AI validation cannot be owned by one function.

Effective governance may require cooperation between:

  • Quality Assurance;
  • IT;
  • Data Science;
  • Validation;
  • system and process owners;
  • subject-matter experts.

This is consistent with the draft Annex 22 approach, which explicitly calls for cooperation between process SMEs, QA, data scientists, IT and other relevant parties.

The Real Problem: Model Updates Can Break the Validated State

One of the most important points in the Bio-IT World article concerns model updates.

For traditional systems, software changes are generally identifiable and controlled.

With GenAI, behaviour may change because:

  • the underlying model is replaced;
  • the provider updates the model;
  • system prompts change;
  • retrieval data changes;
  • configuration changes;
  • context handling changes.

Even where the visible application remains unchanged, the output behaviour may be different.

This creates a difficult lifecycle problem:

How can innovation continue without repeatedly destroying the validated state?

The article proposes several mechanisms:

  • version-locked models;
  • formal change control;
  • predetermined update strategies;
  • risk-based revalidation;
  • continuous performance monitoring.

From Exact Repeatability to Acceptable Performance

Perhaps the most important conceptual idea in the article is that validation may need to move away from expecting identical text outputs.

For GenAI, the relevant acceptance criterion may instead be whether different acceptable outputs remain within predefined quality boundaries.

For example, performance could potentially be evaluated using:

  • correctness;
  • completeness;
  • critical-error rate;
  • hallucination rate;
  • source-grounding;
  • human correction rate;
  • failure-to-escalate rate;
  • performance on predefined challenge cases.

The objective would therefore not necessarily be:

same input → identical wording

but rather:

defined intended use → controlled variability → acceptable and measurable performance.

This represents a major conceptual shift from classical software testing.

Continuous Monitoring Becomes Part of Validation

The article also argues that AI performance should continue to be monitored after deployment.

This is closely aligned with the current draft Annex 22, which requires regular monitoring of model performance and input-data drift for AI models within its scope.

For future GenAI applications, similar principles may become essential.

A lifecycle approach could therefore look like:

Initial Evaluation

Acceptance Testing

Controlled Deployment

Performance Monitoring

Drift / Failure Detection

Change Assessment

Revalidation Where Required

This begins to resemble a maintained AI state of control rather than a one-time validation event.

A Practical Place for GenAI Today

The commentary identifies documentation and knowledge workflows as particularly promising areas.

Examples could include:

  • summarising controlled information;
  • knowledge search;
  • drafting documentation;
  • supporting investigation review;
  • querying validated datasets;
  • assisting with scientific or technical interpretation.

These use cases are more suitable when the AI does not make the final critical GMP decision and qualified personnel remain responsible for reviewing the result.

This remains very different from allowing an LLM to autonomously:

  • release a batch;
  • make a critical product-quality decision;
  • change manufacturing parameters;
  • approve a deviation;
  • close a CAPA.

Why This Article Matters

The significance of this commentary is not that it creates a new regulatory framework.

It does not.

Its value is that it clearly describes the fundamental problem that the industry now needs to solve:

GMP validation was historically designed around deterministic systems, while Generative AI is inherently probabilistic.

Trying to force GenAI into an unchanged deterministic validation model may therefore be unrealistic.

But abandoning validation principles is equally unacceptable.

The likely direction is a middle ground based on:

bounded use → measurable performance → controlled variability → traceability → human oversight → change control → continuous monitoring.

This may ultimately be one of the key concepts determining whether Generative AI can move from GxP-adjacent experimentation into genuinely regulated pharmaceutical workflows.

Regulatory Status Note

The Bio-IT World article is contributed commentary written by representatives of Genedata. It is not regulatory guidance and does not represent the official position of FDA, EMA or another competent authority.

The discussion of Annex 22 should also be understood in the context of an evolving regulatory process. The current draft excludes GenAI and probabilistic models from critical GMP applications, but EMA is actively considering stakeholder input on whether future risk-based controls and guardrails could support broader use.

Sources

Bio-IT World – What It Takes to Make Generative AI Fit for GMP, published 11 September 2026:

https://www.bio-itwo … ative-ai-fit-for-gmp

European Commission – Draft EU GMP Annex 22: Artificial Intelligence:

https://health.ec.eu … ion_guideline_en.pdf

EMA – Multistakeholder workshop on AI guidance development for Annex 22:

https://www.ema.euro … development-annex-22

Tuesday, September 8, 2026

PMDA and PIC/S Put Draft Annex 22 at the Centre of GMP Inspector Training


Introduction PMDA and PIC/S are making draft GMP Annex 22 on Artificial Intelligence a focus of regulator-only inspector training, signalling growing international inspection readiness for AI in pharmaceutical manufacturing.

Artificial intelligence in pharmaceutical manufacturing is increasingly becoming not only a guidance-development topic but also an inspection capability topic.

From 15 to 17 September 2026, Japan’s Pharmaceuticals and Medical Devices Agency (PMDA) will hold its PMDA-ATC GMP Inspection Seminar 2026 in Tokyo.

The seminar is intended exclusively for regulators involved in pharmaceutical GMP inspections.

One of its primary objectives is explicit:

to build a shared understanding of the draft PIC/S GMP Annex 22 on Artificial Intelligence.

The event is aligned with the PIC/S training programme and is also described as a pre-learning opportunity for the 2026 PIC/S Seminar in Istanbul.

For pharmaceutical manufacturers, the event does not introduce a new GMP requirement.

However, it provides an important signal about where international inspector attention is moving.

Annex 22 Is Becoming an Inspector Training Topic

Until recently, much of the discussion around AI in GMP focused on:

  • drafting regulatory guidance;
  • public consultation;
  • industry comments;
  • regulatory workshops;
  • interpretation of existing computerized-system requirements.

The PMDA seminar represents another stage:

training inspectors to understand the technology and discuss how it should be approached during GMP oversight.

According to PMDA, the seminar aims to create shared understanding among participating regulators through discussion of:

  • the draft PIC/S GMP Annex 22;
  • advanced technologies in pharmaceutical manufacturing;
  • GMP inspection approaches;
  • inspection reliance in Asia.

This is significant because implementation of any new GMP expectation ultimately depends not only on the wording of guidance but also on how inspectors interpret and apply it.

The Seminar Is for Experienced GMP Inspectors

The audience is deliberately restricted.

PMDA specifies that participants should be regulatory authority personnel with substantial practical inspection experience.

The criteria include inspectors who have:

  • significant on-site GMP inspection experience;
  • more than three years of work in GMP regulatory activities;
  • intermediate knowledge of GMP regulations.

Approximately 30 regulators are expected to participate.

This is therefore not a general educational webinar about artificial intelligence.

It is a focused GMP inspector capacity-building activity.

Why This Matters for Pharmaceutical Companies

The event itself does not change the legal status of Annex 22.

The Annex remains a draft.

Nevertheless, companies should distinguish between two questions:

Is Annex 22 already an enforceable GMP requirement?

No.

and:

Are GMP inspectorates already developing knowledge and inspection capability around AI?

Clearly yes.

That distinction is important.

Waiting until the final Annex 22 becomes formally effective before establishing any AI governance could leave organisations behind the regulatory learning curve.

Companies already using AI in GMP-relevant processes should therefore be able to explain:

  • where AI is being used;
  • what its intended use is;
  • what its GxP impact is;
  • how the system was evaluated;
  • what data it depends on;
  • how its performance is controlled;
  • what human oversight exists;
  • how changes are managed;
  • how ongoing performance is monitored.

These are not presented here as new requirements created by the PMDA seminar.

They are sensible readiness questions arising from the broader lifecycle-control approach already visible in draft Annex 22 and international AI regulatory discussions.

From Regulatory Text to Regulatory Capability

There are several stages in the development of a new regulatory expectation.

A simplified sequence could be:

Regulatory Discussion

Draft Guidance

Public Consultation

Regulator and Industry Workshops

Inspector Training and Harmonisation

Final Guidance

Inspection Experience

The PMDA-ATC seminar is important because it sits in the middle of this transition.

It shows that regulatory authorities are not only discussing what AI guidance should say.

They are also developing the capability necessary to understand AI during future GMP oversight.

PIC/S Alignment Makes the Event More Significant

PMDA describes the seminar as aligned with the PIC/S training framework.

PIC/S separately lists the PMDA event and confirms that its programme will focus on:

  • the draft PIC/S GMP Annex 22 on Artificial Intelligence;
  • advanced technologies in pharmaceutical manufacturing;
  • GMP inspection reliance.

This matters because PIC/S provides an important international platform for harmonisation of pharmaceutical inspection practices.

The value of Annex 22 therefore potentially extends beyond the European Union.

A common inspector understanding of AI could influence GMP expectations across multiple PIC/S participating authorities.

For multinational pharmaceutical companies, this could gradually reduce the possibility of treating AI governance as an isolated EU compliance topic.

AI Governance May Become Part of Inspection Readiness

Historically, inspection-readiness programmes typically focus on topics such as:

  • quality systems;
  • data integrity;
  • validation;
  • deviations and CAPA;
  • change control;
  • computerized systems;
  • process validation;
  • supplier management.

As AI becomes embedded in these processes, a new layer may emerge.

Inspectors may increasingly need to understand whether AI is involved when reviewing:

  • GMP documentation;
  • deviation investigations;
  • CAPA proposals;
  • quality trending;
  • process monitoring;
  • predictive maintenance;
  • laboratory activities;
  • manufacturing control;
  • decision-support systems.

This does not necessarily mean that inspectors will begin conducting separate “AI inspections”.

More likely, AI governance will become part of existing GMP system inspection whenever AI materially affects the regulated activity.

An AI Inventory Becomes Increasingly Important

One practical implication is particularly straightforward.

A company cannot adequately govern AI if it does not know where AI is being used.

A useful first step is therefore a controlled inventory covering:

  • AI system or model;
  • system owner;
  • intended use;
  • GxP relevance;
  • supplier;
  • model type;
  • data sources;
  • validation status;
  • human oversight;
  • change-management approach;
  • performance monitoring.

This becomes even more important when AI functionality is embedded inside existing commercial software and may not initially be recognised as a separate AI application.

Do Not Confuse Training Activity with a New Requirement

It is equally important not to overinterpret the PMDA event.

The seminar:

  • does not make draft Annex 22 effective;
  • does not create a new Japanese GMP requirement;
  • does not demonstrate that every participating regulator will interpret Annex 22 identically;
  • does not establish an implementation date.

Its importance lies elsewhere.

It demonstrates that experienced inspectors are already being prepared to understand and discuss AI in pharmaceutical manufacturing.

That is a meaningful regulatory signal even before the final Annex is published.

What Pharma Should Watch Next

Several developments will now be particularly important:

  • the finalisation of EU GMP Annex 22;
  • any changes to the treatment of Generative AI and probabilistic models;
  • the relationship between Annex 22 and revised Annex 11;
  • PIC/S adoption and implementation;
  • future inspector training materials;
  • early inspection experience involving AI-enabled systems.

The most valuable evidence will eventually come not only from the final wording of Annex 22 but from how regulators apply it in practice.

Why This Development Is Important

The pharmaceutical AI discussion is entering a new phase.

The question is moving from:

“Should GMP regulators develop requirements for AI?”

towards:

“How should inspectors evaluate AI when they encounter it in pharmaceutical manufacturing?”

The PMDA/PIC/S training activity suggests that regulatory authorities are beginning to prepare for that second question.

For pharmaceutical companies, this is a good reason to begin treating AI governance as part of routine GMP inspection readiness rather than as a future digital-transformation issue.

Regulatory Status Note

The PMDA-ATC GMP Inspection Seminar is a regulator training and capacity-building activity.

Draft PIC/S/EU GMP Annex 22 remains under development and is not currently an effective GMP requirement.

The observations above concerning inspection readiness are therefore an interpretation of the regulatory direction and should not be read as new formal requirements arising from the seminar.

Sources

PMDA – PMDA-ATC GMP Inspection Seminar 2026:

https://www.pmda.go. … h/symposia/0351.html

PIC/S – Meetings and Training Events:

https://picscheme.org/en/events

UK Government Publishes AI Risk Management Toolkit – A Practical Model Pharma Can Adapt for AI Governance


Introduction The UK Government has published a practical AI Risk Management Toolkit covering lifecycle risk, performance, human oversight, drift and incident response, with useful concepts for pharmaceutical AI governance.

On 8 September 2026, the UK Department for Science, Innovation and Technology published a new AI Risk Management Toolkit.

The toolkit is intended for organisations designing, procuring, operating or deploying artificial intelligence systems.

It is not pharmaceutical guidance and it does not establish GMP requirements.

Nevertheless, it is particularly interesting for pharmaceutical Quality, IT, Validation and AI Governance teams because it translates general responsible-AI principles into a relatively concrete operational risk-management process.

For companies currently trying to answer questions such as:

  • How should an AI use case be risk assessed?
  • How should different AI risks be prioritised?
  • How should acceptable risk be defined?
  • How should AI performance be monitored?
  • When should human intervention be mandatory?
  • What should happen when AI performance becomes unacceptable?

the toolkit provides several useful concepts that could complement established pharmaceutical Quality Risk Management.

Risk Management Should Cover the Entire AI Lifecycle

One of the strongest aspects of the UK approach is that AI risk assessment is not presented as a one-time activity performed before deployment.

The toolkit describes risk management across the AI lifecycle, from use-case identification through deployment and ultimately retirement.

It also recognises that AI risks can change because of:

  • changes in data;
  • model updates;
  • changes in the operating environment;
  • changes in intended or actual use;
  • new technical vulnerabilities;
  • performance degradation;
  • new regulatory requirements.

This leads to a principle that is highly relevant to pharmaceutical AI:

AI risk assessment should be continuously reassessed rather than treated as a static validation document.

The toolkit specifically highlights the need to reconsider risk when models are updated and to monitor technical performance during live operation.

For GxP AI, this concept fits naturally with lifecycle management, change control, periodic review and continued performance monitoring.

A Multidisciplinary AI Risk Management Team

The UK toolkit also proposes that AI risk should not be owned exclusively by IT or data science.

It recommends multidisciplinary involvement including:

  • AI governance;
  • senior management;
  • data specialists;
  • AI and machine-learning specialists;
  • security;
  • legal and compliance;
  • business or domain experts;
  • end users.

This is highly relevant for pharmaceutical implementation.

For a GxP application, the equivalent governance structure would normally also require appropriate involvement of:

  • Quality;
  • system/process owner;
  • Validation or CSV/CSA;
  • IT;
  • Information Security;
  • Data Governance;
  • AI/ML specialists;
  • relevant SMEs.

The important principle is that AI risk cannot be adequately assessed from a purely technical perspective.

The significance of an incorrect AI output depends on its pharmaceutical context.

An incorrect recommendation generated by an AI system supporting general knowledge retrieval is fundamentally different from an incorrect result influencing a deviation investigation, process control or product-quality decision.

The Toolkit Introduces the Important Concept of AI Risk Appetite

An interesting feature is the explicit use of risk appetite.

Before determining mitigation measures, organisations are encouraged to define how much risk they are prepared to tolerate.

The toolkit recommends establishing risk appetite at organisational level where possible rather than deciding separately for every AI project.

This could be useful for pharmaceutical AI governance.

A pharmaceutical company might, for example, establish fundamentally different levels of tolerance for:

  • AI used for administrative productivity;
  • AI used for document search;
  • AI supporting GMP investigations;
  • AI supporting manufacturing decisions;
  • AI influencing critical process controls;
  • AI involved in product disposition.

This creates an important distinction between:

AI risk

and

acceptable AI risk for a particular intended use.

For GMP applications, patient safety, product quality and data integrity would normally impose much lower risk tolerance than ordinary business applications.

Risk Is Quantified Using Likelihood and Impact

The toolkit proposes scoring both:

  • Likelihood – from 1 to 5;
  • Impact – from 1 to 5.

The resulting risk score is calculated from likelihood × impact.

This is conceptually familiar to pharmaceutical Quality Risk Management.

However, the toolkit also highlights an important challenge specific to AI:

historical failure data may be limited.

Therefore, likelihood estimates may need to combine:

  • historical information;
  • model performance analysis;
  • expert judgement;
  • experimentation;
  • real-world monitoring.

This is particularly relevant for novel GenAI systems where traditional failure-frequency data may simply not exist.

The toolkit therefore encourages organisations to update assumptions as operational evidence becomes available.

For pharmaceutical companies this could mean that an initial AI risk assessment should evolve as:

validation data → pilot data → operational data → incidents → monitoring results

become available.

Technical Robustness Is Treated as a Lifecycle Requirement

One of the most GMP-relevant sections concerns technical robustness.

The toolkit asks organisations to consider:

  • whether the model is suitable for the intended context;
  • how required performance will be measured;
  • whether performance remains acceptable after deployment;
  • how data quality can affect performance;
  • whether models are tested under adverse or unexpected conditions;
  • how drift will be detected;
  • how often robustness will be reassessed;
  • what happens when performance problems are identified.

This is particularly important because AI performance cannot always be treated as a fixed characteristic.

A system can meet acceptance criteria during initial validation and later become less reliable because the operating environment changes.

This supports a broader lifecycle concept:

Initial Validation → Operational Monitoring → Drift Detection → Reassessment → Change Control / Revalidation

For AI used in GMP, maintaining performance may therefore become as important as demonstrating initial performance.

The Toolkit Goes Beyond “Human-in-the-Loop”

Human oversight is frequently mentioned in AI guidance, but often only as a general principle.

The UK toolkit makes the concept more practical.

For poor AI accuracy or performance, suggested controls include establishing policies defining the required level of human involvement in AI-supported decision-making.

This raises an important question for pharmaceutical companies:

What exactly must the human reviewer do?

Simply requiring someone to click “Approve” is not meaningful human oversight.

Effective oversight may require that the reviewer:

  • understands the purpose of the AI;
  • understands its known limitations;
  • has access to the underlying evidence;
  • can recognise implausible outputs;
  • has authority to reject the AI recommendation;
  • can independently determine the final decision where necessary.

For critical GMP applications, human oversight therefore needs to be designed and demonstrated as an effective control rather than merely documented as a workflow step.

Performance Monitoring Needs Defined Acceptance Limits

Another particularly useful recommendation is that organisations establish:

  • performance metrics;
  • ongoing testing;
  • defined acceptable performance limits.

This is highly compatible with pharmaceutical validation principles.

An AI system should therefore not simply be described as being “monitored”.

The organisation should know:

What are we monitoring?

What constitutes acceptable performance?

What constitutes deterioration?

What action is triggered when the limit is exceeded?

This could include metrics such as:

  • accuracy;
  • false-positive rate;
  • false-negative rate;
  • critical-error frequency;
  • human correction rate;
  • override rate;
  • hallucination frequency;
  • drift indicators;
  • failure to defer when uncertain.

The appropriate measures will depend on the intended use.

A Particularly Important Control: Be Able to Switch the AI Off

One of the more practical recommendations concerns system failure.

The toolkit recommends procedures to:

  • bypass the AI;
  • deactivate the AI;
  • maintain redundant or backup processes;
  • define thresholds for activating these controls.

This is an important but sometimes overlooked aspect of AI governance.

A pharmaceutical organisation should not become dependent on an AI system to such an extent that the regulated process can no longer function safely when the AI becomes unavailable or unreliable.

This creates a useful control principle:

A critical AI-enabled process should have a defined safe state and a defined response to loss of AI capability.

For some applications this could mean returning temporarily to a validated manual process.

For others it could require stopping the affected operation.

Red-Teaming and Stress Testing

The toolkit also recommends adversarial or stress testing to deliberately search for failure modes and vulnerabilities.

For pharmaceutical AI, this could be particularly useful for GenAI systems.

Instead of testing only expected use cases, validation could deliberately challenge the system with:

  • ambiguous instructions;
  • incomplete data;
  • conflicting information;
  • incorrect assumptions;
  • unusual records;
  • edge cases;
  • prompt injection;
  • misleading source documents.

This shifts validation from:

“Can the system produce the correct result under normal conditions?”

towards:

“How does the system fail when conditions are abnormal?”

For probabilistic AI, understanding failure behaviour may be at least as important as measuring average accuracy.

How Could Pharma Use This Toolkit?

The UK AI Risk Management Toolkit should not replace ICH Q9 Quality Risk Management, computerized-system validation or specific GxP requirements.

However, pharmaceutical companies could use some of its concepts as an additional AI-specific layer.

A practical pharmaceutical framework could look like:

Intended Use

GxP Impact

AI Risk Identification

Risk Appetite / Tolerance

Likelihood × Impact

Risk Controls

Validation / Qualification

Operational Performance Monitoring

Incident / Deviation / CAPA

Change Control / Revalidation

This could provide a useful bridge between traditional pharmaceutical Quality Risk Management and risks specific to modern AI.

Why This Publication Is Important

Many existing AI governance documents describe broad principles such as fairness, transparency, accountability and human oversight.

The new UK toolkit goes one step further by asking organisations to define:

  • actual risk scenarios;
  • risk likelihood;
  • risk impact;
  • risk tolerance;
  • mitigation measures;
  • monitoring;
  • failure response.

For pharmaceutical companies, this is useful because GMP already operates according to a similar philosophy:

identify risk → establish controls → verify effectiveness → monitor → respond to failure.

The terminology is different, but the underlying quality-management logic is familiar.

Regulatory Status Note

The AI Risk Management Toolkit was published by the UK Department for Science, Innovation and Technology.

It is general AI risk-management guidance and is not an MHRA GMP requirement, pharmaceutical guideline or replacement for ICH Q9, EU GMP or FDA CGMP requirements.

Its relevance to pharmaceutical AI therefore lies in the practical governance concepts that could supplement existing GxP risk-management frameworks.

Source

UK Government – AI Risk Management Toolkit, published 8 September 2026:

https://www.gov.uk/g … k-management-toolkit

Detailed guidance:

https://www.gov.uk/g … ent-toolkit-guidance

Wednesday, September 2, 2026

AI in Medicines: Accuracy and Reliability Emerge as the Top Regulatory Science Priorities


Introduction A European regulatory science study identifies the top priorities for trustworthy AI in medicines, led by accuracy, reliability, robustness, data governance and lifecycle control, with direct implications for pharmaceutical manufacturing and GMP.

A new paper published on 13 August 2026 in Clinical Pharmacology & Therapeutics provides an important indication of where future regulatory science for artificial intelligence in medicines may be heading.

The study, Regulatory Research Priorities for AI Use in the Medicine Lifecycle: A European Perspective with Global Relevance, was developed within the European medicines regulatory network and involved authors affiliated with the European Medicines Agency and several European regulatory organisations.

The researchers asked a particularly important question:

Which scientific problems need to be solved to enable trustworthy use of AI throughout the medicines lifecycle?

The results suggest that stakeholders are increasingly looking beyond general AI governance principles.

The major unresolved challenge is becoming much more practical:

How can we demonstrate that AI-generated results are sufficiently accurate, reliable, robust and trustworthy for regulated pharmaceutical use?

273 Stakeholders Across the Medicines Ecosystem

The study was based on a European-wide survey conducted by the Network Data Steering Group of the European Medicines Regulatory Network.

A total of 273 responses were collected from groups including:

  • national competent authorities;
  • pharmaceutical industry professionals;
  • small and medium-sized enterprises;
  • patients and consumers;
  • healthcare professionals;
  • contract research organisations;
  • academic researchers;
  • EU agency professionals.

The largest groups were national competent authority employees, representing 24% of respondents, and pharmaceutical industry professionals, representing 22%.

Importantly, most respondents already had at least some professional familiarity with AI.

The study therefore provides more than a general public perception of AI. It captures views from people directly involved in medicines development, regulation and evaluation.

Seven Areas of AI Regulatory Science Were Evaluated

The researchers identified 28 research questions grouped into seven domains:

  • accuracy and reliability of AI tools;
  • data governance, confidentiality and consent;
  • ethics, fairness and bias prevention;
  • regulation and oversight;
  • research integrity and intellectual property;
  • resources and support for AI use;
  • impact on jobs and skills.

Respondents were asked to rank the most important challenges.

The result was particularly interesting.

Accuracy and reliability of AI tools emerged as the highest-priority domain by a substantial margin.

The Number-One Question: Can We Trust the AI Output?

The highest-ranked research question concerned how the accuracy and reliability of AI models can be ensured when they generate evidence or inform regulatory decisions, and how their limitations can be identified.

Within the accuracy and reliability domain, this question was ranked highest by 45% of respondents.

The second major issue was closely related:

How can AI remain robust and trustworthy when data changes, performance deteriorates over time, information is incomplete or inputs are misleading?

This brings several familiar AI lifecycle risks directly into the regulatory science discussion:

  • data drift;
  • model drift;
  • performance degradation;
  • incomplete information;
  • unexpected inputs;
  • misleading inputs;
  • uncertainty;
  • changing operating environments.

The implication for pharmaceutical companies is important.

Initial AI validation may demonstrate acceptable performance at one point in time, but it does not automatically demonstrate that performance will remain acceptable throughout the system lifecycle.

From Validation to Maintaining an AI State of Control

This finding is particularly relevant for GMP applications.

Traditional computerized system validation often asks:

Did the system meet its predefined requirements when it was tested?

For AI, another question becomes equally important:

Does the system continue to perform reliably under changing conditions?

This leads naturally toward lifecycle controls such as:

  • ongoing performance monitoring;
  • predefined performance limits;
  • drift detection;
  • periodic re-evaluation;
  • change control;
  • model version control;
  • revalidation or requalification triggers;
  • monitoring of failure modes.

The regulatory science priorities identified in the study therefore support a broader concept of AI assurance:

Validation should not only demonstrate initial fitness for intended use. It should also establish how continued fitness for use will be demonstrated.

For GMP applications, this is closely related to the emerging concept of maintaining AI in an algorithmic state of control.

Explainability Is Important – But the Question Is “When and How Much?”

The third priority within the accuracy and reliability domain concerned explainability.

The study asks:

When is AI explainability important for regulatory decision-making, and which methods can provide it without unnecessarily affecting model performance?

This is an important distinction.

The future regulatory expectation may not necessarily be:

Every AI model must always be completely explainable.

A more risk-based question could be:

What level of explainability is necessary for this particular intended use, decision and associated risk?

For example, explainability requirements could reasonably differ between AI used to:

  • search regulatory information;
  • identify possible trends;
  • support deviation investigations;
  • predict process behaviour;
  • support product quality decisions;
  • generate evidence used in regulatory submissions.

This would be consistent with a context-of-use and risk-based approach to AI assurance.

Data Governance Is the Second Major Priority

The second-highest overall research domain was:

Data governance, confidentiality and consent.

Among the important questions identified were how data used for AI-based medicines development can be made:

  • secure;
  • auditable;
  • traceable;
  • legally and ethically usable.

For pharmaceutical companies, this reinforces a fundamental principle:

AI assurance starts with assurance of the data used by AI.

An advanced AI model cannot compensate for data that are incomplete, poorly controlled, untraceable or inappropriate for the intended purpose.

In GxP environments this connects directly with established expectations for:

  • data integrity;
  • data provenance;
  • traceability;
  • security;
  • access control;
  • retention;
  • auditability.

Bias and Transparency Remain Major Concerns

Ethics, fairness and bias prevention ranked third among the seven domains.

The study identified two particularly important questions:

  • How can bias in AI models used in medicines development and evaluation be identified and reduced?
  • How should the limitations and uncertainties of AI use be communicated transparently?

The second point may be particularly important in regulated environments.

AI output should not create an impression of certainty that the underlying model cannot justify.

A mature pharmaceutical AI system may therefore need mechanisms for communicating:

  • uncertainty;
  • limitations;
  • confidence;
  • conditions under which the model should not be used;
  • situations requiring escalation to a human expert.

A Surprising Result: “More Regulation” Was Not the Highest Priority

One of the most interesting findings is that Regulation and oversight ranked only fourth among the seven domains.

This should not be interpreted as suggesting that regulation is unimportant.

The authors propose another possible explanation.

Horizontal frameworks such as the EU Artificial Intelligence Act already provide important elements of AI governance.

What remains unresolved are many of the scientific and methodological questions that legislation alone cannot answer.

For example:

How accurate is accurate enough?

How should robustness be demonstrated?

How should performance degradation be detected?

When is explainability necessary?

What validation methodology should be used?

How should uncertainty be communicated?

These are fundamentally different questions from simply determining whether an AI system is legally permitted.

For pharmaceutical companies, this distinction is important.

The next major challenge may not be another general AI policy. It may be the development of accepted scientific methods for demonstrating trustworthy AI performance.

The Ten Highest Regulatory Research Priorities

After weighting the research questions according to the importance of their respective domains, the study identified ten priority areas.

They concern:

  • accuracy and reliability of AI-generated evidence;
  • robustness under changing data, performance degradation and misleading inputs;
  • appropriate AI explainability;
  • legal basis and consent for secondary use of health data;
  • security, auditability and traceability of AI data;
  • identification and reduction of AI bias;
  • transparent communication of AI limitations and uncertainties;
  • gaps in guidance for benefit-risk assessment, pharmacovigilance and evidence generation;
  • transparency and reproducibility of AI-enabled research;
  • technical checks and quality controls required to ensure that AI systems are validated, secure and reliable.

One Priority Is Directly Relevant to Pharmaceutical Manufacturing

The tenth research priority is especially important from a GMP perspective.

The study asks which:

technical checks and quality controls are essential to ensure that AI systems used in clinical trials, manufacturing, safety monitoring and regulatory submissions are properly validated, secure and reliable.

This explicitly places manufacturing within the regulatory science agenda for AI assurance.

The question goes beyond conventional software validation.

For an AI application in pharmaceutical manufacturing, appropriate control could potentially require consideration of:

  • intended use;
  • model performance;
  • independent test data;
  • data integrity;
  • robustness;
  • security;
  • human oversight;
  • model changes;
  • performance monitoring;
  • drift;
  • revalidation triggers;
  • failure management.

These are likely to become increasingly important as AI moves from experimental applications toward routine GxP processes.

The Real Question Is Moving from “Can We Use AI?” to “Can We Trust It?”

Perhaps the most important conclusion of the paper is captured by the overall pattern of the results.

Stakeholders do not appear primarily concerned with whether artificial intelligence should be used in medicines development.

The more important question is:

How can AI-generated outputs be trusted when they influence evidence generation and regulatory decision-making?

This represents an important evolution of the pharmaceutical AI discussion.

The early discussion was largely:

Can AI be used in regulated pharmaceutical processes?

The emerging discussion is increasingly:

What evidence is necessary to demonstrate that AI is sufficiently trustworthy for a particular regulated use?

That requires moving from generic principles toward measurable scientific criteria.

What This Means for Pharmaceutical Companies

For pharmaceutical organisations developing AI governance frameworks, the study provides a useful indication of where attention should be focused.

A practical AI assurance model could increasingly be built around:

Intended Use → Risk → Data → Performance → Robustness → Explainability → Controls → Human Oversight → Change Control → Continuous Monitoring

This is broader than traditional computerized system validation.

It also suggests that the future of AI assurance in pharma may be less about creating one universal “AI validation procedure” and more about establishing a multi-layered control strategy appropriate to the intended use and associated risk.

The fundamental regulatory science question therefore becomes:

Can we demonstrate – with evidence – that this AI remains sufficiently reliable, secure and controlled for the decision or process it supports?

Regulatory Status Note

This publication is a scientific research article and does not constitute formal EMA guidance or a new regulatory requirement.

The authors explicitly state that the views expressed are their personal views and should not be interpreted as representing the official position of the regulatory agencies or organisations with which they are affiliated.

The study should therefore be understood as an important indicator of regulatory science priorities rather than as an enforceable regulatory framework.

Nevertheless, its importance is increased by the involvement of the European Medicines Regulatory Network and by the participation of regulators, pharmaceutical industry professionals and other stakeholders.

Source

Pinheiro L.C. et al. – Regulatory Research Priorities for AI Use in the Medicine Lifecycle: A European Perspective with Global Relevance. Clinical Pharmacology & Therapeutics. First published 13 August 2026.

https://ascpt.online … oi/10.1002/cpt.70400

DOI:

https://doi.org/10.1002/cpt.70400

Monday, August 31, 2026

FDA Scientists Propose “Computable Pharmacovigilance”: Automate What Can Be Verified, Keep Human Judgment Where It Cannot


Introduction FDA scientists propose “computable pharmacovigilance,” a framework for deciding which regulated pharmaceutical tasks can be automated by AI and which still require expert human judgment.

A new perspective paper published on 18 August 2026 in Frontiers in Drug Safety and Regulation provides an important new way of thinking about AI automation in regulated pharmaceutical processes.

The authors are scientists from the U.S. Food and Drug Administration, including the National Center for Toxicological Research and the Center for Drug Evaluation and Research.

The paper focuses specifically on pharmacovigilance, but the underlying concept may have much broader relevance for pharmaceutical Quality Systems and GMP applications.

The central question is not simply:

“Can AI perform this task?”

Instead, the authors propose asking:

“Can this task be translated into sufficiently explicit, auditable and verifiable computational steps?”

They call this concept “computable pharmacovigilance.”

Not Every Pharmaceutical Task Is Equally Suitable for AI Automation

The framework distinguishes between tasks that can be clearly formalised and tasks that depend heavily on professional judgment.

Examples of relatively computable activities include:

  • checking whether required information is missing;
  • detecting potentially duplicated safety reports;
  • extracting predefined information from records;
  • coding structured information;
  • retrieving relevant information;
  • performing defined rule-based checks.

Such activities have relatively clear:

inputs → processing rules → outputs → acceptance criteria

and can therefore be tested, benchmarked and audited.

Other activities are fundamentally more difficult.

The authors use case-level causality assessment as an important example.

Determining whether a medicine actually caused a particular adverse event can require integration of:

  • temporality;
  • alternative explanations;
  • clinical context;
  • pharmacological mechanisms;
  • rechallenge information;
  • professional medical judgment.

These activities cannot yet be reduced reliably to fully explicit computational rules.

The consequence is important:

AI suitability should be assessed task by task rather than system by system.

“Computability” Could Be a Useful Concept Beyond Pharmacovigilance

Although the paper concerns pharmacovigilance, the same reasoning could potentially be useful when evaluating AI use in GMP.

For example, some Quality System activities may be highly computable:

  • checking whether required fields are completed;
  • comparing a document against predefined requirements;
  • identifying duplicate or recurring deviations;
  • extracting equipment or batch information;
  • checking dates, limits or predefined conditions;
  • classifying documents against controlled taxonomies.

Other activities may remain much less computable:

  • determining the true root cause of a complex deviation;
  • assessing whether an unexplained event affects product quality;
  • evaluating the adequacy of a CAPA strategy;
  • making a final batch disposition decision;
  • determining whether an unusual observation represents an unacceptable patient risk.

This suggests a potentially useful AI risk-assessment question for pharma:

How much of the intended task can actually be expressed as explicit, testable and auditable logic?

The less computable the task, the greater the likely need for expert oversight.

Large Language Models Are Not Automatically the Best Solution

Another important message from the paper concerns the current tendency to use large language models for almost every AI application.

The authors distinguish between:

  • large models – general-purpose LLMs and foundation models;
  • small models – task-specific machine-learning models;
  • deterministic tools – rules, dictionaries, workflow logic and conventional software.

Large models offer flexibility and can work across multiple text-heavy tasks.

But they also introduce important disadvantages:

  • hallucinations;
  • variable outputs;
  • prompt sensitivity;
  • possible information loss;
  • more difficult validation;
  • more difficult auditability.

Smaller task-specific models can be easier to validate when the task has constrained inputs and clearly defined outputs.

The paper therefore argues that the future may not belong to one universal AI model.

Instead, the most practical architecture may be hybrid.

A Hybrid Architecture: LLM + Small Models + Deterministic Tools

One particularly interesting proposal is an architecture in which an LLM acts as an orchestrator but does not perform every regulated function itself.

The architecture could look like:

User / Reviewer

LLM or AI Agent

Controlled Tools / Retrieval / Rules / Small Models

Structured Result

Expert Review and Decision

The LLM can interpret the request and coordinate different tools.

But specific operations can be delegated to more deterministic components such as:

  • validated rule engines;
  • controlled terminology;
  • structured retrieval;
  • database queries;
  • schema validation;
  • task-specific ML models.

This architecture can reduce uncontrolled generative behaviour.

It can also improve auditability because intermediate actions and outputs can be recorded.

For regulated pharmaceutical applications, this may be a much more attractive architecture than asking one general-purpose LLM to perform the entire process.

An Interesting Finding: Bigger AI Was Not Always Better

The authors also report an illustrative comparison between a task-specific BioBERT model and a general-purpose Llama-4 model for extraction of safety-related information.

For the evaluated structured extraction tasks, the task-specific model performed better.

The authors do not present this as a universal comparison between the two technologies.

Instead, the example demonstrates a broader principle:

The most powerful general-purpose model is not necessarily the best model for a tightly defined regulated task.

For pharmaceutical companies, this is significant.

The correct question may therefore not be:

“Which is the best AI model?”

but:

“Which technical architecture provides the most reliable, verifiable and auditable performance for this particular task?”

“Trust but Verify”

The paper describes the boundary of AI automation using a “trust but verify” philosophy.

AI can increasingly perform routine, well-specified operations.

However, as the process moves from:

pattern recognition → interpretation → judgment → regulated decision

the requirement for structured verification and expert review becomes increasingly important.

The authors conclude that Generative AI does not currently mean the end of human pharmacovigilance.

Instead, the likely future is a:

Human–AI system

in which AI performs increasingly sophisticated operations while humans remain responsible for judgment-intensive decisions.

Why This Matters for GMP and Pharmaceutical Quality

This concept could be particularly useful when companies perform AI use-case assessments.

Instead of simply classifying an application as:

AI / non-AI

or:

GxP / non-GxP

an additional question could be introduced:

How computable is the task?

A possible assessment sequence could therefore become:

Intended Use → GxP Impact → Task Computability → AI Architecture → Verification Strategy → Human Oversight → Lifecycle Monitoring

This could help organisations determine whether a task should be performed using:

  • deterministic software;
  • a task-specific ML model;
  • an LLM;
  • a hybrid architecture;
  • or primarily human expert judgment.

This may ultimately be more useful than assuming that Generative AI should replace existing software simply because it is more flexible.

Regulatory Status Note

The article was written by FDA scientists, but it is a scientific perspective paper and does not represent formal FDA guidance or FDA policy.

The authors explicitly state that the views expressed are their own and do not necessarily represent the official position of the U.S. Food and Drug Administration.

Nevertheless, the paper provides valuable insight into how regulatory scientists are thinking about AI automation, validation, auditability and human judgment in regulated pharmaceutical activities.

Source

Wu L., Xu J., Dang O., Ball R. – Does generative AI mean the “end of history” for pharmacovigilance automation? Towards a framework for the future of human-AI systems. Frontiers in Drug Safety and Regulation, 18 August 2026:

https://www.frontier … fr.2026.1846339/full