Thursday, August 20, 2026

FDA Proposes a New Way to Evaluate Generative AI: From Exhaustive Testing to Competency-Based Lifecycle Assurance


Introduction FDA proposes a competency-based approach for evaluating Generative AI-enabled medical devices, including risk-based benchmarking, clinical confirmation, lifecycle monitoring, model drift control and re-evaluation after AI changes.

The U.S. FDA has opened an important regulatory discussion on how Generative AI-enabled medical devices could be evaluated when traditional software testing is no longer sufficient.

On 18 August 2026, FDA’s Center for Devices and Radiological Health (CDRH) released the discussion paper Considerations for the Regulation of Generative AI-Enabled Medical Devices. The Agency is requesting stakeholder feedback until 19 October 2026.

The document applies specifically to medical devices, not to pharmaceutical GMP systems. Nevertheless, several concepts discussed by FDA are highly relevant to the broader debate about how probabilistic and generative AI can be validated and controlled in regulated environments.

The most significant idea may be surprisingly simple:

For sufficiently complex GenAI systems, it may be unrealistic to test every possible input and output. Instead, regulators may need to assess whether the system can demonstrate defined competencies and continue to demonstrate them throughout its lifecycle.

Why Traditional Software Testing May Not Be Enough

FDA recognizes that GenAI-enabled devices differ fundamentally from traditional software and even from many conventional AI systems.

They may:

  • accept open-ended inputs;
  • generate variable outputs for similar inputs;
  • perform multiple different subtasks;
  • use third-party foundation models;
  • change because of updates to models, prompts, retrieval mechanisms, guardrails or orchestration logic;
  • operate with increasing levels of autonomy.

FDA also explicitly identifies risks such as confabulations or hallucinations, limited transparency into third-party foundation models and performance degradation over time.

This creates a fundamental validation challenge.

For conventional software with bounded inputs and predetermined outputs, extensive input-output testing can provide strong evidence that the system operates as intended.

For GenAI, the possible combination of inputs, conversations and outputs can become effectively unlimited. FDA therefore acknowledges that evaluating every conceivable situation may simply not be practical.

Risk Depends on Both Autonomy and Consequences

FDA proposes a possible two-axis framework for thinking about GenAI risk.

One dimension considers what the AI actually does — ranging from providing non-directive information through directing an action to taking an action autonomously.

The second dimension considers the consequence of relying on an incorrect output.

Risk therefore increases as AI becomes more autonomous and as the potential consequences of an incorrect output become more serious.

This is an important distinction because the same underlying AI technology could represent very different levels of regulatory risk depending on its intended use.

An AI system that provides general information is fundamentally different from one that recommends a specific clinical action — and different again from an agentic system capable of executing that action.

FDA also emphasizes that simply adding wording such as “talk to your doctor” or “I am not a medical professional” may not necessarily make an otherwise action-directing AI function less directive.

A Competency-Based Approach to GenAI Evaluation

Perhaps the most innovative part of the discussion paper is FDA’s consideration of a competency-based approach inspired, at a high level, by the way human clinicians are evaluated.

Doctors are not qualified by testing every possible clinical situation they could ever encounter. Instead, they demonstrate competencies through examinations, supervised practice and continuing assessment.

FDA is exploring whether a related concept could be adapted for GenAI-enabled medical devices.

The proposed approach consists of two major components:

1. Device benchmarking

The deployed or representative final system would be tested against predefined competencies.

2. Clinical confirmation

Evidence would then confirm that the device performs appropriately under real or clinically representative conditions.

Importantly, FDA proposes evaluating the final user-facing device as configured for deployment, rather than assessing only the underlying foundation model such as an LLM.

This distinction is critical.

The performance of an AI application does not depend only on the model. It may also depend on the system prompt, retrieval architecture, knowledge sources, guardrails, user interface, orchestration logic and other controls surrounding the model.

What Would Be Tested?

FDA identifies several possible areas of competency for benchmarking GenAI systems.

They include:

Safety

  • recognition and escalation of safety-critical situations;
  • maintaining the defined scope and operational boundaries;
  • appropriate communication of uncertainty and deferral when the system cannot provide a reliable answer.

Clinical proficiency

  • knowledge and task fidelity;
  • information gathering and analysis;
  • quantitative reasoning;
  • quality and comprehensibility of communication.

Generalizability

  • robustness, reliability and reproducibility;
  • performance across relevant subgroups.

Agentic capabilities

  • additional competencies where AI can autonomously plan, use tools or execute multi-step actions.

FDA also discusses testing boundary adherence using techniques such as adversarial prompting, prompt injection and multi-turn conversations that gradually move outside the intended scope of the system.

This illustrates an important change in thinking about AI validation:

Validation is not only about whether AI produces correct answers. It is also about whether it behaves safely when it does not know the answer, when it is challenged, and when a user attempts to push it beyond its intended use.

Acceptance Criteria Still Matter

The competency-based approach does not mean abandoning predefined validation requirements.

Quite the opposite.

FDA considers that manufacturers should define the scope of testing based on intended use and risk, prespecify evaluation methods, justify scoring methods and establish acceptance criteria before testing.

Where human expert assessment is needed, FDA also discusses the use of appropriately qualified and structurally independent adjudicators.

This could become particularly important for GenAI because many outputs cannot simply be classified as mathematically “correct” or “incorrect.” Evaluation may require structured expert judgement supported by predefined scoring criteria.

Validation Does Not End at Deployment

FDA places substantial emphasis on postmarket performance monitoring.

Possible approaches include:

  • periodic re-benchmarking against predefined performance thresholds;
  • periodic review of real-world AI outputs by qualified independent clinicians;
  • monitoring for model or data drift and other forms of performance degradation.

Reassessment could occur periodically and following defined triggering events, including changes to the underlying model or other components of the AI architecture.

FDA even asks whether, under appropriate circumstances, greater uncertainty could be accepted before market authorization if it were compensated by stronger postmarket monitoring.

That is a potentially important regulatory concept.

It shifts the emphasis from demonstrating that an AI system was acceptable at one point in time toward demonstrating that it remains acceptable throughout operation.

What Happens When the Foundation Model Changes?

Another difficult issue addressed by FDA concerns applications built on third-party foundation models.

A medical-device manufacturer may build a validated application using an external foundation model, but the model provider may subsequently change that model.

The device manufacturer may therefore face a GMP-like change-control problem even though the change originated outside its own organization.

FDA discusses mechanisms ranging from documentation within the manufacturer’s Quality Management System to regulatory authorization and the use of Predetermined Change Control Plans (PCCPs). Re-benchmarking against the original competency criteria could provide evidence that the modified system continues to perform acceptably.

FDA is also considering the concept of voluntary Foundation Model Device Master Files, through which foundation-model developers could provide FDA with information about model architecture, training-data provenance, known limitations, failure modes, performance, guardrails, model updates and audit-log availability.

Agentic AI Raises the Risk Further

FDA specifically addresses agentic AI — systems capable of autonomously planning and executing multi-step tasks, using external tools and taking actions.

The Agency asks how increased autonomy and the reduced opportunity for human review should influence acceptance criteria and regulatory oversight.

This could become increasingly important as AI moves from:

providing information → recommending decisions → executing decisions.

The regulatory significance of AI therefore may increasingly depend not simply on the sophistication of the model, but on how much authority the system is given to act.

Why This Matters Beyond Medical Devices

The FDA discussion paper is not GMP guidance and does not establish requirements for AI used in pharmaceutical manufacturing or Pharmaceutical Quality Systems. FDA explicitly states that the document is for discussion only and does not represent draft or final guidance or proposed regulatory expectations.

Nevertheless, the regulatory thinking behind the paper deserves attention from pharmaceutical companies.

Several concepts could be highly relevant to future approaches for GenAI used in GxP environments:

Intended use → Risk classification → Defined competencies → Predefined acceptance criteria → Qualification/benchmarking → Human confirmation → Continuous monitoring → Re-benchmarking after change

This may ultimately prove more suitable for probabilistic AI than attempting to force GenAI into a traditional deterministic software-validation model.

The particularly important message is that probabilistic output does not necessarily mean that an AI system cannot be controlled.

Instead, control may need to be demonstrated differently: through risk-based performance requirements, validated boundaries and guardrails, representative challenge testing, human oversight, predefined acceptance criteria and continuous lifecycle monitoring.

In this sense, the FDA paper may represent another important step toward a regulatory model based not on demanding that AI behave like deterministic software, but on demonstrating that its performance remains acceptably controlled for its intended use.

Source

U.S. Food and Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback, 18 August 2026.
https://www.fda.gov/ … on-paper-and-request

Friday, August 7, 2026

New MHRA-backed survey: pharma is exploring AI in CMC, but regulated adoption remains limited


Introduction An MHRA-backed survey finds pharmaceutical AI adoption remains exploratory, with regulatory uncertainty limiting AI use in cGMP, GDP and CMC submissions.

A new open-access industry survey provides one of the clearest recent pictures of how digital and AI-enabled technologies are actually being used in pharmaceutical development and manufacturing.

Published online in July 2026 in the International Journal of Pharmaceutics, the study was conducted by the Digital CMC Centre of Excellence in Regulatory Science and Innovation (CERSI), led by the University of Strathclyde and supported by the UK Medicines and Healthcare products Regulatory Agency (MHRA).

The results reveal an important gap between interest in AI and its actual deployment in regulated pharmaceutical environments.

While many respondents reported experience with AI-enabled tools, only 6% of responses indicated current AI use in cGMP/GDP environments. At the same time, 69% indicated that such use was being considered.

This suggests that pharma interest in AI is already substantial, but companies remain cautious about moving AI from development or exploratory applications into regulated manufacturing and quality operations.

Regulatory uncertainty appears to be a major barrier

One of the most interesting findings concerns regulation.

According to the study, 69% of respondents somewhat or strongly agreed that the current pharmaceutical regulatory environment represents a barrier to implementation of AI tools.

The authors identify several areas where uncertainty remains particularly important: explainability, validation, model drift, lifecycle management and the evidence regulators may expect during submissions or inspections.

The study also found a major difference between internal implementation and regulatory use: fewer than 15% of reported digital and AI-enabled tools had been included in regulatory submissions.

The authors appropriately caution that not every internally used digital tool would be expected to appear in a regulatory submission. Nevertheless, the difference indicates that translating useful internal models into regulator-facing applications remains difficult.

Why this matters for pharmaceutical companies

The findings challenge a common assumption that the main obstacle to pharmaceutical AI adoption is technology.

The models may already be capable of performing useful tasks. The bigger challenge may increasingly be demonstrating that they can operate within a controlled pharmaceutical lifecycle.

For an AI system influencing CMC or GMP activities, companies may need to demonstrate:

  • a clearly defined intended use and context of use;
  • data quality and data-governance controls;
  • appropriate validation or qualification;
  • performance acceptance criteria;
  • explainability appropriate to the risk and intended use;
  • management of model drift and performance deterioration;
  • human oversight and accountability;
  • change control for models, configuration and data;
  • ongoing lifecycle monitoring;
  • evidence suitable for regulatory submissions and inspections.

The study therefore supports an important distinction between AI that works and AI that is regulatory-ready.

A model may provide impressive technical performance and still be difficult to deploy in GMP if the company cannot demonstrate its validation status, data provenance, lifecycle controls and behaviour under foreseeable failure conditions.

Quality by Digital Design

The paper also discusses the emerging concept of Quality by Digital Design (QbDD), building on the established Quality by Design philosophy.

This may become an important direction for pharmaceutical digital transformation. Instead of adding digital tools or AI after a process has already been designed, digital models, data architecture and control strategies could increasingly become part of process and product development from the beginning.

Such an approach could make regulatory justification easier because model purpose, data requirements, risk controls and lifecycle expectations could be designed together with the pharmaceutical process rather than added retrospectively.

Practical takeaway

The survey suggests that pharma does not primarily lack interest in AI. It lacks sufficient regulatory certainty and practical implementation experience for higher-impact regulated applications.

For companies considering AI in manufacturing or quality systems, the most valuable next step may therefore not be another AI proof-of-concept.

It may be to take one well-defined use case and demonstrate the complete regulated lifecycle: intended use, risk assessment, data governance, challenge testing, qualification, implementation, monitoring, change control and requalification.

That experience may ultimately be more valuable than operating numerous disconnected AI pilots that never progress into regulated use.

Source

Cook G. et al. “Digital and AI-enabled models in pharmaceutical development and manufacturing: a regulatory-focused industry survey.” International Journal of Pharmaceutics, 2026. Open access.
https://www.scienced … ii/S0378517326006253

Monday, March 9, 2026

EU regulators formalize industry dialogue on AI across the medicines lifecycle (HMA-EMA AI group)


Introduction HMA and EMA formalize industry dialogue on AI across the medicines lifecycle, covering clinical development, pharmacovigilance, GMP Annex 22 and risk-based AI governance.

EMA has published the event page and supporting documents for the HMA–EMA AI group meeting with industry stakeholders (February 2026) and subsequently posted summary notes. This is a concrete regulatory mechanism for aligning expectations on acceptable AI use, governance, and evidence—also in areas that can impact GMP/CMC and lifecycle data.

Notably, the current guidance development activities mentioned include:

  • Guidance on AI in clinical development (a concept paper is expected before a full draft guideline).
  • Guidance on AI in pharmacovigilance, to be developed jointly with PRAC as a Q&A-style document.
  • EU GMP Annex 22 on AI in manufacturing: following public consultation (≈1,300 comments received), it is now under revision, with the final document expected by the end of the year.
  • Several industry interventions were noted, including an expectation for a more flexible approach to the use of AI in GMP. This includes the potential use of generative AI and large language models (LLMs) in critical GMP applications, provided this is supported by a robust, risk-based framework.
2AI_forPharmaCommunicationPlatform.png

Event page (EMA): https://www.ema.euro … olders-february-2026
Summary notes PDF (EMA): https://www.ema.euro … february-2026_en.pdf

Thursday, February 5, 2026

Key AI GMP-relevant documents


Introduction Explore key AI and GMP regulatory documents from FDA, EMA, EU GMP, PIC/S and MHRA covering validation, data integrity, risk management and AI lifecycle governance.

Key AI GMP-relevant documents where regulators explicitly address AI / ML or the core compliance expectations that govern AI used in manufacturing (computerized systems, validation/assurance, data integrity, lifecycle control) are listed below with relevant links

    Disclaimer:

  • Last review and links update : 2026-08-31, note: links to draft documents may stop working after the consultation phase ends.
  • The documents listed below have different regulatory status. Some are binding requirements, while others are guidance, discussion papers, or draft documents under consultation or revision.

FDA (US) — AI in pharma manufacturing & quality

• Artificial Intelligence in Drug Manufacturing (Discussion Paper, 2023) — FDA CDER discussion paper focused on AI use in drug manufacturing (not guidance, but important signal of expectations). https://www.fda.gov/ … edia/165743/download - It also includes many other useful links.
• Considerations for the Use of AI to Support Regulatory Decision-Making for Drug and Biological Products (Draft Guidance, Jan 2025) — FDA risk-based credibility assessment framework for AI models used to support regulatory decisions regarding safety, effectiveness or quality. The framework is based on the model’s specific Context of Use (COU). https://www.fda.gov/ … -drug-and-biological
• Guiding Principles of Good AI Practice in Drug Development (Jan 2026) — FDA + EMA aligned principles; explicitly spans lifecycle including manufacturing. https://www.fda.gov/ … edia/189581/download
• FDA “Artificial Intelligence for Drug Development” hub page (collects the above + related FDA AI resources). https://www.fda.gov/ … nce-drug-development

GMP/quality foundations that apply to AI systems

• Data Integrity and Compliance With Drug CGMP: Q&A (Dec 2018) — the core FDA data integrity expectations that AI systems must meet (ALCOA+, audit trail, controls, governance). https://www.fda.gov/ … uestions-and-answers
• PAT — A Framework for Innovative Pharmaceutical Development, Manufacturing and Quality Assurance (Guidance; PDF) — not “AI”, but foundational for model-based control/monitoring and advanced analytics in manufacturing. https://www.fda.gov/media/71012/download
• Process Validation: General Principles and Practices (Guidance; PDF) — validation lifecycle principles that also govern AI-enabled control/monitoring when it impacts product quality. https://www.fda.gov/ … es-and-Practices.pdf
• Emerging Technology Program (ETP) (CDER) — FDA program supporting innovative manufacturing technologies (relevant pathway when AI is part of novel manufacturing control/automation). https://www.fda.gov/ … chnology-program-etp
• Advanced Manufacturing Technologies (AMT) Designation Program (Final Guidance, Dec 2024) — FDA programme supporting early adoption of advanced manufacturing technologies. Relevant where AI/ML forms part of a novel manufacturing or control technology.
https://www.fda.gov/ … -designation-program

EMA / EU medicines regulators — AI + EU GMP updates

• EMA Reflection paper on the use of AI in the medicinal product lifecycle (final, 9 Sept 2024; PDF) — covers principles across lifecycle and regulatory expectations when AI outputs are used in regulated submissions (incl. manufacturing-related evidence). https://www.ema.euro … uct-lifecycle_en.pdf
• Network Data Steering Group Workplan 2026–2028: Data and Artificial Intelligence in Medicines Regulation (Version 2.0, Feb 2026) — current HMA–EMA strategic workplan for the use of data and AI in medicines regulation. It includes development of AI guidance for clinical development and pharmacovigilance, regulatory AI tools, data standards and coordinated AI implementation across the European medicines regulatory network.
https://www.ema.euro … -steering-group-ndsg

• EMA/FDA: Guiding principles of good AI practice in drug development (Jan 2026; PDF) — joint high-level principles (explicitly spanning manufacturing).https://www.ema.euro … g-development_en.pdf

EU GMP (EudraLex Volume 4) — Computerised Systems and Artificial Intelligence

• EU GMP Annex 11: Computerised Systems (current) — the core existing GMP framework applicable to computerised systems, including AI-enabled systems used in GMP processes.

https://health.ec.eu … x11_01-2011_en_0.pdf

• Revision of Chapter 4, Annex 11 and New Annex 22 – Artificial Intelligence (draft) — the European Commission consultation closed on 7 October 2025. Annex 22 remains under development and is not currently an effective GMP requirement.

https://health.ec.eu … s-chapter-4-annex_en

• Draft Annex 22: Artificial Intelligence — proposes GMP expectations covering intended use, model selection and training, performance metrics, independent test data, validation, operation, change control, performance monitoring and human review.

Status: EMA Annex 22 Multistakeholder Expert Workshop, 30 June–1 July 2026 — EMA collected additional expert evidence to support further development of Annex 22, including discussion of GenAI/LLMs, guardrails, human oversight, lifecycle monitoring, cybersecurity and outsourced/cloud AI.

https://www.ema.euro … development-annex-22

Horizontal EU AI Legislation

• EU Artificial Intelligence Act – Regulation (EU) 2024/1689 — binding horizontal EU legislation governing AI systems. It is not a GMP regulation, but pharmaceutical companies may need to consider its requirements in parallel with sector-specific GxP requirements depending on the AI system, its intended purpose and the organisation’s role in the AI value chain.

https://eur-lex.euro … eu/eli/reg/2024/1689

PIC/S (global GMP inspection cooperation)

• PIC/S PI 041-1: Good Practices for Data Management and Integrity in regulated GMP/GDP environments (final; PDF) — widely relied upon by inspectorates; very relevant for AI data pipelines and governance. https://picscheme.org/docview/4234
• PIC/S PI 011-3: Good Practices for Computerised Systems in Regulated “GxP” Environments (PDF) — inspector-oriented expectations for computerised systems (validation, supplier management, control). https://picscheme.org/docview/3444
• PIC/S PI 006-4: Recommendations on Qualification and Validation (published 30 Jul 2026; effective 1 Oct 2026) — revised PIC/S recommendations covering qualification and validation lifecycle principles. Not AI-specific, but an important supporting framework when AI-enabled technologies form part of qualified equipment, validated processes or GMP systems.
https://picscheme.org/docview/11277

UK MHRA

• Use of AI for GxP Inspection Responses: Setting Standards Without Stifling Innovation (MHRA Inspectorate, 29 Jun 2026) — important practical regulatory position based on actual inspection experience. MHRA reported AI-assisted inspection responses containing non-existent regulatory references and materially inaccurate information. The regulator emphasises human verification, organisational accountability and effective CAPA rather than generic AI-generated responses.
https://mhrainspecto … stifling-innovation/
• MHRA GxP Data Integrity Guidance and Definitions (Rev. 1, March 2018; PDF) — strong practical expectations for data integrity controls that apply directly to AI toolchains. https://assets.publi … rch_edited_Final.pdf
Health Canada
• Annex 11 (GUI-0050): Computerized Systems — Health Canada adoption of Annex 11 principles for GMP computerized systems (useful for “regulatory convergence” arguments). https://www.canada.c … ystems-gui-0050.html

Tuesday, January 27, 2026

What is Artificial Intelligence (AI)?


Introduction Artificial Intelligence (AI) definitions

(EU)
‘AI system’ means a machine-based system that is designed to operate with varying levels of autonomy and that may exhibit adaptiveness after deployment, and that, for explicit or implicit objectives, infers, from the input it receives, how to generate outputs such as predictions, content, recommendations, or decisions that can influence physical or virtual environments;

Source: EU Regulation 2024/1689

***

(US / FDA) A machine-based system that can, for a given set of human-defined objectives, make predictions, recommendations, or decisions influencing real or virtual environments. Artificial intelligence systems use machine- and human-based inputs to perceive real and virtual environments; abstract such perceptions into models through analysis in an automated manner; and use model inference to formulate options for information or action.

Source: 15 U.S.C. 9401(3).