Thursday, August 20, 2026

FDA Proposes a New Way to Evaluate Generative AI: From Exhaustive Testing to Competency-Based Lifecycle Assurance


Introduction FDA proposes a competency-based approach for evaluating Generative AI-enabled medical devices, including risk-based benchmarking, clinical confirmation, lifecycle monitoring, model drift control and re-evaluation after AI changes.

The U.S. FDA has opened an important regulatory discussion on how Generative AI-enabled medical devices could be evaluated when traditional software testing is no longer sufficient.

On 18 August 2026, FDA’s Center for Devices and Radiological Health (CDRH) released the discussion paper Considerations for the Regulation of Generative AI-Enabled Medical Devices. The Agency is requesting stakeholder feedback until 19 October 2026.

The document applies specifically to medical devices, not to pharmaceutical GMP systems. Nevertheless, several concepts discussed by FDA are highly relevant to the broader debate about how probabilistic and generative AI can be validated and controlled in regulated environments.

The most significant idea may be surprisingly simple:

For sufficiently complex GenAI systems, it may be unrealistic to test every possible input and output. Instead, regulators may need to assess whether the system can demonstrate defined competencies and continue to demonstrate them throughout its lifecycle.

Why Traditional Software Testing May Not Be Enough

FDA recognizes that GenAI-enabled devices differ fundamentally from traditional software and even from many conventional AI systems.

They may:

  • accept open-ended inputs;
  • generate variable outputs for similar inputs;
  • perform multiple different subtasks;
  • use third-party foundation models;
  • change because of updates to models, prompts, retrieval mechanisms, guardrails or orchestration logic;
  • operate with increasing levels of autonomy.

FDA also explicitly identifies risks such as confabulations or hallucinations, limited transparency into third-party foundation models and performance degradation over time.

This creates a fundamental validation challenge.

For conventional software with bounded inputs and predetermined outputs, extensive input-output testing can provide strong evidence that the system operates as intended.

For GenAI, the possible combination of inputs, conversations and outputs can become effectively unlimited. FDA therefore acknowledges that evaluating every conceivable situation may simply not be practical.

Risk Depends on Both Autonomy and Consequences

FDA proposes a possible two-axis framework for thinking about GenAI risk.

One dimension considers what the AI actually does — ranging from providing non-directive information through directing an action to taking an action autonomously.

The second dimension considers the consequence of relying on an incorrect output.

Risk therefore increases as AI becomes more autonomous and as the potential consequences of an incorrect output become more serious.

This is an important distinction because the same underlying AI technology could represent very different levels of regulatory risk depending on its intended use.

An AI system that provides general information is fundamentally different from one that recommends a specific clinical action — and different again from an agentic system capable of executing that action.

FDA also emphasizes that simply adding wording such as “talk to your doctor” or “I am not a medical professional” may not necessarily make an otherwise action-directing AI function less directive.

A Competency-Based Approach to GenAI Evaluation

Perhaps the most innovative part of the discussion paper is FDA’s consideration of a competency-based approach inspired, at a high level, by the way human clinicians are evaluated.

Doctors are not qualified by testing every possible clinical situation they could ever encounter. Instead, they demonstrate competencies through examinations, supervised practice and continuing assessment.

FDA is exploring whether a related concept could be adapted for GenAI-enabled medical devices.

The proposed approach consists of two major components:

1. Device benchmarking

The deployed or representative final system would be tested against predefined competencies.

2. Clinical confirmation

Evidence would then confirm that the device performs appropriately under real or clinically representative conditions.

Importantly, FDA proposes evaluating the final user-facing device as configured for deployment, rather than assessing only the underlying foundation model such as an LLM.

This distinction is critical.

The performance of an AI application does not depend only on the model. It may also depend on the system prompt, retrieval architecture, knowledge sources, guardrails, user interface, orchestration logic and other controls surrounding the model.

What Would Be Tested?

FDA identifies several possible areas of competency for benchmarking GenAI systems.

They include:

Safety

  • recognition and escalation of safety-critical situations;
  • maintaining the defined scope and operational boundaries;
  • appropriate communication of uncertainty and deferral when the system cannot provide a reliable answer.

Clinical proficiency

  • knowledge and task fidelity;
  • information gathering and analysis;
  • quantitative reasoning;
  • quality and comprehensibility of communication.

Generalizability

  • robustness, reliability and reproducibility;
  • performance across relevant subgroups.

Agentic capabilities

  • additional competencies where AI can autonomously plan, use tools or execute multi-step actions.

FDA also discusses testing boundary adherence using techniques such as adversarial prompting, prompt injection and multi-turn conversations that gradually move outside the intended scope of the system.

This illustrates an important change in thinking about AI validation:

Validation is not only about whether AI produces correct answers. It is also about whether it behaves safely when it does not know the answer, when it is challenged, and when a user attempts to push it beyond its intended use.

Acceptance Criteria Still Matter

The competency-based approach does not mean abandoning predefined validation requirements.

Quite the opposite.

FDA considers that manufacturers should define the scope of testing based on intended use and risk, prespecify evaluation methods, justify scoring methods and establish acceptance criteria before testing.

Where human expert assessment is needed, FDA also discusses the use of appropriately qualified and structurally independent adjudicators.

This could become particularly important for GenAI because many outputs cannot simply be classified as mathematically “correct” or “incorrect.” Evaluation may require structured expert judgement supported by predefined scoring criteria.

Validation Does Not End at Deployment

FDA places substantial emphasis on postmarket performance monitoring.

Possible approaches include:

  • periodic re-benchmarking against predefined performance thresholds;
  • periodic review of real-world AI outputs by qualified independent clinicians;
  • monitoring for model or data drift and other forms of performance degradation.

Reassessment could occur periodically and following defined triggering events, including changes to the underlying model or other components of the AI architecture.

FDA even asks whether, under appropriate circumstances, greater uncertainty could be accepted before market authorization if it were compensated by stronger postmarket monitoring.

That is a potentially important regulatory concept.

It shifts the emphasis from demonstrating that an AI system was acceptable at one point in time toward demonstrating that it remains acceptable throughout operation.

What Happens When the Foundation Model Changes?

Another difficult issue addressed by FDA concerns applications built on third-party foundation models.

A medical-device manufacturer may build a validated application using an external foundation model, but the model provider may subsequently change that model.

The device manufacturer may therefore face a GMP-like change-control problem even though the change originated outside its own organization.

FDA discusses mechanisms ranging from documentation within the manufacturer’s Quality Management System to regulatory authorization and the use of Predetermined Change Control Plans (PCCPs). Re-benchmarking against the original competency criteria could provide evidence that the modified system continues to perform acceptably.

FDA is also considering the concept of voluntary Foundation Model Device Master Files, through which foundation-model developers could provide FDA with information about model architecture, training-data provenance, known limitations, failure modes, performance, guardrails, model updates and audit-log availability.

Agentic AI Raises the Risk Further

FDA specifically addresses agentic AI — systems capable of autonomously planning and executing multi-step tasks, using external tools and taking actions.

The Agency asks how increased autonomy and the reduced opportunity for human review should influence acceptance criteria and regulatory oversight.

This could become increasingly important as AI moves from:

providing information → recommending decisions → executing decisions.

The regulatory significance of AI therefore may increasingly depend not simply on the sophistication of the model, but on how much authority the system is given to act.

Why This Matters Beyond Medical Devices

The FDA discussion paper is not GMP guidance and does not establish requirements for AI used in pharmaceutical manufacturing or Pharmaceutical Quality Systems. FDA explicitly states that the document is for discussion only and does not represent draft or final guidance or proposed regulatory expectations.

Nevertheless, the regulatory thinking behind the paper deserves attention from pharmaceutical companies.

Several concepts could be highly relevant to future approaches for GenAI used in GxP environments:

Intended use → Risk classification → Defined competencies → Predefined acceptance criteria → Qualification/benchmarking → Human confirmation → Continuous monitoring → Re-benchmarking after change

This may ultimately prove more suitable for probabilistic AI than attempting to force GenAI into a traditional deterministic software-validation model.

The particularly important message is that probabilistic output does not necessarily mean that an AI system cannot be controlled.

Instead, control may need to be demonstrated differently: through risk-based performance requirements, validated boundaries and guardrails, representative challenge testing, human oversight, predefined acceptance criteria and continuous lifecycle monitoring.

In this sense, the FDA paper may represent another important step toward a regulatory model based not on demanding that AI behave like deterministic software, but on demonstrating that its performance remains acceptably controlled for its intended use.

Source

U.S. Food and Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback, 18 August 2026.
https://www.fda.gov/ … on-paper-and-request

Friday, August 7, 2026

New MHRA-backed survey: pharma is exploring AI in CMC, but regulated adoption remains limited


Introduction An MHRA-backed survey finds pharmaceutical AI adoption remains exploratory, with regulatory uncertainty limiting AI use in cGMP, GDP and CMC submissions.

A new open-access industry survey provides one of the clearest recent pictures of how digital and AI-enabled technologies are actually being used in pharmaceutical development and manufacturing.

Published online in July 2026 in the International Journal of Pharmaceutics, the study was conducted by the Digital CMC Centre of Excellence in Regulatory Science and Innovation (CERSI), led by the University of Strathclyde and supported by the UK Medicines and Healthcare products Regulatory Agency (MHRA).

The results reveal an important gap between interest in AI and its actual deployment in regulated pharmaceutical environments.

While many respondents reported experience with AI-enabled tools, only 6% of responses indicated current AI use in cGMP/GDP environments. At the same time, 69% indicated that such use was being considered.

This suggests that pharma interest in AI is already substantial, but companies remain cautious about moving AI from development or exploratory applications into regulated manufacturing and quality operations.

Regulatory uncertainty appears to be a major barrier

One of the most interesting findings concerns regulation.

According to the study, 69% of respondents somewhat or strongly agreed that the current pharmaceutical regulatory environment represents a barrier to implementation of AI tools.

The authors identify several areas where uncertainty remains particularly important: explainability, validation, model drift, lifecycle management and the evidence regulators may expect during submissions or inspections.

The study also found a major difference between internal implementation and regulatory use: fewer than 15% of reported digital and AI-enabled tools had been included in regulatory submissions.

The authors appropriately caution that not every internally used digital tool would be expected to appear in a regulatory submission. Nevertheless, the difference indicates that translating useful internal models into regulator-facing applications remains difficult.

Why this matters for pharmaceutical companies

The findings challenge a common assumption that the main obstacle to pharmaceutical AI adoption is technology.

The models may already be capable of performing useful tasks. The bigger challenge may increasingly be demonstrating that they can operate within a controlled pharmaceutical lifecycle.

For an AI system influencing CMC or GMP activities, companies may need to demonstrate:

  • a clearly defined intended use and context of use;
  • data quality and data-governance controls;
  • appropriate validation or qualification;
  • performance acceptance criteria;
  • explainability appropriate to the risk and intended use;
  • management of model drift and performance deterioration;
  • human oversight and accountability;
  • change control for models, configuration and data;
  • ongoing lifecycle monitoring;
  • evidence suitable for regulatory submissions and inspections.

The study therefore supports an important distinction between AI that works and AI that is regulatory-ready.

A model may provide impressive technical performance and still be difficult to deploy in GMP if the company cannot demonstrate its validation status, data provenance, lifecycle controls and behaviour under foreseeable failure conditions.

Quality by Digital Design

The paper also discusses the emerging concept of Quality by Digital Design (QbDD), building on the established Quality by Design philosophy.

This may become an important direction for pharmaceutical digital transformation. Instead of adding digital tools or AI after a process has already been designed, digital models, data architecture and control strategies could increasingly become part of process and product development from the beginning.

Such an approach could make regulatory justification easier because model purpose, data requirements, risk controls and lifecycle expectations could be designed together with the pharmaceutical process rather than added retrospectively.

Practical takeaway

The survey suggests that pharma does not primarily lack interest in AI. It lacks sufficient regulatory certainty and practical implementation experience for higher-impact regulated applications.

For companies considering AI in manufacturing or quality systems, the most valuable next step may therefore not be another AI proof-of-concept.

It may be to take one well-defined use case and demonstrate the complete regulated lifecycle: intended use, risk assessment, data governance, challenge testing, qualification, implementation, monitoring, change control and requalification.

That experience may ultimately be more valuable than operating numerous disconnected AI pilots that never progress into regulated use.

Source

Cook G. et al. “Digital and AI-enabled models in pharmaceutical development and manufacturing: a regulatory-focused industry survey.” International Journal of Pharmaceutics, 2026. Open access.
https://www.scienced … ii/S0378517326006253

Wednesday, July 15, 2026

MHRA sets clear expectations for AI-assisted GxP inspection responses


Introduction MHRA sets expectations for AI-assisted GxP inspection responses, requiring accurate evidence, technical review, human approval and accountable CAPA development.

The UK Medicines and Healthcare products Regulatory Agency has published an important statement on the use of artificial intelligence when preparing responses following GxP inspections.

The MHRA confirmed that companies are already using AI tools to draft inspection responses. The regulator recognises that AI can help explain complex technical matters, improve consistency and accelerate routine drafting. However, it has also identified cases in which inappropriate AI use created real regulatory and patient-safety risks.

According to the MHRA, some AI-assisted inspection responses included references to guidance that did not exist, citations of inappropriate regulatory frameworks and inaccurate information. In one case, a response exceeded 90 pages but still failed to address the identified deficiencies. In another case involving a serious patient-safety concern, inaccurate AI-generated information increased the regulator’s review time from approximately four hours to more than 20 hours.

This is an important development because it moves the risk of AI hallucination in GxP documentation from theory into documented regulatory experience.

MHRA is not prohibiting AI

The MHRA’s position is not that companies should stop using AI. Its concern is whether information submitted to the regulator is accurate, verifiable, technically reviewed and prepared under appropriate oversight.

The regulator states that all inspection responses and related submissions must be:

  • factually accurate and verifiable;
  • technically reviewed by appropriately experienced personnel;
  • approved by a person with sufficient authority and accountability;
  • supported by evidence for factual claims;
  • appropriate to the specific regulatory and organisational context.

These expectations apply regardless of whether the document was prepared using AI, templates, consultants or conventional manual drafting.

Voluntary disclosure of AI use

The MHRA is also offering companies the option to disclose when AI has been used to support an inspection response. Disclosure is currently voluntary.

Where a company chooses to disclose AI use, the MHRA recommends:

  • including a brief statement at the beginning of the response;
  • identifying the sections in which AI assistance was used;
  • confirming that the content was verified and approved by humans.

The MHRA indicates that transparent disclosure, combined with effective verification, may demonstrate a mature and open quality culture.

This does not mean that disclosure compensates for inaccurate information. Responsibility remains with the organisation and the accountable personnel approving the response.

What may indicate inadequate AI oversight?

The MHRA identifies several warning signs that may indicate weak verification or quality-system controls:

  • incorrect statements or non-existent references;
  • generic language that does not address the specific deficiency;
  • lack of organisation-specific information;
  • citations of regulations without explaining their relevance;
  • inconsistent technical terminology;
  • excessively long responses that fail to address the issue clearly.

Where inspection responses are inaccurate, incomplete or unnecessarily verbose, the MHRA may reject them or return them for revision.

Weak responses may also affect the regulator’s assessment of the company’s CAPA system and future inspection risk. Serious or repeated problems may be referred for further regulatory action.

Practical implications for pharmaceutical companies

Companies using AI to prepare responses to inspections, audits or regulatory deficiencies should establish a controlled review process.

At minimum, this should include:

  • use of approved regulatory and company source documents;
  • verification of every cited regulation, guidance document and factual claim;
  • confirmation that the response addresses the specific observation;
  • technical review by subject-matter experts;
  • Quality Unit review and formal approval;
  • removal of generic or unsupported AI-generated statements;
  • retention of evidence supporting proposed CAPAs;
  • clear accountability for the final submitted response.

AI may help organise information and improve drafting efficiency, but it cannot replace root-cause analysis, technical understanding or knowledge of the company’s processes.

A well-written response that does not address the actual cause of a deficiency is not an effective CAPA response.

Key takeaway

The MHRA’s message is pragmatic and important: regulators do not need to prohibit AI to control its risks.

The regulatory expectation remains focused on outcomes. Inspection responses must be accurate, evidence-based, organisation-specific and approved by accountable personnel.

For pharmaceutical companies, this means that AI-assisted regulatory writing should be treated as a controlled quality process—not as an informal administrative shortcut.

Source:
https://mhrainspecto … stifling-innovation/

banner2.png

Saturday, July 11, 2026

EMA Annex 22 AI workshop: awaiting official conclusions, but the direction of discussion is important


Introduction EMA’s Annex 22 AI workshop highlights a possible shift from prohibiting GenAI and LLMs in GMP toward risk-based control using guardrails, human oversight, traceability, validation, lifecycle monitoring and pharmaceutical quality system governance.

Following the EMA GMP multistakeholder workshop organised on 30 June and 1 July 2026 on expert contributions to the development of EU GMP Annex 22 on Artificial Intelligence, the pharmaceutical industry is still awaiting official feedback, conclusions or a post-workshop report from EMA.

The workshop was organised to support the development of Annex 22 and to collect expert input on the possible use of artificial intelligence in medicines manufacturing. The first day was available as a live broadcast and included expert opinions on the future use of generative AI, large language models and other probabilistic or adaptive AI models in GMP environments.

One important message emerging from the expert discussion was that a simple, general prohibition of GenAI or LLMs in pharmaceutical GMP environments may not be the most effective or future-proof regulatory approach. Such a prohibition could quickly become outdated, especially as AI models, control mechanisms, guardrails, validation methods and monitoring tools continue to improve.

At the same time, this does not mean that GenAI or LLMs should be freely accepted in critical GMP applications. The more balanced and practical direction appears to be a risk-based approach: AI should be considered according to its intended use, GMP impact, patient risk, level of human oversight, data quality, traceability, model behaviour, validation evidence and lifecycle controls.

This is particularly important because the original draft Annex 22 indicated that dynamic, adaptive and probabilistic models, including GenAI and LLMs, should not be used in critical GMP applications. However, stakeholder feedback showed support for potentially enabling these technologies in medicines manufacturing if adequate control and mitigation measures can be demonstrated.

For pharmaceutical companies, the practical message is clear: the discussion is moving from “AI should be prohibited” toward “under what conditions can AI be controlled well enough for GMP use?”

The key control areas are likely to include:

  • clear definition of intended use;
  • GMP impact and patient-risk assessment;
  • approved and controlled source data;
  • model/version control;
  • guardrails and their verification;
  • human oversight and accountability;
  • traceability of AI outputs;
  • detection of hallucinations or incorrect recommendations;
  • incident escalation and prevention of GMP impact;
  • performance monitoring and drift detection;
  • supplier qualification and cloud-service oversight;
  • change control for model updates, retraining and configuration changes;
  • documented validation or qualification evidence.

One of the most important concepts discussed in the context of GenAI and LLMs in GMP is the use of “guardrails”. What could “guardrails” mean for AI in GMP?

In simple terms, guardrails are predefined controls built around an AI system to keep its use within safe, intended and acceptable boundaries.

In GMP language, guardrails can be understood as technical, procedural and human controls that prevent AI from being used in the wrong way, reduce the risk of incorrect or unsupported outputs, and stop AI-generated conclusions from becoming GMP decisions without appropriate human review.

Guardrails do not make AI perfect. They do not remove the need for validation, qualification, human oversight or quality risk management. Their purpose is to ensure that AI remains a controlled support tool, not an uncontrolled authority.
Guardrails are predefined technical, procedural and organizational controls that define what an AI system is allowed to do, what it must not do, how it should behave under uncertainty, and how its outputs are reviewed before they affect GMP activities, records or decisions.

Examples of AI guardrails in GMP

  • Intended-use limitation – the AI system may only be used for defined and approved purposes. For example, an AI tool may summarize SOP content, but may not approve a deviation, assign final root cause or make a batch disposition decision.
  • Source grounding – the AI must generate answers only from approved and controlled sources, such as current SOPs, specifications, batch records, validation reports, quality agreements or approved regulatory guidance.
  • Refusal rules – the AI must not guess when evidence is missing. For example, the system should state: “Insufficient information is available to conclude product impact” instead of generating an unsupported conclusion.
  • Output constraints – the AI output should be limited to predefined sections or fields, such as facts identified, source references, missing information, potential questions for review and recommended human follow-up. It should not create final GMP conclusions unless this is explicitly validated and approved for the intended use.
  • Human review gate – AI-generated outputs should be reviewed by qualified personnel before they are used in GMP records or decisions. For example, QA should review an AI-generated deviation summary before it is entered into the official investigation record.
  • Access control – only trained and authorized users should be able to use AI functions with potential GMP impact. For example, a batch review assistant should be available only to trained QA or manufacturing reviewers.
  • Audit trail – the system should retain evidence of AI use, including prompt, output, source documents, model or knowledge-base version, reviewer edits and final approval.
  • Change control – changes to the AI model, prompts, configuration, guardrails or knowledge base should be assessed for GMP impact. A new model version or major knowledge-base update may require impact assessment and requalification.
  • Performance monitoring – AI performance should be monitored in routine use. Examples include tracking false citations, unsupported claims, reviewer corrections, repeated failure modes and cases where AI output was rejected by users.

Practical example: AI support for deviation investigations

For an AI tool supporting deviation investigations, guardrails could be defined as follows:

The AI may identify relevant facts, summarize the event chronology, list missing information, identify potentially similar historical deviations and suggest questions for the investigator to consider. The AI must use only approved QMS records and controlled source documents. It must not assign the final root cause, conclude product impact, determine CAPA effectiveness or recommend batch disposition. If the available information is insufficient, the AI must clearly state that no conclusion can be made. Any AI-generated text must be reviewed, corrected where needed and approved by the investigation owner and QA before inclusion in the GMP record.

This example illustrates the practical meaning of guardrails: the AI can support the process, but it cannot replace GMP responsibility. The final decision remains with qualified personnel and must be justified by evidence.

The final Annex 22 position is not yet known. However, the workshop confirms that regulators are actively considering how to balance innovation with GMP control. For industry, this is a strong signal to start preparing practical AI governance frameworks now, rather than waiting until the final Annex 22 text is published.

AI in GMP should not be treated as an informal tool or uncontrolled black box. If AI supports or influences regulated manufacturing, quality or compliance decisions, it should be governed within the pharmaceutical quality system.

Previous note on the planned EMA workshop:
https://www.aiforpha … ment-of-ai-annex-22/

EMA workshop page:
https://www.ema.euro … development-annex-22

Friday, July 10, 2026

Digital standards and machine-readable compendial methods may become a key enabler for AI in QC laboratories


Introduction Machine-readable compendial methods and digital standards may become key enablers for AI-ready pharmaceutical QC laboratories by improving traceability, version control, GMP data integrity and integration with LIMS, LES and ELN systems.

A recent Pharmaceutical Technology article discusses how digital standards can modernise pharmaceutical quality assurance without compromising confidence. This follows the launch of USP MethodConnect, a machine-readable library of USP-NF test methods designed for integration into LIMS, LES and ELN systems.

This may look like a laboratory informatics topic, but it is highly relevant to AI in pharma. AI and automation depend on reliable, structured and controlled data. If compendial methods are manually retyped from PDF or paper into laboratory systems, there is a risk of transcription errors, inconsistent interpretation and weak traceability.

Machine-readable standards reduce that risk by allowing trusted quality requirements to be integrated directly into digital laboratory workflows. This creates a stronger foundation for automation, advanced analytics and future AI-supported QC processes.

Why this matters for pharma and GMP:

  • AI-ready laboratories need structured, trusted and machine-readable quality data.
  • Manual transcription of compendial methods into digital systems creates compliance risk.
  • Digital standards can improve traceability, version control and consistency.
  • Machine-readable methods can support LIMS, LES, ELN and automated review workflows.
  • Trusted digital standards may help AI systems remain anchored to approved scientific and compendial sources.

Practical takeaway:

For QC laboratories, AI readiness is not only about buying AI software. It is also about building a controlled digital foundation.

Companies should assess:

  • how compendial methods are transferred into laboratory systems;
  • how method versions are controlled;
  • whether manual transcription creates data integrity risks;
  • whether LIMS, LES, ELN and CDS systems can use structured method content;
  • how future AI tools will access controlled and approved source information;
  • whether digital standards can reduce ambiguity during method execution, review and inspection.

The long-term message is important: reliable AI in QC will depend on reliable digital standards. If the source content is controlled, machine-readable and traceable, AI-supported workflows become easier to justify and govern.

Source:
https://www.pharmtec … promising-confidence

Additional source on USP MethodConnect:
https://www.labmanag … y-to-the-bench-35443