Friday, August 28, 2026

PDA: AI in Quality Metrics Should Complement – Not Replace – the Quality System


Introduction PDA outlines how AI can strengthen pharmaceutical quality metrics and predictive oversight while remaining embedded within validated data, Quality Systems and human decision-making.

A new article published by the Parenteral Drug Association on 20 August 2026 provides a particularly practical view of how artificial intelligence could be integrated into pharmaceutical quality management.

The article, Using Artificial Intelligence and Metrics, was prepared as part of the PDA Quality Management Maturity Team’s series on metrics implementation. Its authors represent Bristol Myers Squibb, HarborView and Eli Lilly.

Rather than focusing on AI as a standalone technology, the article places AI inside an existing pharmaceutical Quality System.

This distinction is important.

AI as a Quality Intelligence Layer

The authors describe how increasingly contextualized and high-integrity data can allow AI to support more proactive quality and regulatory decision-making.

Potential applications include:

  • real-time monitoring of quality metrics;
  • earlier detection of deviations, trends and anomalies;
  • predictive identification of potential quality problems;
  • continuous process insights;
  • dynamic prioritization of investigations based on emerging trends.

The concept therefore goes beyond using AI simply to summarize documents or automate administrative work.

AI could become an analytical layer over established quality data, helping Quality organisations identify signals earlier and potentially move from retrospective review toward predictive oversight.

But AI Cannot Compensate for Weak Data

One of the strongest messages in the PDA article is that AI-driven metrics depend on structured and reliable data.

The authors emphasize the importance of structured, validated data models so that metrics remain standardized, traceable and auditable.

This is particularly relevant in GMP environments.

An advanced algorithm applied to inconsistent, poorly contextualized or unreliable data does not create a mature quality system. It can instead accelerate incorrect conclusions.

For pharmaceutical companies, this suggests that AI implementation should often start not with selecting an AI model, but with evaluating:

  • data quality;
  • data structure;
  • data ownership;
  • metric definitions;
  • traceability;
  • data integration between systems;
  • reliability of the underlying quality indicators.

AI Should Reinforce Existing Metrics, Not Replace Them

The PDA authors propose that AI-generated insights should be linked to existing quality metrics rather than creating an independent decision framework.

They highlight several important control principles:

  • AI models should be validated as decision-support tools under defined conditions;
  • outputs should be sufficiently explainable for quality and regulatory review;
  • inputs, processing and outcomes should remain traceable;
  • appropriate audit trails and system controls should be maintained;
  • AI should be integrated into existing governance and management processes.

This creates an important boundary:

AI should complement the Pharmaceutical Quality System rather than become a parallel Quality System.

An Interesting Warning: Do Not Optimize the Wrong Metric

One particularly important part of the article concerns the risk of relying on a narrow set of metrics.

AI systems are very effective at optimizing measurable objectives.

But if an organisation measures the wrong thing – or relies too heavily on one indicator – optimisation can produce unintended behaviour.

For example, reducing investigation closure time could appear positive.

However, if the metric becomes dominant, it could unintentionally encourage faster but less thorough investigations.

Similarly, reducing deviations is not necessarily evidence of improved process performance if problems are being classified or reported differently.

The authors therefore recommend a multi-metric, context-aware approach, with metrics periodically reassessed and stress-tested against possible gaming or unintended effects.

This concept is particularly important for AI because algorithms may amplify the consequences of poorly designed performance indicators.

What This Means for Pharmaceutical Quality Systems

The article suggests a potentially important direction for AI implementation in pharma:

AI should not replace established Quality Management Review – it should make it more intelligent.

A future AI-supported Quality System could continuously analyze:

  • deviations;
  • CAPAs;
  • complaints;
  • OOS and OOT results;
  • process capability;
  • environmental monitoring;
  • audit observations;
  • supplier performance;
  • training effectiveness;
  • change controls;
  • recurring quality signals.

Instead of reviewing each metric independently, AI could identify relationships between them and detect weak signals that are difficult to see during conventional periodic review.

But such a system would still require validated data, appropriate model controls, human oversight and clear accountability.

Why This Publication Is Important

Much of the current discussion around pharmaceutical AI focuses on validating algorithms.

The PDA article points toward another important dimension:

AI effectiveness may depend as much on the maturity of the surrounding Quality System and data architecture as on the performance of the AI model itself.

This is consistent with an emerging view that pharmaceutical AI should not be implemented as an isolated technology project.

Instead, it should become part of:

Data Governance → Quality Metrics → AI Analytics → Human Review → Management Decision → Continuous Improvement

The AI model is only one component of that chain.

Regulatory Perspective

The PDA publication is an industry article and does not establish a new regulatory requirement.

However, its recommendations align with several themes increasingly visible in regulatory discussions: defined context of use, trustworthy data, validation, traceability, explainability, human oversight and lifecycle governance.

For pharmaceutical companies considering AI in Quality Systems, the article provides a useful practical bridge between AI technology and established ICH Q10-style quality management principles.

Source

PDA – Using Artificial Intelligence and Metrics, published 20 August 2026:

https://pda.org/pda- … lligence-and-metrics

FDA Draft Guidance – Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products:

https://www.fda.gov/ … -drug-and-biological

Saturday, August 22, 2026

Different Concepts of AI Control in GMP and Pharma


Introduction Six different concepts of AI control in GMP and pharma: CSV/GAMP validation, context of use, algorithmic state of control, guardrails, competency-based evaluation and AI governance.

Several clearly different concepts of AI control can be distinguished in the GMP/pharma environment. They share a common foundation, but [b]they are not the same approach

They differ mainly in what they consider to be the primary object of control: the computerized system, the model, the model output, AI behaviour, or the entire system of “AI + guardrails + human + process”.

In my view, we can currently distinguish at least six approaches:

  • Traditional CSV / GAMP lifecycle – control of system functions; classical validation
  • Context-of-Use / Model Credibility – credibility of the model for a specific application; risk-based
  • Performance / State of Control – maintaining model performance over time; lifecycle monitoring
  • Guardrails + Human Oversight – control of risky GenAI behaviours
  • Competency-Based Evaluation – assessment of AI’s ability to perform defined classes of tasks
  • Enterprise AI Governance – control of the entire AI ecosystem and organisational accountability


1. Traditional CSV/GAMP Approach: “AI as a Computerized System”

This is the most conservative approach.

The starting point is:

Intended Use → Requirements → Risk Assessment → Specification → Verification/Validation → Change Control → Periodic Review

A large part of traditional CSV principles still applies: security, access control, audit trail, data integrity, interfaces, backup, configuration management, etc.

The ISPE GAMP approach to machine learning also emphasizes that much of the traditional lifecycle for computerized systems remains applicable to ML. At the same time, some ML applications require validation to be extended to include performance metrics, an appropriate validation dataset, version control and monitoring.

This approach can be summarized as:

Prove that the system is fit for intended use.

The problem appears with GenAI: how do we verify all possible responses if the output is probabilistic and practically unlimited?

Source:

https://ispe.org/pha … pts-machine-learning


2. FDA Context-of-Use / Model Credibility: “Do Not Ask Whether AI Is Good – Ask Whether It Is Sufficiently Credible for a Specific Decision”

This is already a different philosophy.

FDA, in its 2025 draft guidance, proposes a risk-based credibility assessment framework.

The most important element is Context of Use (COU).

The same AI model may require a completely different level of evidence if:

  • in one application it only helps search for information;
  • in another it influences a quality assessment;
  • in a third it provides information used in regulatory decision-making.

FDA indicates that the stringency of credibility assessment, acceptance criteria, oversight and documentation should be proportional to model risk and the specific COU.

This approach can be summarized as:

Prove that the model output is sufficiently credible for this particular decision.

This is subtly different from classical system validation.

Source:

https://www.fda.gov/ … -drug-and-biological


3. “AI State of Control”: Validation Is Not Enough – We Need to Demonstrate That AI Continues to Perform Properly

This is a concept that, in my view, will become increasingly important.

Traditional validation answers:

Was the system operating correctly at the time of validation?

For AI, a more important question becomes:

Is it still operating correctly after 3, 6 or 12 months?

This introduces:

  • model drift;
  • data drift;
  • performance degradation;
  • override rate;
  • false positives;
  • false negatives;
  • uncertainty;
  • anomalous outputs;
  • retraining triggers.

ISPE already points to the need to update verification and validation when ML changes and to use change management, version control and monitoring.

This approach can be described as:

Validate once + continuously demonstrate maintained performance.

In traditional GMP, the analogy would be Continued Process Verification.

This is where a very attractive concept emerges:

Algorithmic State of Control.


4. EMA Annex 22: Guardrails + Human Oversight

This is another way of thinking, particularly important for GenAI and LLMs.

In the Annex 22 workshop, EMA did not ask only:

Does the model provide correct answers?

but also:

What happens when the model provides an incorrect answer?

Therefore, the object of control becomes the entire protective system:

AI → guardrail → uncertainty detection → escalation → human review → final GMP decision

EMA has explicitly raised questions concerning:

  • the ability to detect hallucinations and fabricated information;
  • effectiveness of guardrails;
  • human-in-the-loop oversight;
  • escalation;
  • stress testing;
  • failure analysis;
  • drift;
  • behaviour of guardrails after model updates;
  • situations in which even guardrails + human oversight may be insufficient.

This can be summarized as:

Do not rely solely on making AI perfect. Build a controlled environment around imperfect AI.

This is a very important philosophical difference.

We are not trying to prove that GenAI will never make a mistake.

We are trying to prove that:

the error will be detected, stopped or corrected sufficiently often before it creates a GMP impact.

Source:

https://www.ema.euro … development-annex-22


5. Competency-Based Evaluation: Probably the Most Radically Different Approach

This is a new FDA concept from August 2026 concerning GenAI-enabled medical devices.

It is not GMP guidance, but conceptually it is very interesting for pharma.

FDA identifies the problem: for GenAI, it may be unrealistic to test all possible input-output combinations.

It therefore considers:

Device Benchmarking → Clinical Confirmation → Postmarket Monitoring

and assessment of AI “competencies”, such as:

  • knowledge/task fidelity;
  • safety behaviour;
  • quantitative reasoning;
  • communication;
  • generalizability;
  • appropriate deferral;
  • adherence to operational boundaries.

At a high level, FDA compares this idea with the way we assess the competence of physicians – we do not test every possible situation that may occur during the next 30 years of clinical practice.

A similar concept could potentially be adapted for GenAI.

This can be summarized as:

Do not validate every possible answer; demonstrate that the AI possesses and maintains the required competencies.

This is very different from traditional CSV.

And, in my view, potentially very interesting for GMP as well.

Source:

https://www.fda.gov/ … on-paper-and-request


6. Governance / PQS Approach: We Control Not Only the AI, but the Organisation Using AI

Another perspective says that validating the model itself is insufficient.

We need to control:

AI inventory → classification → ownership → supplier → data → intended use → risk → validation → access → human accountability → change → monitoring → incidents → CAPA → retirement

In other words, AI becomes part of the Pharmaceutical Quality System.

In this approach, we are focusing not if the model have eg. 95% accuracy but if the the organisation have a system ensuring that AI is used only where it has been approved, by the right people, for the right purpose, under appropriate controls and with clear accountability?

This fits very well with both QRM/PQS and the direction of Annex 22, where EMA discusses lifecycle QRM, accountability, supplier qualification, change-control visibility and outsourced/cloud AI.


Are These Concepts Competing With Each Other?

Probably not. They overlap and complement each other.

They can actually be arranged as layers:

Governance

Intended Use / Context of Use / Risk

System Validation

Model Performance / Competency

Guardrails

Human Oversight

Operational Monitoring / State of Control

This may represent the most mature model of AI control in GMP.

Not simply:

“Validate AI”

but:

“Establish a multi-layered control system ensuring that AI remains appropriate for a specific intended use and that its potential errors do not lead to unacceptable GMP impact.”


The Common Denominator of All Approaches

Despite different names, practically all contemporary concepts converge around several principles:

Intended Use → Risk → Data → Performance → Controls → Human Accountability → Change Control → Continuous Monitoring

They differ mainly in where they place the centre of gravity.

The industry may therefore gradually move away from the question:

“How do we validate AI?”

towards:

“How do we keep AI in a demonstrable state of control?”

That is a much broader and potentially more appropriate concept for AI used in pharmaceutical GMP environments.


Regulatory Status Note

The concepts described above originate from sources with different regulatory status. Some are based on established GxP and GAMP principles, while others originate from draft guidance, regulatory workshops or discussion papers. In particular, the FDA competency-based evaluation concept relates to GenAI-enabled medical devices and does not constitute a GMP requirement. These concepts should therefore be considered as different approaches and emerging regulatory thinking rather than as one harmonised regulatory framework for AI in GMP.

Friday, August 21, 2026

FDA Experience Shows Model-Based Pharmaceutical Manufacturing Is Already Regulatory Reality


Introduction FDA experience with 15 approved drug applications shows how manufacturing process models are assessed according to intended use, risk and impact, with important lessons for AI in GMP.

Artificial intelligence in pharmaceutical manufacturing is often discussed as something regulators will need to address in the future.

A new FDA-authored paper provides an important reminder that model-based manufacturing is already part of regulatory reality.

Published online on 28 July 2026 in the International Journal of Pharmaceutics, the paper FDA regulatory experience with drug manufacturing process models in approved applications reviews FDA experience with manufacturing models submitted as part of approved pharmaceutical applications between 2012 and 2025.

The findings provide unusually concrete insight into how FDA evaluates models that influence pharmaceutical manufacturing and control.

At Least 15 Approved Applications Used Manufacturing Process Models

According to the FDA authors, process models were successfully used in at least:

15 FDA-approved applications from 8 companies between 2012 and 2025.

The applications covered both drug substance and drug product manufacturing and included NDAs and BLAs, including biosimilars, original applications and supplements.

Eleven of the model-supported applications involved continuous manufacturing, where understanding residence-time distribution is important for tracking material through the process.

Models were used for activities including:

  • establishment of design spaces;
  • real-time process control;
  • process monitoring;
  • in-process control;
  • diversion of non-conforming material;
  • release-related decisions.

FDA encountered different types of models, including mechanistic models, empirical models such as NIR/Raman chemometric models, and hybrid approaches.

The Most Important Message: Model Risk Depends on What the Model Does

Perhaps the most useful GMP lesson from the paper is that FDA’s assessment is linked to the risk or impact of the model’s intended use.

The authors describe examples across different impact levels.

High impact:

  • release decisions;
  • parametric control.

Medium impact:

  • in-process controls;
  • diversion of non-conforming material.

Lower impact:

  • definition of design space;
  • investigation and monitoring.

The amount of information included in regulatory submissions was generally commensurate with the determined risk or impact of the model.

This principle has major implications for future AI applications in GMP.

The fundamental regulatory question may not be:

“Is the technology AI?”

but rather:

“What GMP decision does the model influence, and what happens if the model is wrong?”

That distinction is critical.

An AI model identifying unusual process patterns for investigation is not equivalent in risk to an AI model automatically determining whether material is acceptable for release.

The technology may be similar, but the required level of assurance should not necessarily be the same.

FDA Can Examine the Model Very Deeply

The paper also reveals an interesting aspect of regulatory review.

FDA modelling experts are frequently consulted when reviewing applicants’ modelling and control strategies.

In some cases, FDA experts have even developed their own internal models calibrated using experimental data submitted by the applicant.

This demonstrates that model-based regulatory submissions should not be viewed as a black box accompanied only by a headline performance figure.

A regulator may need to understand:

  • model structure;
  • assumptions;
  • input variables;
  • data used to establish the model;
  • operating range;
  • predictive performance;
  • relationship to the control strategy;
  • consequences of model failure.

For future AI/ML applications, explainability may therefore need to mean more than providing an attractive dashboard showing accuracy.

The organisation must be able to demonstrate why the model is appropriate for the specific manufacturing decision it supports.

Does Modelling Slow Regulatory Approval?

Interestingly, FDA’s experience does not suggest that the use of manufacturing models necessarily creates regulatory delay.

After excluding three submissions with issues unrelated to modelling, the model-supported submissions reviewed in the study were approved an average of 20.3 days before their regulatory goal dates, including several applications under accelerated review.

This observation should be interpreted carefully.

It does not prove that using models accelerated approval.

However, it provides useful evidence against the assumption that sophisticated manufacturing models automatically create regulatory obstacles.

When modelling is scientifically justified and properly integrated into a manufacturing and control strategy, regulators clearly have experience assessing and approving such approaches.

Is This an AI Paper?

Not exactly — and this distinction is important.

The FDA paper covers pharmaceutical manufacturing process models, including mechanistic, empirical and hybrid models. These should not all be described as artificial intelligence.

However, the paper is highly relevant to AI and machine learning because it provides practical regulatory precedent for the broader question of how models that influence GMP manufacturing decisions can be assessed.

AI does not eliminate the established regulatory principles used for modelling.

Instead, it makes several of them even more important:

context of use, model risk, data quality, predictive performance, control strategy and lifecycle management.

Lessons for AI/ML Used in GMP Manufacturing

1. What exactly is the intended use?

Does the model:

  • monitor?
  • detect anomalies?
  • recommend action?
  • control a process?
  • divert material?
  • contribute to release?

The answer determines the potential GMP impact.

2. What happens if the model is wrong?

A model can have impressive average accuracy and still create unacceptable risk if rare errors affect critical quality decisions.

Performance criteria therefore should reflect the consequences of different types of error rather than relying only on one overall accuracy figure.

3. What is the validated operating region?

A model should not automatically be assumed to remain reliable for data or manufacturing conditions substantially different from those used during development and validation.

The boundaries of reliable operation need to be understood.

4. Can the model and its decisions be reconstructed?

For regulated applications, companies should be capable of identifying:

  • the model version;
  • relevant input data;
  • output or prediction;
  • configuration;
  • decision subsequently taken;
  • applicable human review.

Traceability becomes increasingly important as models influence higher-risk decisions.

5. How will changes be controlled?

Changes to:

  • algorithms;
  • training data;
  • process data;
  • sensors;
  • model parameters;
  • software infrastructure;
  • manufacturing processes

may affect model performance.

AI/ML therefore needs to be integrated into the Pharmaceutical Quality System rather than managed solely as a data-science project.

6. How will continued performance be demonstrated?

Initial validation shows that a model was suitable when tested.

Lifecycle monitoring needs to demonstrate that it remains suitable.

This becomes particularly important for data-driven models that may be sensitive to changes in raw materials, equipment, sensors, operating practices or process distributions.

A Useful Regulatory Principle for AI in GMP

The FDA experience suggests a practical principle for future AI implementation:

The level of model assurance should be proportional to the influence the model has on product quality and GMP decisions.

This is more useful than treating every AI application as equally risky.

For example:

AI for trend detection → AI for decision support → AI for process control → AI for release

represents increasing influence on the final GMP decision and therefore potentially increasing requirements for validation, transparency, monitoring and governance.

That is consistent with long-established Quality Risk Management principles and may provide one of the most practical foundations for integrating AI into pharmaceutical manufacturing.

The latest FDA experience demonstrates something important:

Regulators are not starting from zero when evaluating advanced models in pharmaceutical manufacturing.

A substantial regulatory foundation already exists.

The challenge for AI will be to extend those principles to models that may be more complex, data-dependent and probabilistic — without losing the fundamental GMP requirement that the process remains understood and controlled.

Source

Fisher AC, Chatterjee S, Madurawe R, Tian G, Tran R, Lee SL. FDA regulatory experience with drug manufacturing process models in approved applications. International Journal of Pharmaceutics. Available online 28 July 2026. Article 127249.

https://www.scienced … ii/S0378517326006976

EU AI Act Enforcement Has Started: What the 2 August 2026 Milestone Means for Pharmaceutical Companies


Introduction EU AI Act enforcement started on 2 August 2026. What pharmaceutical companies should know about AI transparency, GPAI oversight and integration with GxP governance.

The European Union has reached an important milestone in the regulation of artificial intelligence.

From 2 August 2026, the European Commission’s AI Office, together with national competent authorities, has begun enforcing relevant provisions of the EU Artificial Intelligence Act. At the same time, important new transparency requirements for certain AI systems have become applicable.

For pharmaceutical companies, this development deserves attention even though the AI Act is not a GMP regulation.

As artificial intelligence becomes increasingly embedded in pharmaceutical manufacturing, quality systems, regulatory affairs, medical information, pharmacovigilance and other business processes, companies may need to manage two parallel regulatory layers:
AI regulation and pharmaceutical/GxP regulation.

Compliance with one does not automatically demonstrate compliance with the other.

What Changed on 2 August 2026?

The European Commission confirmed that from 2 August the AI Office and national authorities started enforcement of the AI Act.

The same date also marked the application of transparency obligations under Article 50 of the Act. These requirements apply to defined categories of AI systems and are intended to make it clear when individuals interact with AI or encounter AI-generated or manipulated content.

For example, certain interactive AI systems such as chatbots must inform users that they are interacting with AI rather than a human.

The rules also introduce requirements concerning machine-readable marking of certain AI-generated or manipulated content and disclosure requirements for areas such as deepfakes and specified AI-generated content.

These obligations can be relevant to pharmaceutical companies operating public-facing AI applications, including potentially:

  • patient or healthcare-professional chatbots;
  • automated medical-information interfaces;
  • AI-supported customer-service systems;
  • externally published AI-generated material;
  • interactive digital-health applications.

The precise obligation depends on the role of the company, the AI system and its intended use.

General-Purpose AI Is Also Entering a More Serious Enforcement Phase

There is another important change.

Obligations for providers of general-purpose AI (GPAI) models have applied since August 2025. However, from 2 August 2026, the European Commission’s enforcement powers concerning these obligations became applicable, including the possibility of fines.

This primarily affects companies that provide general-purpose AI models rather than ordinary users of commercially available LLM services.

Nevertheless, pharmaceutical companies should understand where they sit in the AI value chain.

A company using a third-party foundation model will normally be in a very different regulatory position from the original model provider. However, sufficiently significant modifications of a model may affect whether an organisation itself becomes a provider under the AI Act.

This distinction may become increasingly relevant as pharmaceutical companies move from simply using commercial AI services toward fine-tuning, adapting or integrating models into proprietary systems.

Why This Matters for GMP

The EU AI Act and GMP answer different regulatory questions.

The AI Act primarily addresses risks associated with placing AI systems and models on the European market and using them in the EU.

GMP, by contrast, asks whether a system used in pharmaceutical operations is suitable and controlled for its intended GxP use and whether product quality, patient safety and data integrity are protected.

A pharmaceutical company could therefore theoretically have an AI system that meets applicable AI Act requirements but is still unsuitable for a critical GMP process.

The reverse is also possible: a technically well-validated internal GxP application may still need assessment against applicable AI Act obligations.

This creates a new governance challenge.

Instead of asking only:

“Is this AI validated?”

pharmaceutical companies increasingly need to ask:

“Which regulatory frameworks apply to this particular AI use case, and what evidence is required under each?”

One AI Inventory – Several Regulatory Assessments

A practical consequence is that pharmaceutical companies should avoid maintaining completely separate inventories for AI Act compliance, IT governance and GxP validation.

A single corporate AI inventory can identify, for every AI use case:

  • intended use;
  • system owner;
  • provider and underlying model;
  • users and affected persons;
  • GxP relevance;
  • potential impact on product quality and patient safety;
  • company role under the AI Act;
  • applicable transparency requirements;
  • level of human oversight;
  • model version and configuration;
  • third-party dependencies;
  • required validation or qualification;
  • lifecycle monitoring and change-control requirements.

This provides a common starting point from which Legal, Quality, IT, Data Privacy, Cybersecurity and business functions can perform their respective assessments.

Supplier Governance May Become Even More Important

Most pharmaceutical companies will not build frontier AI models themselves. They will increasingly rely on external providers.

This makes supplier governance critical.

For GxP-relevant AI, pharmaceutical companies may need sufficient information to understand matters such as:

  • which model and version are being used;
  • when the model changes;
  • known limitations;
  • security controls;
  • data handling;
  • retention of prompts and outputs;
  • auditability;
  • availability of technical documentation;
  • incident notification;
  • mechanisms for monitoring performance.

The AI Act strengthens regulatory attention to transparency throughout the AI value chain. For pharmaceutical Quality organisations, this reinforces a familiar GMP principle:

An outsourced technology does not outsource the regulated company’s responsibility for its intended use.

Important: The AI Act Is Not Fully Applicable All at Once

The 2 August 2026 milestone should not be interpreted as meaning that every requirement of the AI Act suddenly became applicable to every AI system.

The Act follows a phased implementation timetable, and requirements depend on the type of system, the role of the organisation and the relevant provision.

Companies should therefore avoid simplistic statements such as:

“All AI systems must now comply with the full AI Act.”

Instead, each use case should be classified against the applicable provisions and implementation dates.

What Pharmaceutical Companies Should Do Now

For organisations already implementing AI, seven actions appear particularly useful:

  • Maintain a controlled inventory of AI use cases.
  • Determine the organisation’s regulatory role for each system.
  • Identify whether Article 50 transparency requirements apply.
  • Separately determine whether the use case is GxP-relevant.
  • Define supplier, model-version and change-control requirements.
  • Document human oversight and accountability.
  • Maintain evidence demonstrating both regulatory classification and ongoing control.

The central lesson is that AI governance in pharma can no longer be treated solely as an IT or innovation activity.

As AI regulation matures alongside emerging pharmaceutical-specific guidance, companies will need an integrated governance model connecting:

AI regulation + GMP + data integrity + cybersecurity + privacy + supplier management + quality risk management.

For pharmaceutical organisations operating in Europe, 2 August 2026 is therefore more than another AI Act implementation date.

It marks the transition from preparation toward active regulatory enforcement.

Sources

European Commission – Commission starts enforcing AI Act rules and new transparency requirements on 2 August:

https://digital-stra … equirements-2-august

European Commission – Guidelines on transparency obligations for providers and deployers of AI systems:

https://digital-stra … deployers-ai-systems

European Commission – Guidelines for providers of general-purpose AI models:

https://digital-stra … lines-gpai-providers

Thursday, August 20, 2026

FDA Proposes a New Way to Evaluate Generative AI: From Exhaustive Testing to Competency-Based Lifecycle Assurance


Introduction FDA proposes a competency-based approach for evaluating Generative AI-enabled medical devices, including risk-based benchmarking, clinical confirmation, lifecycle monitoring, model drift control and re-evaluation after AI changes.

The U.S. FDA has opened an important regulatory discussion on how Generative AI-enabled medical devices could be evaluated when traditional software testing is no longer sufficient.

On 18 August 2026, FDA’s Center for Devices and Radiological Health (CDRH) released the discussion paper Considerations for the Regulation of Generative AI-Enabled Medical Devices. The Agency is requesting stakeholder feedback until 19 October 2026.

The document applies specifically to medical devices, not to pharmaceutical GMP systems. Nevertheless, several concepts discussed by FDA are highly relevant to the broader debate about how probabilistic and generative AI can be validated and controlled in regulated environments.

The most significant idea may be surprisingly simple:

For sufficiently complex GenAI systems, it may be unrealistic to test every possible input and output. Instead, regulators may need to assess whether the system can demonstrate defined competencies and continue to demonstrate them throughout its lifecycle.

Why Traditional Software Testing May Not Be Enough

FDA recognizes that GenAI-enabled devices differ fundamentally from traditional software and even from many conventional AI systems.

They may:

  • accept open-ended inputs;
  • generate variable outputs for similar inputs;
  • perform multiple different subtasks;
  • use third-party foundation models;
  • change because of updates to models, prompts, retrieval mechanisms, guardrails or orchestration logic;
  • operate with increasing levels of autonomy.

FDA also explicitly identifies risks such as confabulations or hallucinations, limited transparency into third-party foundation models and performance degradation over time.

This creates a fundamental validation challenge.

For conventional software with bounded inputs and predetermined outputs, extensive input-output testing can provide strong evidence that the system operates as intended.

For GenAI, the possible combination of inputs, conversations and outputs can become effectively unlimited. FDA therefore acknowledges that evaluating every conceivable situation may simply not be practical.

Risk Depends on Both Autonomy and Consequences

FDA proposes a possible two-axis framework for thinking about GenAI risk.

One dimension considers what the AI actually does — ranging from providing non-directive information through directing an action to taking an action autonomously.

The second dimension considers the consequence of relying on an incorrect output.

Risk therefore increases as AI becomes more autonomous and as the potential consequences of an incorrect output become more serious.

This is an important distinction because the same underlying AI technology could represent very different levels of regulatory risk depending on its intended use.

An AI system that provides general information is fundamentally different from one that recommends a specific clinical action — and different again from an agentic system capable of executing that action.

FDA also emphasizes that simply adding wording such as “talk to your doctor” or “I am not a medical professional” may not necessarily make an otherwise action-directing AI function less directive.

A Competency-Based Approach to GenAI Evaluation

Perhaps the most innovative part of the discussion paper is FDA’s consideration of a competency-based approach inspired, at a high level, by the way human clinicians are evaluated.

Doctors are not qualified by testing every possible clinical situation they could ever encounter. Instead, they demonstrate competencies through examinations, supervised practice and continuing assessment.

FDA is exploring whether a related concept could be adapted for GenAI-enabled medical devices.

The proposed approach consists of two major components:

1. Device benchmarking

The deployed or representative final system would be tested against predefined competencies.

2. Clinical confirmation

Evidence would then confirm that the device performs appropriately under real or clinically representative conditions.

Importantly, FDA proposes evaluating the final user-facing device as configured for deployment, rather than assessing only the underlying foundation model such as an LLM.

This distinction is critical.

The performance of an AI application does not depend only on the model. It may also depend on the system prompt, retrieval architecture, knowledge sources, guardrails, user interface, orchestration logic and other controls surrounding the model.

What Would Be Tested?

FDA identifies several possible areas of competency for benchmarking GenAI systems.

They include:

Safety

  • recognition and escalation of safety-critical situations;
  • maintaining the defined scope and operational boundaries;
  • appropriate communication of uncertainty and deferral when the system cannot provide a reliable answer.

Clinical proficiency

  • knowledge and task fidelity;
  • information gathering and analysis;
  • quantitative reasoning;
  • quality and comprehensibility of communication.

Generalizability

  • robustness, reliability and reproducibility;
  • performance across relevant subgroups.

Agentic capabilities

  • additional competencies where AI can autonomously plan, use tools or execute multi-step actions.

FDA also discusses testing boundary adherence using techniques such as adversarial prompting, prompt injection and multi-turn conversations that gradually move outside the intended scope of the system.

This illustrates an important change in thinking about AI validation:

Validation is not only about whether AI produces correct answers. It is also about whether it behaves safely when it does not know the answer, when it is challenged, and when a user attempts to push it beyond its intended use.

Acceptance Criteria Still Matter

The competency-based approach does not mean abandoning predefined validation requirements.

Quite the opposite.

FDA considers that manufacturers should define the scope of testing based on intended use and risk, prespecify evaluation methods, justify scoring methods and establish acceptance criteria before testing.

Where human expert assessment is needed, FDA also discusses the use of appropriately qualified and structurally independent adjudicators.

This could become particularly important for GenAI because many outputs cannot simply be classified as mathematically “correct” or “incorrect.” Evaluation may require structured expert judgement supported by predefined scoring criteria.

Validation Does Not End at Deployment

FDA places substantial emphasis on postmarket performance monitoring.

Possible approaches include:

  • periodic re-benchmarking against predefined performance thresholds;
  • periodic review of real-world AI outputs by qualified independent clinicians;
  • monitoring for model or data drift and other forms of performance degradation.

Reassessment could occur periodically and following defined triggering events, including changes to the underlying model or other components of the AI architecture.

FDA even asks whether, under appropriate circumstances, greater uncertainty could be accepted before market authorization if it were compensated by stronger postmarket monitoring.

That is a potentially important regulatory concept.

It shifts the emphasis from demonstrating that an AI system was acceptable at one point in time toward demonstrating that it remains acceptable throughout operation.

What Happens When the Foundation Model Changes?

Another difficult issue addressed by FDA concerns applications built on third-party foundation models.

A medical-device manufacturer may build a validated application using an external foundation model, but the model provider may subsequently change that model.

The device manufacturer may therefore face a GMP-like change-control problem even though the change originated outside its own organization.

FDA discusses mechanisms ranging from documentation within the manufacturer’s Quality Management System to regulatory authorization and the use of Predetermined Change Control Plans (PCCPs). Re-benchmarking against the original competency criteria could provide evidence that the modified system continues to perform acceptably.

FDA is also considering the concept of voluntary Foundation Model Device Master Files, through which foundation-model developers could provide FDA with information about model architecture, training-data provenance, known limitations, failure modes, performance, guardrails, model updates and audit-log availability.

Agentic AI Raises the Risk Further

FDA specifically addresses agentic AI — systems capable of autonomously planning and executing multi-step tasks, using external tools and taking actions.

The Agency asks how increased autonomy and the reduced opportunity for human review should influence acceptance criteria and regulatory oversight.

This could become increasingly important as AI moves from:

providing information → recommending decisions → executing decisions.

The regulatory significance of AI therefore may increasingly depend not simply on the sophistication of the model, but on how much authority the system is given to act.

Why This Matters Beyond Medical Devices

The FDA discussion paper is not GMP guidance and does not establish requirements for AI used in pharmaceutical manufacturing or Pharmaceutical Quality Systems. FDA explicitly states that the document is for discussion only and does not represent draft or final guidance or proposed regulatory expectations.

Nevertheless, the regulatory thinking behind the paper deserves attention from pharmaceutical companies.

Several concepts could be highly relevant to future approaches for GenAI used in GxP environments:

Intended use → Risk classification → Defined competencies → Predefined acceptance criteria → Qualification/benchmarking → Human confirmation → Continuous monitoring → Re-benchmarking after change

This may ultimately prove more suitable for probabilistic AI than attempting to force GenAI into a traditional deterministic software-validation model.

The particularly important message is that probabilistic output does not necessarily mean that an AI system cannot be controlled.

Instead, control may need to be demonstrated differently: through risk-based performance requirements, validated boundaries and guardrails, representative challenge testing, human oversight, predefined acceptance criteria and continuous lifecycle monitoring.

In this sense, the FDA paper may represent another important step toward a regulatory model based not on demanding that AI behave like deterministic software, but on demonstrating that its performance remains acceptably controlled for its intended use.

Source

U.S. Food and Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback, 18 August 2026.
https://www.fda.gov/ … on-paper-and-request