The U.S. FDA has opened an important regulatory discussion on how Generative AI-enabled medical devices could be evaluated when traditional software testing is no longer sufficient.
On 18 August 2026, FDA’s Center for Devices and Radiological Health (CDRH) released the discussion paper Considerations for the Regulation of Generative AI-Enabled Medical Devices. The Agency is requesting stakeholder feedback until 19 October 2026.
The document applies specifically to medical devices, not to pharmaceutical GMP systems. Nevertheless, several concepts discussed by FDA are highly relevant to the broader debate about how probabilistic and generative AI can be validated and controlled in regulated environments.
The most significant idea may be surprisingly simple:
For sufficiently complex GenAI systems, it may be unrealistic to test every possible input and output. Instead, regulators may need to assess whether the system can demonstrate defined competencies and continue to demonstrate them throughout its lifecycle.
Why Traditional Software Testing May Not Be Enough
FDA recognizes that GenAI-enabled devices differ fundamentally from traditional software and even from many conventional AI systems.
They may:
- accept open-ended inputs;
- generate variable outputs for similar inputs;
- perform multiple different subtasks;
- use third-party foundation models;
- change because of updates to models, prompts, retrieval mechanisms, guardrails or orchestration logic;
- operate with increasing levels of autonomy.
FDA also explicitly identifies risks such as confabulations or hallucinations, limited transparency into third-party foundation models and performance degradation over time.
This creates a fundamental validation challenge.
For conventional software with bounded inputs and predetermined outputs, extensive input-output testing can provide strong evidence that the system operates as intended.
For GenAI, the possible combination of inputs, conversations and outputs can become effectively unlimited. FDA therefore acknowledges that evaluating every conceivable situation may simply not be practical.
Risk Depends on Both Autonomy and Consequences
FDA proposes a possible two-axis framework for thinking about GenAI risk.
One dimension considers what the AI actually does — ranging from providing non-directive information through directing an action to taking an action autonomously.
The second dimension considers the consequence of relying on an incorrect output.
Risk therefore increases as AI becomes more autonomous and as the potential consequences of an incorrect output become more serious.
This is an important distinction because the same underlying AI technology could represent very different levels of regulatory risk depending on its intended use.
An AI system that provides general information is fundamentally different from one that recommends a specific clinical action — and different again from an agentic system capable of executing that action.
FDA also emphasizes that simply adding wording such as “talk to your doctor” or “I am not a medical professional” may not necessarily make an otherwise action-directing AI function less directive.
A Competency-Based Approach to GenAI Evaluation
Perhaps the most innovative part of the discussion paper is FDA’s consideration of a competency-based approach inspired, at a high level, by the way human clinicians are evaluated.
Doctors are not qualified by testing every possible clinical situation they could ever encounter. Instead, they demonstrate competencies through examinations, supervised practice and continuing assessment.
FDA is exploring whether a related concept could be adapted for GenAI-enabled medical devices.
The proposed approach consists of two major components:
1. Device benchmarking
The deployed or representative final system would be tested against predefined competencies.
2. Clinical confirmation
Evidence would then confirm that the device performs appropriately under real or clinically representative conditions.
Importantly, FDA proposes evaluating the final user-facing device as configured for deployment, rather than assessing only the underlying foundation model such as an LLM.
This distinction is critical.
The performance of an AI application does not depend only on the model. It may also depend on the system prompt, retrieval architecture, knowledge sources, guardrails, user interface, orchestration logic and other controls surrounding the model.
What Would Be Tested?
FDA identifies several possible areas of competency for benchmarking GenAI systems.
They include:
Safety
- recognition and escalation of safety-critical situations;
- maintaining the defined scope and operational boundaries;
- appropriate communication of uncertainty and deferral when the system cannot provide a reliable answer.
Clinical proficiency
- knowledge and task fidelity;
- information gathering and analysis;
- quantitative reasoning;
- quality and comprehensibility of communication.
Generalizability
- robustness, reliability and reproducibility;
- performance across relevant subgroups.
Agentic capabilities
- additional competencies where AI can autonomously plan, use tools or execute multi-step actions.
FDA also discusses testing boundary adherence using techniques such as adversarial prompting, prompt injection and multi-turn conversations that gradually move outside the intended scope of the system.
This illustrates an important change in thinking about AI validation:
Validation is not only about whether AI produces correct answers. It is also about whether it behaves safely when it does not know the answer, when it is challenged, and when a user attempts to push it beyond its intended use.
Acceptance Criteria Still Matter
The competency-based approach does not mean abandoning predefined validation requirements.
Quite the opposite.
FDA considers that manufacturers should define the scope of testing based on intended use and risk, prespecify evaluation methods, justify scoring methods and establish acceptance criteria before testing.
Where human expert assessment is needed, FDA also discusses the use of appropriately qualified and structurally independent adjudicators.
This could become particularly important for GenAI because many outputs cannot simply be classified as mathematically “correct” or “incorrect.” Evaluation may require structured expert judgement supported by predefined scoring criteria.
Validation Does Not End at Deployment
FDA places substantial emphasis on postmarket performance monitoring.
Possible approaches include:
- periodic re-benchmarking against predefined performance thresholds;
- periodic review of real-world AI outputs by qualified independent clinicians;
- monitoring for model or data drift and other forms of performance degradation.
Reassessment could occur periodically and following defined triggering events, including changes to the underlying model or other components of the AI architecture.
FDA even asks whether, under appropriate circumstances, greater uncertainty could be accepted before market authorization if it were compensated by stronger postmarket monitoring.
That is a potentially important regulatory concept.
It shifts the emphasis from demonstrating that an AI system was acceptable at one point in time toward demonstrating that it remains acceptable throughout operation.
What Happens When the Foundation Model Changes?
Another difficult issue addressed by FDA concerns applications built on third-party foundation models.
A medical-device manufacturer may build a validated application using an external foundation model, but the model provider may subsequently change that model.
The device manufacturer may therefore face a GMP-like change-control problem even though the change originated outside its own organization.
FDA discusses mechanisms ranging from documentation within the manufacturer’s Quality Management System to regulatory authorization and the use of Predetermined Change Control Plans (PCCPs). Re-benchmarking against the original competency criteria could provide evidence that the modified system continues to perform acceptably.
FDA is also considering the concept of voluntary Foundation Model Device Master Files, through which foundation-model developers could provide FDA with information about model architecture, training-data provenance, known limitations, failure modes, performance, guardrails, model updates and audit-log availability.
Agentic AI Raises the Risk Further
FDA specifically addresses agentic AI — systems capable of autonomously planning and executing multi-step tasks, using external tools and taking actions.
The Agency asks how increased autonomy and the reduced opportunity for human review should influence acceptance criteria and regulatory oversight.
This could become increasingly important as AI moves from:
providing information → recommending decisions → executing decisions.
The regulatory significance of AI therefore may increasingly depend not simply on the sophistication of the model, but on how much authority the system is given to act.
Why This Matters Beyond Medical Devices
The FDA discussion paper is not GMP guidance and does not establish requirements for AI used in pharmaceutical manufacturing or Pharmaceutical Quality Systems. FDA explicitly states that the document is for discussion only and does not represent draft or final guidance or proposed regulatory expectations.
Nevertheless, the regulatory thinking behind the paper deserves attention from pharmaceutical companies.
Several concepts could be highly relevant to future approaches for GenAI used in GxP environments:
Intended use → Risk classification → Defined competencies → Predefined acceptance criteria → Qualification/benchmarking → Human confirmation → Continuous monitoring → Re-benchmarking after change
This may ultimately prove more suitable for probabilistic AI than attempting to force GenAI into a traditional deterministic software-validation model.
The particularly important message is that probabilistic output does not necessarily mean that an AI system cannot be controlled.
Instead, control may need to be demonstrated differently: through risk-based performance requirements, validated boundaries and guardrails, representative challenge testing, human oversight, predefined acceptance criteria and continuous lifecycle monitoring.
In this sense, the FDA paper may represent another important step toward a regulatory model based not on demanding that AI behave like deterministic software, but on demonstrating that its performance remains acceptably controlled for its intended use.
Source
U.S. Food and Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback, 18 August 2026.
https://www.fda.gov/ … on-paper-and-request
