A new perspective paper published on 18 August 2026 in Frontiers in Drug Safety and Regulation provides an important new way of thinking about AI automation in regulated pharmaceutical processes.
The authors are scientists from the U.S. Food and Drug Administration, including the National Center for Toxicological Research and the Center for Drug Evaluation and Research.
The paper focuses specifically on pharmacovigilance, but the underlying concept may have much broader relevance for pharmaceutical Quality Systems and GMP applications.
The central question is not simply:
“Can AI perform this task?”
Instead, the authors propose asking:
“Can this task be translated into sufficiently explicit, auditable and verifiable computational steps?”
They call this concept “computable pharmacovigilance.”
Not Every Pharmaceutical Task Is Equally Suitable for AI Automation
The framework distinguishes between tasks that can be clearly formalised and tasks that depend heavily on professional judgment.
Examples of relatively computable activities include:
- checking whether required information is missing;
- detecting potentially duplicated safety reports;
- extracting predefined information from records;
- coding structured information;
- retrieving relevant information;
- performing defined rule-based checks.
Such activities have relatively clear:
inputs → processing rules → outputs → acceptance criteria
and can therefore be tested, benchmarked and audited.
Other activities are fundamentally more difficult.
The authors use case-level causality assessment as an important example.
Determining whether a medicine actually caused a particular adverse event can require integration of:
- temporality;
- alternative explanations;
- clinical context;
- pharmacological mechanisms;
- rechallenge information;
- professional medical judgment.
These activities cannot yet be reduced reliably to fully explicit computational rules.
The consequence is important:
AI suitability should be assessed task by task rather than system by system.
“Computability” Could Be a Useful Concept Beyond Pharmacovigilance
Although the paper concerns pharmacovigilance, the same reasoning could potentially be useful when evaluating AI use in GMP.
For example, some Quality System activities may be highly computable:
- checking whether required fields are completed;
- comparing a document against predefined requirements;
- identifying duplicate or recurring deviations;
- extracting equipment or batch information;
- checking dates, limits or predefined conditions;
- classifying documents against controlled taxonomies.
Other activities may remain much less computable:
- determining the true root cause of a complex deviation;
- assessing whether an unexplained event affects product quality;
- evaluating the adequacy of a CAPA strategy;
- making a final batch disposition decision;
- determining whether an unusual observation represents an unacceptable patient risk.
This suggests a potentially useful AI risk-assessment question for pharma:
How much of the intended task can actually be expressed as explicit, testable and auditable logic?
The less computable the task, the greater the likely need for expert oversight.
Large Language Models Are Not Automatically the Best Solution
Another important message from the paper concerns the current tendency to use large language models for almost every AI application.
The authors distinguish between:
- large models – general-purpose LLMs and foundation models;
- small models – task-specific machine-learning models;
- deterministic tools – rules, dictionaries, workflow logic and conventional software.
Large models offer flexibility and can work across multiple text-heavy tasks.
But they also introduce important disadvantages:
- hallucinations;
- variable outputs;
- prompt sensitivity;
- possible information loss;
- more difficult validation;
- more difficult auditability.
Smaller task-specific models can be easier to validate when the task has constrained inputs and clearly defined outputs.
The paper therefore argues that the future may not belong to one universal AI model.
Instead, the most practical architecture may be hybrid.
A Hybrid Architecture: LLM + Small Models + Deterministic Tools
One particularly interesting proposal is an architecture in which an LLM acts as an orchestrator but does not perform every regulated function itself.
The architecture could look like:
User / Reviewer
↓
LLM or AI Agent
↓
Controlled Tools / Retrieval / Rules / Small Models
↓
Structured Result
↓
Expert Review and Decision
The LLM can interpret the request and coordinate different tools.
But specific operations can be delegated to more deterministic components such as:
- validated rule engines;
- controlled terminology;
- structured retrieval;
- database queries;
- schema validation;
- task-specific ML models.
This architecture can reduce uncontrolled generative behaviour.
It can also improve auditability because intermediate actions and outputs can be recorded.
For regulated pharmaceutical applications, this may be a much more attractive architecture than asking one general-purpose LLM to perform the entire process.
An Interesting Finding: Bigger AI Was Not Always Better
The authors also report an illustrative comparison between a task-specific BioBERT model and a general-purpose Llama-4 model for extraction of safety-related information.
For the evaluated structured extraction tasks, the task-specific model performed better.
The authors do not present this as a universal comparison between the two technologies.
Instead, the example demonstrates a broader principle:
The most powerful general-purpose model is not necessarily the best model for a tightly defined regulated task.
For pharmaceutical companies, this is significant.
The correct question may therefore not be:
“Which is the best AI model?”
but:
“Which technical architecture provides the most reliable, verifiable and auditable performance for this particular task?”
“Trust but Verify”
The paper describes the boundary of AI automation using a “trust but verify” philosophy.
AI can increasingly perform routine, well-specified operations.
However, as the process moves from:
pattern recognition → interpretation → judgment → regulated decision
the requirement for structured verification and expert review becomes increasingly important.
The authors conclude that Generative AI does not currently mean the end of human pharmacovigilance.
Instead, the likely future is a:
Human–AI system
in which AI performs increasingly sophisticated operations while humans remain responsible for judgment-intensive decisions.
Why This Matters for GMP and Pharmaceutical Quality
This concept could be particularly useful when companies perform AI use-case assessments.
Instead of simply classifying an application as:
AI / non-AI
or:
GxP / non-GxP
an additional question could be introduced:
How computable is the task?
A possible assessment sequence could therefore become:
Intended Use → GxP Impact → Task Computability → AI Architecture → Verification Strategy → Human Oversight → Lifecycle Monitoring
This could help organisations determine whether a task should be performed using:
- deterministic software;
- a task-specific ML model;
- an LLM;
- a hybrid architecture;
- or primarily human expert judgment.
This may ultimately be more useful than assuming that Generative AI should replace existing software simply because it is more flexible.
Regulatory Status Note
The article was written by FDA scientists, but it is a scientific perspective paper and does not represent formal FDA guidance or FDA policy.
The authors explicitly state that the views expressed are their own and do not necessarily represent the official position of the U.S. Food and Drug Administration.
Nevertheless, the paper provides valuable insight into how regulatory scientists are thinking about AI automation, validation, auditability and human judgment in regulated pharmaceutical activities.
Source
Wu L., Xu J., Dang O., Ball R. – Does generative AI mean the “end of history” for pharmacovigilance automation? Towards a framework for the future of human-AI systems. Frontiers in Drug Safety and Regulation, 18 August 2026: