Preprint
Large Language Models

Evaluation of large language models’ ability to identify clinically relevant drug-drug interactions and generate high-quality clinical pharmacotherapy recommendations

Aaron Chase(Department of Pharmacy, Augusta University Medical Center/UGA College of Pharmacy , Augusta, GA ,), Amoreena Most(Department of Pharmacy, Augusta University Medical Center/UGA College of Pharmacy , Augusta, GA ,), Andrea Sikora(Department of Clinical and Administrative Pharmacy, University of Georgia College of Pharmacy, Augusta, GA, and University of Colorado School of Medicine , Aurora, CO ,), Susan E Smith(Department of Clinical and Administrative Pharmacy, University of Georgia College of Pharmacy , Athens, GA ,), John W Devlin(Northeastern University School of Pharmacy, Boston, MA, and Division of Pulmonary and Critical Care Medicine, Brigham and Women’s Hospital , Boston, MA ,), Shaochen Xu(University of Georgia School of Computing , Athens, GA ,), Tianming Liu(University of Georgia School of Computing , Athens, GA ,), Brian Murray(University of Colorado Skaggs School of Pharmacy , Aurora, CO ,)
July 1, 2025American Journal of Health-System Pharmacy4 citations

4

Citations

0

Influential Citations

American Journal of Health-System Pharmacy

Venue

2025

Year

Abstract

Abstract Purpose Large language models (LLMs) are promising artificial intelligence (AI) tools to support clinical decision-making. The ability of LLMs to evaluate medication regimens, identify drug-drug interactions (DDIs), and provide clinical recommendations has undergone limited evaluation. The purpose of this study was to compare the performance of 3 LLMs in recognizing DDIs, determining clinical relevance, and generating management recommendations. Methods A total of 15 patient cases with medication regimens were created; each contained a commonly encountered DDI. Two separate study phases were developed: (1) DDI identification and determination of clinical relevance; and (2) DDI identification and generation of a clinical recommendation. The primary outcome was the ability of the LLMs (GPT-4, Gemini 1.5, and Claude 3) to identify the DDI within each medication regimen. Secondary outcomes included the ability of the LLMs to identify the clinical relevance of each DDI and generate a recommendation of high quality relative to ground truth. Results Claude 3 identified all DDIs, followed by GPT-4 (14/15, 93.3%) and Gemini 1.5 (12/15, 80.0%). All LLMs were significantly more likely than clinical experts to categorize the DDI as clinically relevant (P < 0.01). DDI management recommendations provided by GPT-4 were rated as optimal in 8 of 13 (61.5%) of the cases (P = 0.05 for comparison to ground truth). Two recommendations from GPT-4 and one recommendation from Gemini 1.5 were deemed to result in potential patient harm. Conclusion While LLMs demonstrate promising potential to identify DDIs, application to clinical cases requires ongoing development. Findings from this study may assist in future development and refinement of LLMs for clinical decision-making related to DDIs.

Analysis

Why This Paper Matters

This paper addresses a critical gap in the application of large language models (LLMs) to clinical pharmacology: the ability to identify drug-drug interactions (DDIs) and generate actionable recommendations. While LLMs have shown promise in medical question answering, their performance on complex, multi-step clinical reasoning tasks like medication management is underexplored. The study directly compares three leading LLMs (GPT-4, Gemini 1.5, Claude 3) on a realistic task, providing valuable insights for AI practitioners and healthcare IT developers.

The findings are timely given the rapid adoption of LLMs in healthcare settings. The paper's rigorous methodology—using 15 constructed cases with ground truth—offers a reproducible framework for evaluating LLMs in clinical decision support. The result that all LLMs overestimate clinical relevance of DDIs compared to experts is a crucial cautionary note, as it could lead to alert fatigue or unnecessary treatment changes if deployed without oversight.

Technical Contributions

  • Comparative evaluation: First head-to-head comparison of GPT-4, Gemini 1.5, and Claude 3 on DDI identification and recommendation generation.
  • Two-phase study design: Separates DDI identification/relevance assessment from recommendation generation, allowing granular analysis of LLM capabilities.
  • Ground truth and expert comparison: Uses clinical experts as a benchmark, revealing that LLMs differ significantly from human judgment in relevance classification.
  • Safety assessment: Evaluates recommendations for potential patient harm, a critical metric often missing in LLM evaluations.
  • Reproducible case set: Provides 15 patient cases with common DDIs, enabling future benchmarking.

Results

Claude 3 achieved perfect DDI identification (15/15), followed by GPT-4 (14/15, 93.3%) and Gemini 1.5 (12/15, 80%). However, all LLMs were significantly more likely than experts to label DDIs as clinically relevant (P<0.01), suggesting a tendency toward over-caution. For management recommendations, GPT-4 produced optimal recommendations in 8/13 cases (61.5%), with a borderline significant difference from ground truth (P=0.05). Critically, two GPT-4 and one Gemini 1.5 recommendations were deemed potentially harmful, underscoring the risk of relying on LLMs without human oversight.

Significance

The study provides a balanced view of LLM potential in clinical pharmacology. While identification accuracy is high, the overestimation of clinical relevance and the generation of harmful recommendations in a small subset of cases highlight the need for robust validation and safety mechanisms. For AI practitioners, this work emphasizes the importance of domain-specific evaluation metrics (e.g., harm potential) beyond simple accuracy. It also suggests that LLMs could be used as assistive tools rather than autonomous decision-makers, with human experts making final judgments. Future work should expand the case set, include more diverse DDIs, and explore fine-tuning or retrieval-augmented generation to improve recommendation quality and safety.