ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
4
Citations
0
Influential Citations
American Journal of Health-System Pharmacy
Venue
2025
Year
Abstract Purpose Large language models (LLMs) are promising artificial intelligence (AI) tools to support clinical decision-making. The ability of LLMs to evaluate medication regimens, identify drug-drug interactions (DDIs), and provide clinical recommendations has undergone limited evaluation. The purpose of this study was to compare the performance of 3 LLMs in recognizing DDIs, determining clinical relevance, and generating management recommendations. Methods A total of 15 patient cases with medication regimens were created; each contained a commonly encountered DDI. Two separate study phases were developed: (1) DDI identification and determination of clinical relevance; and (2) DDI identification and generation of a clinical recommendation. The primary outcome was the ability of the LLMs (GPT-4, Gemini 1.5, and Claude 3) to identify the DDI within each medication regimen. Secondary outcomes included the ability of the LLMs to identify the clinical relevance of each DDI and generate a recommendation of high quality relative to ground truth. Results Claude 3 identified all DDIs, followed by GPT-4 (14/15, 93.3%) and Gemini 1.5 (12/15, 80.0%). All LLMs were significantly more likely than clinical experts to categorize the DDI as clinically relevant (P < 0.01). DDI management recommendations provided by GPT-4 were rated as optimal in 8 of 13 (61.5%) of the cases (P = 0.05 for comparison to ground truth). Two recommendations from GPT-4 and one recommendation from Gemini 1.5 were deemed to result in potential patient harm. Conclusion While LLMs demonstrate promising potential to identify DDIs, application to clinical cases requires ongoing development. Findings from this study may assist in future development and refinement of LLMs for clinical decision-making related to DDIs.
This paper addresses a critical gap in the application of large language models (LLMs) to clinical pharmacology: the ability to identify drug-drug interactions (DDIs) and generate actionable recommendations. While LLMs have shown promise in medical question answering, their performance on complex, multi-step clinical reasoning tasks like medication management is underexplored. The study directly compares three leading LLMs (GPT-4, Gemini 1.5, Claude 3) on a realistic task, providing valuable insights for AI practitioners and healthcare IT developers.
The findings are timely given the rapid adoption of LLMs in healthcare settings. The paper's rigorous methodology—using 15 constructed cases with ground truth—offers a reproducible framework for evaluating LLMs in clinical decision support. The result that all LLMs overestimate clinical relevance of DDIs compared to experts is a crucial cautionary note, as it could lead to alert fatigue or unnecessary treatment changes if deployed without oversight.
Claude 3 achieved perfect DDI identification (15/15), followed by GPT-4 (14/15, 93.3%) and Gemini 1.5 (12/15, 80%). However, all LLMs were significantly more likely than experts to label DDIs as clinically relevant (P<0.01), suggesting a tendency toward over-caution. For management recommendations, GPT-4 produced optimal recommendations in 8/13 cases (61.5%), with a borderline significant difference from ground truth (P=0.05). Critically, two GPT-4 and one Gemini 1.5 recommendations were deemed potentially harmful, underscoring the risk of relying on LLMs without human oversight.
The study provides a balanced view of LLM potential in clinical pharmacology. While identification accuracy is high, the overestimation of clinical relevance and the generation of harmful recommendations in a small subset of cases highlight the need for robust validation and safety mechanisms. For AI practitioners, this work emphasizes the importance of domain-specific evaluation metrics (e.g., harm potential) beyond simple accuracy. It also suggests that LLMs could be used as assistive tools rather than autonomous decision-makers, with human experts making final judgments. Future work should expand the case set, include more diverse DDIs, and explore fine-tuning or retrieval-augmented generation to improve recommendation quality and safety.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba