ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
114
Citations
2
Influential Citations
International Journal of Medical Sciences
Venue
2025
Year
In recent years, large language models (LLMs) represented by GPT-4 have developed rapidly and performed well in various natural language processing tasks, showing great potential and transformative impact. The medical field, due to its vast data information as well as complex diagnostic and treatment processes, is undoubtedly one of the most promising areas for the application of LLMs. At present, LLMs has been gradually implemented in clinical practice, medical research, and medical education. However, in practical applications, medical LLMs still face numerous challenges, including the phenomenon of hallucination, interpretability, and ethical concerns. Therefore, in-depth exploration is still needed in areas of standardized evaluation frameworks, multimodal LLMs, and multidisciplinary collaboration in the future, so as to realize the widespread application of medical LLMs and promote the development and transformation in the field of global healthcare. This review offers a comprehensive overview of applications, challenges, and future directions of LLMs in medicine, providing new insights for the sustained development of medical LLMs.
This review arrives at a pivotal moment when LLMs like GPT-4 are being rapidly adopted across industries, yet their integration into high-stakes medical settings remains fraught with risk. The paper systematically maps the landscape of medical LLM applications—from clinical decision support to medical education—while honestly confronting the core obstacles that prevent widespread deployment. For AI practitioners, this is essential reading because it moves beyond hype to identify concrete failure modes (hallucination, lack of interpretability) that must be addressed before LLMs can be trusted in patient care. The emphasis on standardized evaluation frameworks is particularly timely, as the field currently lacks rigorous benchmarks for medical LLM performance.
As a review, the paper does not report new experimental results. It cites 114 references to support its claims, noting that current medical LLMs achieve high performance on NLP tasks (e.g., question answering, summarization) but still exhibit hallucination rates that are unacceptable in clinical settings. The paper does not provide specific accuracy or error metrics from individual studies.
For the AI community, this paper serves as a call to action. It underscores that while LLMs have demonstrated remarkable capabilities in general domains, their application in medicine requires domain-specific rigor. The identified challenges—especially hallucination and interpretability—are active research areas that demand novel solutions. The push for multimodal LLMs aligns with the broader trend toward foundation models that can process diverse medical data types. By framing the path forward in terms of evaluation, multimodality, and collaboration, the paper provides a strategic blueprint that could accelerate safe and effective deployment of LLMs in healthcare, ultimately impacting patient outcomes and clinical workflows.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba