ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
14
Citations
0
Influential Citations
arXiv.org
Venue
2025
Year
The proliferation of Large Language Models (LLMs) in medicine has enabled impressive capabilities, yet a critical gap remains in their ability to perform systematic, transparent, and verifiable reasoning, a cornerstone of clinical practice. This has catalyzed a shift from single-step answer generation to the development of LLMs explicitly designed for medical reasoning. This paper provides the first systematic review of this emerging field. We propose a taxonomy of reasoning enhancement techniques, categorized into training-time strategies (e.g., supervised fine-tuning, reinforcement learning) and test-time mechanisms (e.g., prompt engineering, multi-agent systems). We analyze how these techniques are applied across different data modalities (text, image, code) and in key clinical applications such as diagnosis, education, and treatment planning. Furthermore, we survey the evolution of evaluation benchmarks from simple accuracy metrics to sophisticated assessments of reasoning quality and visual interpretability. Based on an analysis of 60 seminal studies from 2022-2025, we conclude by identifying critical challenges, including the faithfulness-plausibility gap and the need for native multimodal reasoning, and outlining future directions toward building efficient, robust, and sociotechnically responsible medical AI.
This paper addresses a critical gap in the application of Large Language Models (LLMs) to medicine: the lack of systematic, transparent, and verifiable reasoning. While LLMs have demonstrated impressive capabilities in medical tasks, their clinical utility is fundamentally limited by opaque, single-step answer generation that cannot be trusted or audited. The authors provide the first systematic review of the emerging field of medical reasoning in LLMs, offering a much-needed taxonomy that organizes the rapidly growing body of work. This is particularly timely as the AI community moves from benchmark chasing to building models that can reason like clinicians.
The paper's significance lies in its comprehensive framing of the problem space. By categorizing techniques into training-time (supervised fine-tuning, reinforcement learning) and test-time (prompt engineering, multi-agent systems) strategies, it gives practitioners a clear map of available tools. The analysis across data modalities (text, image, code) and clinical applications (diagnosis, education, treatment planning) further grounds the review in real-world use cases. The identification of the faithfulness-plausibility gap—where models produce plausible but unfaithful reasoning—is a crucial insight that will shape future research priorities.
The review synthesizes findings from 60 studies, documenting a clear shift from single-step answer generation to multi-step reasoning in medical LLMs. Key results include the identification of training-time techniques (e.g., SFT on clinical reasoning chains, RL from feedback) as more effective for deep reasoning, while test-time methods (e.g., chain-of-thought prompting, multi-agent debate) improve transparency. The authors note that current benchmarks are evolving to capture reasoning quality, but most models still exhibit a faithfulness-plausibility gap where reasoning chains appear logical but are not grounded in clinical evidence. No single technique has emerged as dominant; rather, combinations of training and test-time strategies show the most promise.
This review provides a foundational framework for the rapidly growing field of medical reasoning in LLMs. For AI practitioners, it offers a clear taxonomy to guide model development and evaluation, highlighting which techniques are most appropriate for different clinical scenarios. The identification of the faithfulness-plausibility gap is particularly important, as it sets a concrete research agenda for building trustworthy medical AI. By emphasizing the need for native multimodal reasoning and sociotechnical responsibility, the paper pushes the field beyond narrow accuracy metrics toward clinically meaningful capabilities. This work will likely influence both academic research and industry product development in medical AI, serving as a reference point for future studies on reasoning in LLMs.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba