Medical Reasoning with Large Language Models: A Systematic Review and Evaluation (April 2026)
FreeComprehensive review of medical reasoning methods + MR-Bench (real-world hospital data); reveals large gap between exam-level performance and authentic clinical decision-making
About Medical Reasoning with Large Language Models: A Systematic Review and Evaluation (April 2026)
This paper presents a comprehensive survey of medical reasoning with large language models (LLMs), grounded in cognitive theories of clinical reasoning and organizing existing methods into seven major technical routes spanning training-based and training-free approaches. The authors introduce MR-Bench, a benchmark derived from real-world hospital data designed to assess clinically grounded reasoning. Evaluations on MR-Bench reveal a pronounced gap between LLMs' strong performance on medical exam-style tasks and their accuracy on authentic clinical decision-making tasks, highlighting the need for robust reasoning beyond factual recall.
Key Features
Pros & Cons
- Provides a systematic and unified overview of existing medical reasoning methods
- Introduces a benchmark built from actual clinical data, not synthetic exams
- Reveals critical gap that is important for safe deployment in healthcare
- Grounds analysis in established cognitive theories of clinical reasoning
- Open access and freely available on arXiv
- Paper is a survey and benchmark, not a deployable clinical tool
- Benchmark may not cover all medical specialties or settings
- Findings may become outdated as models evolve rapidly