ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
Pathology image analysis is a critical application for multimodal large language models (MLLMs), yet existing benchmarks often focus on final diagnostic answers or captions, which do not reveal whether a model truly understands the multiscale visual content necessary for pathology reasoning. PathVU addresses this gap by introducing a vision-anchored benchmark that explicitly evaluates fine-grained and multiscale visual understanding. This is significant because pathology images contain structures at multiple scales—from cellular details to tissue architecture—and models must integrate information across these scales to make accurate decisions.
The paper's emphasis on programmatic scoring and deterministic task targets is a major step toward reproducible evaluation. By converting raw annotations into well-defined tasks, PathVU avoids the subjectivity of free-form answers and enables objective comparison of model capabilities. This is particularly important as MLLMs become more prevalent in medical settings, where reliability and interpretability are paramount.
The paper reports that even advanced MLLMs exhibit substantial limitations on fine-grained visual tasks across multiscale pathology images. While specific numerical results are not detailed in the abstract, the overall finding is that current models struggle with tasks requiring precise localization, quantity estimation, and spatial reasoning. This suggests that existing MLLMs are not yet reliable for pathology applications that demand fine-grained visual understanding.
The benchmark's design allows for detailed comparisons across model categories, likely revealing that pathology-oriented models perform better on some tasks but still fall short of human-level performance. The results underscore the need for further research in multiscale visual reasoning and domain-specific training.
PathVU has the potential to become a standard benchmark for evaluating pathology MLLMs, similar to how ImageNet drove progress in computer vision. By focusing on fine-grained visual understanding rather than just final answers, it encourages the development of models that can reason about pathology images in a clinically meaningful way. This could lead to more trustworthy AI systems in healthcare, where understanding the 'why' behind a diagnosis is as important as the diagnosis itself.
Moreover, the benchmark's reproducibility and comprehensive coverage make it a valuable resource for the research community. It provides a clear roadmap for improving MLLMs in pathology, highlighting specific areas where models fail and guiding future innovations. As MLLMs continue to evolve, benchmarks like PathVU will be essential to ensure that progress is aligned with real-world needs.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba