MMRDoc: A multi-granularity multimodal RAG framework for long-document VQA
Shengxu Xu, Hongsong Wang
MMRDoc introduces a multi-granularity multimodal RAG framework that improves long-document VQA by enhancing semantic representation and retrieval.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Shengxu Xu, Hongsong Wang
MMRDoc introduces a multi-granularity multimodal RAG framework that improves long-document VQA by enhancing semantic representation and retrieval.
Unknown
SpatialVLM endows vision-language models with spatial reasoning capabilities by generating and training on a large-scale spatial VQA dataset.
Zongyi Chen, Yu Liang, Jie Lin, et al.
PathVU is a vision-anchored benchmark for fine-grained multiscale visual understanding in pathology, with 14 VQA tasks, 61,673 images, and 308,070 samples, revealing substantial limitations in current MLLMs.
Unknown
LLaVA 1.5 enhances multimodal AI by integrating CLIP-ViT-L-336px with MLP projection and academic VQA data, achieving state-of-the-art results.
Iryna Hartsock, Ghulam Rasool
This review surveys recent medical vision-language models, covering 18 datasets and 16 models for report generation and VQA, highlighting challenges and future directions.