A comprehensive survey on multimodal rag: All combinations of modalities as input and output
Unknown
A comprehensive survey of multimodal RAG covering all combinations of input and output modalities, providing a taxonomy and analysis of recent work.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
A comprehensive survey of multimodal RAG covering all combinations of input and output modalities, providing a taxonomy and analysis of recent work.
Priyanka Kargupta, S. Li, Haocheng Wang, et al.
This paper synthesizes cognitive science into a taxonomy of 28 cognitive elements, evaluates 192K LLM traces across modalities, and develops test-time guidance improving performance by up to 66.7%.
Woongyeong Yeo, Kangsan Kim, Soyeong Jeong, et al.
UniversalRAG is an any-to-any RAG framework that retrieves and integrates knowledge from heterogeneous sources across modalities and granularities using modality-aware routing to overcome the modality gap.
Wenxuan Wang, Zizhan Ma, Meidan Ding, et al.
First systematic review of LLM reasoning in medicine, proposing a taxonomy of training-time and test-time enhancement techniques across modalities and clinical applications.
Unknown
CogVLM2 introduces a family of visual language models for image and video understanding with improved training recipes, higher resolution, and broader modalities.
Unknown
This paper revisits neural scaling laws for language and vision models, finding that scaling trends hold across modalities but with different exponents and saturation points.
Unknown
FuseMoE introduces a mixture-of-experts framework with a novel gating function for integrating diverse numbers of modalities.
Akash Ghosh, Arkadeep Acharya, Sriparna Saha, et al.
This survey systematically classifies Vision-Language Models (VLMs) based on their input modalities and provides a comprehensive taxonomy of current methodologies and future directions.
Ce Zhou, Qian Li, Chen Li, et al.
A comprehensive survey tracing the evolution of pretrained foundation models from BERT to ChatGPT, covering architectures, training methods, and applications across modalities.
Jun Ma, Yuting He, Feifei Li, et al.
MedSAM is a foundation model for universal medical image segmentation trained on 1.57M image-mask pairs across 10 modalities and 30+ cancer types, outperforming modality-wise specialist models.
Nahian Siddique, Sidike Paheding, Colin Elkin, et al.
A narrative literature review examining U-Net architecture developments, breakthroughs, and applications in medical image segmentation across multiple modalities.