3D-Aware VLMs with Implicit and Explicit Geometries
Wenhao Li, Xueying Jiang, Quanhao Qian, et al.
VLM-IE3D enhances 3D spatial awareness of vision-language models by fusing implicit and explicit 3D geometries learned from RGB videos.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Wenhao Li, Xueying Jiang, Quanhao Qian, et al.
VLM-IE3D enhances 3D spatial awareness of vision-language models by fusing implicit and explicit 3D geometries learned from RGB videos.
Karan Goyal, Afreen Hossain, Debojyoti Das, et al.
ENTRAP-VL is a taxonomically structured, dual-modality dataset and evaluation protocol for studying contextual entrainment in vision-language models.
Unknown
This paper proposes a reliable multimodal RAG method for medical vision-language models to improve factual accuracy.
Unknown
MMed-RAG is a versatile multimodal RAG system designed for medical vision-language models to generate more factual responses.
Unknown
This paper improves vision features for vision-language tasks by developing an enhanced object detection model.
Unknown
This paper proposes encoder-free vision-language models that directly process visual features without a separate vision encoder, simplifying architecture and improving efficiency.
Akash Ghosh, Arkadeep Acharya, Sriparna Saha, et al.
This survey systematically classifies Vision-Language Models (VLMs) based on their input modalities and provides a comprehensive taxonomy of current methodologies and future directions.
Unknown
This paper identifies two key issues in LVLM evaluation: visual content is often unnecessary for many samples, and proposes a more rigorous evaluation framework.
Unknown
This survey comprehensively reviews benchmark evaluations, applications, and challenges of large vision-language models, synthesizing progress across multimodal AI.
Hugo Laurençon, Andrés Marafioti, Victor Sanh, et al.
This paper provides a comprehensive tutorial on building vision-language models, focusing on architecture, data, and training methods.
Unknown
Proposes a simple continuous prompt learning approach to automate prompt engineering for pre-trained vision-language models.
Unknown
A survey evaluating large vision-language models through benchmark assessments, highlighting challenges and future directions.