Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Unknown
SpatialVLM endows vision-language models with spatial reasoning capabilities by generating and training on a large-scale spatial VQA dataset.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
SpatialVLM endows vision-language models with spatial reasoning capabilities by generating and training on a large-scale spatial VQA dataset.
Unknown
Introduces GPT4Scene, a framework for understanding 3D scenes from videos using vision-language models, and ScanAlign, a multimodal dataset with 165K aligned pairs for training smaller VLMs.
Unknown
SHARP is a representation-level intervention framework that modulates hallucination in LVLMs by steering internal representations.
Unknown
Proposes MRRE, a training-free inference-time method using representation engineering to enhance multilingual reasoning in LLMs and LVLMs without additional training data.
Unknown
This paper explores how generative world models can provide priors to vision-language models (VLMs) for improved understanding of world dynamics in common vision tasks.
Unknown
FastVLM introduces an efficient vision encoding method that reduces computational cost while maintaining high performance in text-rich image understanding tasks.
Unknown
This paper introduces LVLM-eHub, a comprehensive evaluation benchmark for Large Vision-Language Models, systematically assessing their capabilities across diverse multimodal tasks.
Unknown
This paper assesses the effectiveness of recent large vision-language models (LVLMs) in achieving artificial general intelligence, focusing on their performance across various tasks.
Unknown
This paper proposes an efficient and unbalanced multi-agent collaboration framework for long-horizon planning, using a final-outcome-based reward and a VLM evaluator to provide stable feedback.
Unknown
This paper evaluates closed- and open-source VLMs as planners, showing that spatially grounded long-horizon planning remains a major unsolved challenge.
Yang Zhou, Zixuan Huang, Sunzhu Li, et al.
SpatialCLI teaches VLMs to reason with spatial tools and progressively internalize specialist perceptual capabilities, raising Qwen3-VL-8B-Instruct from 29.3% to 84.6% on MindCube.
Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati, et al.
UltraViT is a latency-optimized vision encoder for LVLMs using a pyramidal architecture with heterogeneous spatial mixers and a two-stage generative pre-training strategy, achieving 1.7x speedup on-device.