Sharp: Steering hallucination in lvlms via representation engineering
Unknown
SHARP is a representation-level intervention framework that modulates hallucination in LVLMs by steering internal representations.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
SHARP is a representation-level intervention framework that modulates hallucination in LVLMs by steering internal representations.
Unknown
Proposes MRRE, a training-free inference-time method using representation engineering to enhance multilingual reasoning in LLMs and LVLMs without additional training data.
Unknown
This paper introduces LVLM-eHub, a comprehensive evaluation benchmark for Large Vision-Language Models, systematically assessing their capabilities across diverse multimodal tasks.
Unknown
This paper assesses the effectiveness of recent large vision-language models (LVLMs) in achieving artificial general intelligence, focusing on their performance across various tasks.
Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati, et al.
UltraViT is a latency-optimized vision encoder for LVLMs using a pyramidal architecture with heterogeneous spatial mixers and a two-stage generative pre-training strategy, achieving 1.7x speedup on-device.
Andrés Marafioti, Orr Zohar, Miquel Farr'e, et al.
SmolVLM introduces a family of compact vision-language models achieving strong performance on resource-constrained devices through aggressive token compression and optimized encoder-LM balance.
Amirmohammad Izadi, Mohammadali Banayeeanzade, Fatemeh Askari, et al.
VISER augments visual inputs with low-level spatial structures to improve LVLM visual reasoning, achieving gains of 25% in visual search, 26.8% in counting, and 9.5% in spatial relationships.
Unknown
This paper identifies two key issues in LVLM evaluation: visual content is often unnecessary for many samples, and proposes a more rigorous evaluation framework.