ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2024
Year
We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can …
This paper introduces Blink, a benchmark that targets a critical yet underexplored aspect of multimodal large language models (LLMs): core visual perception. While existing benchmarks often focus on high-level reasoning, visual question answering, or commonsense knowledge, Blink isolates fundamental perceptual tasks that humans perform effortlessly but current models struggle with. This distinction between 'seeing' (processing pixels) and 'perceiving' (understanding spatial relationships, object properties, and visual details) is crucial for advancing AI toward human-like visual intelligence.
The significance of Blink lies in its potential to expose a fundamental limitation in current multimodal LLMs. If models cannot reliably perform basic perception tasks—such as identifying which of two objects is closer, counting items, or matching visual patterns—then their performance on more complex tasks may be built on shaky foundations. This benchmark could serve as a wake-up call for the research community, prompting a shift toward improving low-level perception rather than solely scaling up model size or training data.
While the abstract is truncated, it indicates that current multimodal LLMs perform poorly on Blink tasks, often near random chance, whereas humans achieve near-perfect accuracy. This stark contrast underscores the 'seeing but not perceiving' phenomenon. The benchmark likely includes quantitative results for several models, but specific numbers are not available in the provided text. The key takeaway is that despite advances in multimodal LLMs, their perceptual abilities remain far behind human capabilities.
Blink has the potential to influence the direction of multimodal AI research by highlighting a critical weakness. It could encourage the development of models that integrate better visual perception mechanisms, perhaps through more sophisticated vision encoders or training objectives that emphasize low-level features. Moreover, it provides a benchmark that can be used to track progress over time, similar to how ImageNet drove advances in computer vision. For practitioners, Blink offers a tool to evaluate and compare models on perceptual tasks, which is essential for applications requiring precise visual understanding, such as robotics, autonomous driving, and augmented reality. Ultimately, this work underscores that true multimodal intelligence requires not just language understanding but also robust visual perception.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba