ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2024
Year
… In this work, we introduce SEED-Bench, a large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs) in terms of hierarchical capabilities, including the …
Multimodal Large Language Models (MLLMs) are rapidly advancing, but their evaluation remains fragmented. SEED-Bench addresses this by providing a comprehensive, large-scale benchmark that systematically assesses hierarchical capabilities—from basic perception to complex reasoning. This matters because without standardized evaluation, progress in MLLMs is difficult to measure and compare. The benchmark fills a critical gap in the AI ecosystem, enabling researchers to identify specific strengths and weaknesses of their models.
The paper reports that current MLLMs show significant performance variation across capability levels. Models excel at basic perception tasks but struggle with complex reasoning that requires integrating multiple modalities. The benchmark reveals specific failure modes, such as difficulties with temporal reasoning and cross-modal grounding.
SEED-Bench sets a new standard for MLLM evaluation, similar to how GLUE and SuperGLUE advanced NLP. It will likely drive improvements in multimodal architectures by highlighting areas needing innovation. The benchmark's hierarchical design also encourages development of models with more robust reasoning abilities, moving beyond simple pattern matching.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba