ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
85
Citations
2
Influential Citations
arXiv.org
Venue
2024
Year
… Multimodal large language models (MLLMs) have experienced rapid advancements, driven by significant improvements in deep learning techniques [12, 13, 14, 15, 16, 17]. By …
Multimodal large language models (MLLMs) have become central to AI applications, yet their opacity hinders trust and adoption in critical domains. This survey addresses the urgent need for explainability and interpretability in these models, providing a structured overview of the current landscape. By systematically categorizing methods, the paper helps researchers navigate a fragmented field and identifies key gaps that must be filled to build more reliable systems.
The timing is crucial: as MLLMs are deployed in healthcare, autonomous driving, and content moderation, the ability to explain decisions becomes not just desirable but necessary. This survey consolidates knowledge, making it easier for new researchers to enter the field and for practitioners to select appropriate techniques.
The paper's primary contribution is a comprehensive taxonomy that organizes explainability methods along multiple dimensions:
Additionally, the survey discusses the unique challenges of multimodal interpretability, such as aligning explanations across modalities and handling the complexity of fused representations.
As a survey, the paper does not present new experimental results. Instead, it synthesizes findings from 85 cited works, offering a qualitative analysis of the state of the art. It highlights that while many methods exist, there is a lack of standardized evaluation, making comparisons difficult. The survey also notes that most current methods focus on vision-language tasks, with fewer addressing audio or video modalities.
The survey provides a roadmap for future research in explainable multimodal AI. By identifying gaps—such as the need for more robust evaluation metrics and user-centric explanations—it encourages the community to move beyond ad-hoc solutions. This work is likely to influence both academic research and industrial practice, promoting the development of MLLMs that are not only powerful but also transparent and accountable. As regulations around AI transparency tighten, such surveys become essential resources for ensuring compliance and building public trust.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba