Preprint
Large Language Models

Explainable and interpretable multimodal large language models: A comprehensive survey

Yunkai Dang, Kaichen Huang, Jiahao Huo, Yibo Yan, Sirui Huang, Dongrui Liu, Mengxi Gao, Jie Zhang, Chen Qian, Kun Wang, Yong Liu, Jing Shao, Hui Xiong, Xuming Hu
December 1, 2024arXiv.org85 citations

85

Citations

2

Influential Citations

arXiv.org

Venue

2024

Year

Abstract

… Multimodal large language models (MLLMs) have experienced rapid advancements, driven by significant improvements in deep learning techniques [12, 13, 14, 15, 16, 17]. By …

Analysis

Why This Paper Matters

Multimodal large language models (MLLMs) have become central to AI applications, yet their opacity hinders trust and adoption in critical domains. This survey addresses the urgent need for explainability and interpretability in these models, providing a structured overview of the current landscape. By systematically categorizing methods, the paper helps researchers navigate a fragmented field and identifies key gaps that must be filled to build more reliable systems.

The timing is crucial: as MLLMs are deployed in healthcare, autonomous driving, and content moderation, the ability to explain decisions becomes not just desirable but necessary. This survey consolidates knowledge, making it easier for new researchers to enter the field and for practitioners to select appropriate techniques.

Technical Contributions

The paper's primary contribution is a comprehensive taxonomy that organizes explainability methods along multiple dimensions:

  • Stage-based categorization: Distinguishes between intrinsic (built-in) and post-hoc (external) explanation methods.
  • Modality coverage: Addresses methods for vision, language, and cross-modal interactions.
  • Technique families: Covers attention-based explanations, concept-based models, surrogate models, and counterfactual reasoning.
  • Evaluation frameworks: Summarizes metrics and benchmarks used to assess explanation fidelity, consistency, and user satisfaction.

Additionally, the survey discusses the unique challenges of multimodal interpretability, such as aligning explanations across modalities and handling the complexity of fused representations.

Results

As a survey, the paper does not present new experimental results. Instead, it synthesizes findings from 85 cited works, offering a qualitative analysis of the state of the art. It highlights that while many methods exist, there is a lack of standardized evaluation, making comparisons difficult. The survey also notes that most current methods focus on vision-language tasks, with fewer addressing audio or video modalities.

Significance

The survey provides a roadmap for future research in explainable multimodal AI. By identifying gaps—such as the need for more robust evaluation metrics and user-centric explanations—it encourages the community to move beyond ad-hoc solutions. This work is likely to influence both academic research and industrial practice, promoting the development of MLLMs that are not only powerful but also transparent and accountable. As regulations around AI transparency tighten, such surveys become essential resources for ensuring compliance and building public trust.