Preprint
Large Language Models

Multimodal large language models: A survey

Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, Philip S. Yu
January 1, 2023BigData Congress [Services Society]407 citations

407

Citations

8

Influential Citations

BigData Congress [Services Society]

Venue

2023

Year

Abstract

The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models …

Analysis

Why This Paper Matters

This survey arrives at a critical juncture in AI research, where large language models (LLMs) are increasingly being extended beyond text to handle multiple modalities. The integration of vision, audio, and other data types into LLMs promises more robust and human-like understanding, enabling applications from image captioning to audio transcription and beyond. By providing a structured overview of the field, the paper helps researchers navigate the fragmented landscape of multimodal LLMs (MLLMs), which has seen explosive growth in recent years.

The paper's significance lies in its comprehensive taxonomy, which categorizes MLLMs by the types of modalities they handle (e.g., vision-language, audio-language) and the architectural strategies used for fusion (e.g., early fusion, cross-attention, modality-specific encoders). This framework allows practitioners to quickly identify relevant models and techniques for their own work. Moreover, the survey discusses training paradigms, such as multimodal pretraining and instruction tuning, which are crucial for achieving state-of-the-art performance.

Technical Contributions

  • Taxonomy of MLLMs: The paper proposes a clear categorization based on modality types (e.g., image-text, video-text, audio-text) and fusion methods (e.g., encoder-decoder, transformer-based, modular).
  • Comprehensive Review: It covers a wide range of models, including CLIP, Flamingo, BLIP-2, LLaVA, and others, detailing their architectures, training data, and key innovations.
  • Benchmark and Dataset Analysis: The survey compiles commonly used multimodal datasets (e.g., COCO, Flickr30k, AudioSet) and evaluation benchmarks (e.g., VQA, captioning, retrieval), providing a practical resource for model comparison.
  • Open Challenges: It identifies limitations such as modality imbalance, alignment difficulties, and the need for more diverse and high-quality multimodal data.

Results

As a survey paper, the authors do not present new experimental results. Instead, they synthesize findings from dozens of prior works, noting that MLLMs have achieved impressive performance on tasks like visual question answering (e.g., LLaVA achieving ~85% on VQAv2) and image captioning (e.g., BLIP-2 reaching state-of-the-art CIDEr scores). The paper also highlights that multimodal pretraining on large-scale datasets (e.g., LAION-5B) is a key driver of performance gains.

Significance

This survey has broad impact by providing a structured entry point for researchers new to multimodal LLMs, as well as a reference for experts tracking the field's evolution. It underscores the trend toward unified models that can process multiple modalities simultaneously, which is likely to be a cornerstone of future AI systems. The identified open challenges—such as efficient fusion, cross-modal reasoning, and real-world deployment—will guide future research efforts. For practitioners at Neura Market, this paper offers a valuable map of the current landscape, helping to identify promising directions for product development and investment in multimodal AI technologies.