Preprint
Multimodal AI

Perception, reason, think, and plan: A survey on large multimodal reasoning models

January 1, 2025

0

Citations

0

Influential Citations

Venue

2025

Year

Abstract

… Large Multimodal Reasoning Models (LMRMs) have emerged as a promising paradigm, … the conceptual direction of native large multimodal reasoning models (N-LMRMs), which aim to …

Analysis

Why This Paper Matters

This survey addresses a critical gap in multimodal AI: the lack of unified models that seamlessly combine perception, reasoning, and planning. While large language models (LLMs) and vision-language models (VLMs) have advanced separately, their integration remains ad hoc. The paper's proposal of native LMRMs (N-LMRMs) offers a conceptual framework for building more coherent and capable AI systems that can reason across modalities from the ground up.

For practitioners, this work provides a structured overview of the current landscape, helping to identify which approaches are most promising for tasks requiring complex multimodal understanding, such as robotics, autonomous driving, and interactive AI assistants.

Technical Contributions

  • Taxonomy of LMRMs: The paper categorizes existing models based on how they handle perception, reasoning, and planning, distinguishing between modular and native approaches.
  • Concept of N-LMRMs: Introduces the idea of models where reasoning is inherently multimodal, rather than relying on separate encoders and decoders.
  • Survey of Architectures: Reviews key models like Flamingo, GPT-4V, and Gemini, analyzing their strengths and weaknesses in multimodal reasoning tasks.
  • Future Directions: Identifies open challenges, including evaluation benchmarks, training efficiency, and interpretability.

Results

As a survey, the paper does not present new experimental results. However, it synthesizes findings from numerous prior works, noting that current LMRMs often struggle with compositional reasoning and long-horizon planning. The paper suggests that N-LMRMs could achieve better performance on tasks requiring multi-step reasoning across vision and language, though concrete metrics are not provided.

Significance

This survey lays the groundwork for a new research direction in multimodal AI. By advocating for native integration of reasoning, it challenges the prevailing modular paradigm and may influence how future models are designed. For the AI field, this could lead to more robust and versatile systems capable of complex real-world interactions.