Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
FreeAdvancing multi-modal LLMs with interpretable chain-of-thought reasoning
About Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
Visual CoT is a large-scale dataset and benchmark designed to advance chain-of-thought reasoning in multi-modal large language models (MLLMs). The dataset comprises 438k question-answer pairs, each annotated with intermediate bounding boxes that highlight key visual regions relevant to answering the question; approximately 98k of these pairs also include detailed reasoning steps. The authors propose a multi-turn processing pipeline that dynamically focuses on visual inputs and generates interpretable chain-of-thought thoughts. A corresponding benchmark evaluates MLLMs on tasks requiring specific local region identification. The dataset, benchmark, and pre-trained models are publicly available to support further research in improving MLLM interpretability and handling complex visual inputs.
Key Features
Pros & Cons
- Provides interpretable reasoning through intermediate bounding boxes and steps
- Large-scale dataset with diverse annotations
- Specifically addresses model limitations on complex visual inputs
- Includes a dedicated benchmark for standardized evaluation
- Open-source and publicly available for free
- Primarily focuses on local region identification, not general visual reasoning
- Multi-turn processing pipeline may increase computational overhead