M3CoT
Freea benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.
About M3CoT
M3CoT is a comprehensive benchmark designed to evaluate the multimodal reasoning capabilities of large language models (LLMs) across a diverse set of subjects. It features a test split of 2,359 samples covering categories including language, natural science, social science, physical commonsense, social commonsense, temporal reasoning, algebra, geometry, and theory. The benchmark supports multiple evaluation settings such as zero-shot, tool-usage, fine-tuning, and preference-optimization, and provides a leaderboard for comparing model performance. Community contributions to the leaderboard are encouraged via email or GitHub.
Key Features
Pros & Cons
- Comprehensive coverage of multimodal reasoning types
- Open-source benchmark with publicly available test set
- Multiple evaluation settings allow flexible testing
- Community-contributed leaderboard fosters collaboration
- Leaderboard data collected manually may contain errors or ambiguities
- Missing data for some models due to incomplete reporting in papers
- Limited to evaluation; no built-in tool usage or generation capabilities