M3CoT logo

M3CoT

Free

a benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.

FreeFree tier
Inputs: text, image
Type
Open Source

About M3CoT

M3CoT is a comprehensive benchmark designed to evaluate the multimodal reasoning capabilities of large language models (LLMs) across a diverse set of subjects. It features a test split of 2,359 samples covering categories including language, natural science, social science, physical commonsense, social commonsense, temporal reasoning, algebra, geometry, and theory. The benchmark supports multiple evaluation settings such as zero-shot, tool-usage, fine-tuning, and preference-optimization, and provides a leaderboard for comparing model performance. Community contributions to the leaderboard are encouraged via email or GitHub.

Key Features

Benchmark for multimodal reasoning evaluation
Covers nine reasoning categories: language, natural science, social science, physical commonsense, social commonsense, temporal reasoning, algebra, geometry, and theory
Test split of 2,359 samples
Supports evaluation settings: zero-shot, tool-usage, fine-tuning, and preference-optimization
Public leaderboard for comparing model accuracy
Invites community contributions via email or GitHub

Pros & Cons

Pros
  • Comprehensive coverage of multimodal reasoning types
  • Open-source benchmark with publicly available test set
  • Multiple evaluation settings allow flexible testing
  • Community-contributed leaderboard fosters collaboration
Cons
  • Leaderboard data collected manually may contain errors or ambiguities
  • Missing data for some models due to incomplete reporting in papers
  • Limited to evaluation; no built-in tool usage or generation capabilities

Best For

Evaluating multimodal reasoning abilities of large language modelsBenchmarking model performance across diverse reasoning domainsAcademic research on multimodal chain-of-thought reasoningComparing different prompting strategies and fine-tuning methods