We-Math
Freea benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.
About We-Math
We-Math is a benchmark designed to evaluate large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning. It goes beyond end-to-end performance by decomposing composite visual math problems into sub-problems based on hierarchical knowledge concepts. The benchmark includes 6.5K visual math problems spanning 67 knowledge concepts across 5 granularity levels, and introduces a four-dimensional metric (Insufficient Knowledge, Inadequate Generalization, Complete Mastery, Rote Memorization) to assess LMM reasoning processes. Evaluations reveal that GPT-4o transitions from knowledge insufficiency to generalization issues, while other LMMs rely more on rote memorization. The benchmark is openly available with code, dataset, and leaderboard.
Key Features
Pros & Cons
- Provides multi-dimensional evaluation of reasoning beyond simple accuracy
- Hierarchical knowledge structure aligned with real textbook concepts
- Knowledge decomposition enables detailed diagnosis of LMM issues
- Open-source benchmark with available code, dataset, and leaderboard
- First benchmark to explore problem-solving principles in LMMs
- Primarily limited to visual mathematical reasoning from textbook concepts, may not generalize to other domains
- Requires careful decomposition of problems which may be resource-intensive