MMToM-QA
Freea multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs and goals.
About MMToM-QA
MMToM-QA is a multimodal question-answering benchmark designed to systematically evaluate AI models' cognitive ability to understand human beliefs and goals (theory of mind). It assesses models on multimodal data as well as unimodal data (text-only and video-only). Questions span seven categories focusing on belief inference and goal inference in rich and diverse situations. The public leaderboard tracks state-of-the-art performance across models such as AutoToM, BIP-ALM, GPT-4o, and others, with separate rankings for multimodal, text-only, and video-only tasks. The benchmark is open-source, with submission instructions available on GitHub.
Key Features
Pros & Cons
- Systematic evaluation across multiple modalities
- Public leaderboard with transparent rankings
- Covers diverse situations for belief and goal inference
- Open-source and community-driven
- Includes human baseline for comparison
- Limited documentation on benchmark creation and details beyond the leaderboard page
- Requires understanding of theory of mind concepts to interpret results
- May not cover all aspects of human cognition or real-world scenarios