MMToM-QA logo

MMToM-QA

Free

a multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs and goals.

FreeFree tier
Inputs: text, video
Type
Open Source

About MMToM-QA

MMToM-QA is a multimodal question-answering benchmark designed to systematically evaluate AI models' cognitive ability to understand human beliefs and goals (theory of mind). It assesses models on multimodal data as well as unimodal data (text-only and video-only). Questions span seven categories focusing on belief inference and goal inference in rich and diverse situations. The public leaderboard tracks state-of-the-art performance across models such as AutoToM, BIP-ALM, GPT-4o, and others, with separate rankings for multimodal, text-only, and video-only tasks. The benchmark is open-source, with submission instructions available on GitHub.

Key Features

Evaluates cognitive ability for belief and goal inference
Covers seven question categories
Supports multimodal, text-only, and video-only inputs
Public leaderboard for model comparison
Open-source benchmark with GitHub repository
Systematic evaluation of both human and AI performance

Pros & Cons

Pros
  • Systematic evaluation across multiple modalities
  • Public leaderboard with transparent rankings
  • Covers diverse situations for belief and goal inference
  • Open-source and community-driven
  • Includes human baseline for comparison
Cons
  • Limited documentation on benchmark creation and details beyond the leaderboard page
  • Requires understanding of theory of mind concepts to interpret results
  • May not cover all aspects of human cognition or real-world scenarios

Best For

Benchmarking AI theory of mind capabilitiesResearch in cognitive AI and multimodal understandingEvaluating large language models and vision-language modelsComparing human vs AI performance on belief and goal inferenceAnalyzing model performance across different input modalities

FAQ

What does MMToM-QA evaluate?
MMToM-QA evaluates AI models' ability to understand human beliefs and goals (theory of mind) through question answering across seven categories. It assesses both belief inference and goal inference.
What input modalities does MMToM-QA support?
The benchmark covers multimodal data as well as unimodal data: text-only and video-only. The leaderboard has separate rankings for each modality.
How are scores calculated?
Scores are reported as accuracy percentages for belief inference, goal inference, and an overall 'All' score. The leaderboard shows the top-performing methods for each modality.