Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? (April 2026)
FreeSystematic evaluation of agentic capability in multimodal LLMs — decomposes tasks into perception, reasoning, and action levels; reveals where agentic loops help vs. where they add overhead
About Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? (April 2026)
Agentic-MME is a process-verified benchmark for evaluating multimodal agentic capabilities in large language models. It comprises 418 real-world tasks across 6 domains and 3 difficulty levels, featuring over 2,000 stepwise checkpoints that each required 10+ person-hours of manual annotation. The benchmark decomposes tasks into perception, reasoning, and action levels, and uses a dual-axis (S-axis and V-axis) human reference trajectory to enable fine-grained process verification. It also introduces an overthinking metric to quantify efficiency by comparing model trajectories against human references. Evaluations show that the best model (Gemini3-pro) achieves 56.3% overall accuracy, dropping to 23.0% on Level-3 tasks, highlighting the challenge of real-world multimodal agentic problem solving.
Key Features
Pros & Cons
- Comprehensive process verification beyond final-answer evaluation
- Granular stepwise checkpoints with high-quality manual annotation
- Includes efficiency metric (overthinking) to quantify overhead
- Covers diverse real-world tasks with multiple difficulty levels
- Open-source benchmark available for research community
- Best models still achieve only 56.3% overall accuracy, with sharp drop on hard tasks (23.0% on Level-3)
- Benchmark may not fully capture dynamic real-world tool interactions
- Overthinking metric may not generalize to all agentic architectures