FeynmanBench: Diagrammatic Physics Reasoning for MLLMs (April 2026)
FreeFirst benchmark for Feynman diagram tasks — evaluates multistep diagrammatic reasoning requiring conservation laws, symmetry constraints, and graph topology; 2000+ tasks across Standard Model interactions
About FeynmanBench: Diagrammatic Physics Reasoning for MLLMs (April 2026)
FeynmanBench is a benchmark of over 2,000 tasks centered on Feynman diagrams spanning the electromagnetic, weak, and strong interactions of the Standard Model. It evaluates multimodal large language models (MLLMs) on diagrammatic physics reasoning, requiring models to recover the full physical content – vertex inventory, propagator types, topological connectivity, momentum routing, and complete scattering amplitude – from a diagram image coupled with minimal textual conventions. An automated generation and verification pipeline produces the diagrams, annotations, and reference answers under standardized rules. Evaluation of 19 state-of-the-art MLLMs reveals a consistent failure pattern: models perform well on local recognition (70–95%) but collapse to 13–17% on topological reconstruction and near zero on full algebraic derivation, highlighting fundamental limitations in topology-sensitive scientific reasoning.
Key Features
Pros & Cons
- First standardized benchmark specifically for Feynman diagram reasoning
- Reveals clear failure patterns in current MLLMs (local recognition vs. topology/algebra)
- Controlled testbed with automated generation and verification
- Covers multiple interaction types and difficulty levels
- Open source and free to use
- Specialized to Feynman diagrams; does not cover other scientific diagram types
- May require domain knowledge in particle physics to interpret results
- Current MLLMs perform very poorly on higher-level reasoning tasks (topology, algebra)