FeynmanBench: Diagrammatic Physics Reasoning for MLLMs (April 2026) logo

FeynmanBench: Diagrammatic Physics Reasoning for MLLMs (April 2026)

Free

First benchmark for Feynman diagram tasks — evaluates multistep diagrammatic reasoning requiring conservation laws, symmetry constraints, and graph topology; 2000+ tasks across Standard Model interactions

FreeFree tier
Type
Open Source

About FeynmanBench: Diagrammatic Physics Reasoning for MLLMs (April 2026)

FeynmanBench is a benchmark of over 2,000 tasks centered on Feynman diagrams spanning the electromagnetic, weak, and strong interactions of the Standard Model. It evaluates multimodal large language models (MLLMs) on diagrammatic physics reasoning, requiring models to recover the full physical content – vertex inventory, propagator types, topological connectivity, momentum routing, and complete scattering amplitude – from a diagram image coupled with minimal textual conventions. An automated generation and verification pipeline produces the diagrams, annotations, and reference answers under standardized rules. Evaluation of 19 state-of-the-art MLLMs reveals a consistent failure pattern: models perform well on local recognition (70–95%) but collapse to 13–17% on topological reconstruction and near zero on full algebraic derivation, highlighting fundamental limitations in topology-sensitive scientific reasoning.

Key Features

Over 2,000 tasks centered on Feynman diagrams
Covers electromagnetic, weak, and strong interactions of the Standard Model
Automated generation and verification pipeline for diagrams, annotations, and reference answers
Tasks range from local recognition (vertex/propagator ID) to topological reconstruction and full algebraic derivation
Evaluated on 19 state-of-the-art multimodal LLMs
Standardized rules ensure controlled comparison

Pros & Cons

Pros
  • First standardized benchmark specifically for Feynman diagram reasoning
  • Reveals clear failure patterns in current MLLMs (local recognition vs. topology/algebra)
  • Controlled testbed with automated generation and verification
  • Covers multiple interaction types and difficulty levels
  • Open source and free to use
Cons
  • Specialized to Feynman diagrams; does not cover other scientific diagram types
  • May require domain knowledge in particle physics to interpret results
  • Current MLLMs perform very poorly on higher-level reasoning tasks (topology, algebra)

Best For

Evaluating multimodal LLMs on formal scientific diagram reasoningTesting topology-sensitive reasoning in AI modelsResearch in physics-informed AI and diagrammatic understandingBenchmarking progress in multimodal learning for scientific domains

FAQ

What is FeynmanBench?
FeynmanBench is a benchmark of over 2,000 tasks centered on Feynman diagrams that evaluates multimodal LLMs on diagrammatic physics reasoning, requiring them to recover physical content such as vertex inventory, propagator types, and scattering amplitude from diagram images.
What types of interactions are covered?
The benchmark covers Feynman diagrams from the electromagnetic, weak, and strong interactions of the Standard Model.
How are the tasks generated?
Tasks are generated using an automated pipeline that produces diagrams, annotations, and reference answers under standardized rules, ensuring consistency and verifiability.
How do current multimodal LLMs perform on FeynmanBench?
In evaluations, models achieve 70–95% on local recognition (vertex and propagator identification) but drop to 13–17% on topological reconstruction and near zero on full algebraic derivation, indicating a major gap in topology-sensitive reasoning.