EpiBench: Multi-turn Research Workflows for Multimodal Agents (April 2026) logo

EpiBench: Multi-turn Research Workflows for Multimodal Agents (April 2026)

Free

Benchmarks multimodal agents on episodic scientific research workflows — literature search, figure extraction, cross-paper synthesis; built on smolagents with persistent memory and tool use

FreeFree tier
Type
Open Source

About EpiBench: Multi-turn Research Workflows for Multimodal Agents (April 2026)

EpiBench is an episodic multi-turn multimodal benchmark designed to evaluate AI agents on short scientific research workflows. It requires agents to proactively search literature, extract information from figures and tables, integrate evidence across multiple papers, and sustain use of accumulated evidence over multiple turns to answer objective questions that demand cross-paper comparisons and multi-figure integration. The benchmark introduces a process-level evaluation framework for fine-grained testing and diagnosis of research agents. Experiments show that even the leading model achieves only 29.23% accuracy on the hard split, highlighting significant room for improvement in multi-turn, multi-evidence research workflows.

Key Features

Episodic multi-turn benchmark instantiating short research workflows
Requires proactive literature search, figure and table extraction, and cross-paper evidence integration
Process-level evaluation framework for fine-grained testing and diagnosis of research agents
Evaluates sustained use of accumulated evidence over multiple turns
Covers objective questions requiring cross-paper comparisons and multi-figure integration

Pros & Cons

Pros
  • Introduces a novel benchmark for multi-turn, multi-evidence research workflows not covered by existing evaluations
  • Provides process-level evaluation for detailed diagnostic insights
  • Open-access benchmark available on arXiv with downloadable code and data
  • Reveals significant performance gaps in current models (leading model only 29.23% on hard split)
Cons
  • Currently only covers short research workflows; longer or more complex workflows not yet addressed
  • Accuracy on the hard split is very low (29.23%), indicating the benchmark is extremely challenging
  • Limited to scientific research domain; not generalizable to other multi-turn tasks

Best For

Evaluating AI agents on scientific research tasksBenchmarking multimodal agents for literature review and evidence synthesisTesting proactive search and multi-turn reasoning in research contextsDiagnosing weaknesses in multi-evidence integration and memory usage

FAQ

What types of tasks does EpiBench evaluate?
EpiBench evaluates agents on multi-turn scientific research workflows requiring proactive literature search, extraction of information from figures and tables, integration of evidence across multiple papers, and sustained use of accumulated evidence to answer objective questions.
How is EpiBench different from existing benchmarks?
Existing benchmarks largely under-evaluate proactive search, multi-evidence integration, and sustained evidence use over time. EpiBench specifically targets these capabilities with an episodic multi-turn design and a process-level evaluation framework.
What performance did current models achieve on EpiBench?
The leading model achieved only 29.23% accuracy on the hard split of EpiBench, indicating substantial room for improvement in multi-turn, multi-evidence research workflows.