MERRIN: Multimodal Evidence Retrieval in Noisy Web Environments (April 2026) logo

MERRIN: Multimodal Evidence Retrieval in Noisy Web Environments (April 2026)

Free

Benchmark for multimodal evidence retrieval and multi-hop reasoning over noisy web content — even strongest agent (Gemini-3.1-Pro) achieves only 40.1%; finds more search ≠ better performance

FreeFree tier
Type
Open Source

About MERRIN: Multimodal Evidence Retrieval in Noisy Web Environments (April 2026)

MERRIN (Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments) is a human-annotated benchmark designed to evaluate AI search-augmented agents on multimodal evidence retrieval and multi-hop reasoning from noisy web content. Unlike prior benchmarks, it uses natural language queries without explicit modality cues, incorporates underexplored modalities such as video and audio, and requires agents to retrieve and reason over complex, often conflicting, multimodal evidence from web sources. The benchmark evaluates agents across three search settings (no search, native search, and agentic search) and is highly challenging: the best-performing agent achieves only 40.1% accuracy, with an average of 22.3% across ten models (including GPT-5.4-mini, Gemini 3/3.1 Flash/Pro, and Qwen3-4B/30B/235B). Key findings reveal that stronger agents often over-explore and rely too heavily on text, leading to modest gains and highlighting the need for more robust search and reasoning across diverse modalities.

Key Features

Human-annotated benchmark for multimodal evidence retrieval and reasoning
Natural language queries without explicit modality cues
Includes video and audio modalities alongside text and images
Requires multi-hop reasoning over noisy, conflicting web sources
Evaluates agents across three search settings: no search, native search, agentic search
Highly challenging: average accuracy 22.3%, best agent 40.1%
Reveals over-exploration and text modality bias in current agents

Pros & Cons

Pros
  • Human-annotated with high-quality ground truth data
  • Covers underexplored modalities (video, audio) not typically included in similar benchmarks
  • Uses natural language queries, simulating real user search behavior
  • Provides insights into agent inefficiencies like over-exploration and text overreliance
  • Reveals clear performance gaps, driving future research in robust search agents
Cons
  • Research-focused benchmark, not a production-ready tool or application
  • Current evaluations limited to ten models and specific search settings
  • May not fully capture the diversity of real-world web noise and conflicting sources
  • No integrated pipeline or software package provided for easy reuse

Best For

Evaluating AI search agents' multimodal reasoning capabilitiesResearch in multimodal evidence retrieval and web search augmentationTesting robustness of agents to conflicting and partially relevant web contentBenchmarking improvements in agentic search and tool useStudying the impact of modality selection on retrieval accuracy

FAQ

What is MERRIN?
MERRIN (Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments) is a human-annotated benchmark for evaluating AI search-augmented agents on multimodal evidence retrieval and multi-hop reasoning over noisy web content.
What modalities does MERRIN cover?
MERRIN includes text, image, video, and audio modalities, with queries that do not explicitly indicate which modality is relevant.
How challenging is MERRIN?
The average accuracy across all evaluated agents is 22.3%, with the best-performing agent (Gemini Deep Research) achieving only 40.1%, making it a highly difficult benchmark.
What are the key findings from MERRIN?
Stronger agents often over-explore and take more steps, yet gains are modest due to distraction from conflicting content and an overreliance on text modalities.