Preprint
Machine Learning

On the fragility of benchmark contamination detection in reasoning models

January 1, 2026

0

Citations

0

Influential Citations

Venue

2026

Year

Abstract

… In this paper, we present the first systematic study of benchmark contamination in LRMs, structured around two points where contamination can happen. In particular, Stage I (pre-LRM) …

Analysis

Why This Paper Matters

Benchmark contamination—where test data leaks into training data—has long plagued machine learning evaluation, but its impact on large reasoning models (LRMs) is particularly insidious. These models, designed to perform multi-step logical and mathematical reasoning, are often trained on vast internet-scale corpora that may inadvertently include benchmark questions. As LRMs become more capable, the risk of inflated performance due to contamination grows, undermining the validity of leaderboards and scientific conclusions. This paper is the first to systematically study contamination in LRMs, providing a structured analysis that is urgently needed as these models are deployed in high-stakes domains.

The paper's two-stage framework (pre-LRM and post-LRM) offers a clear lens for understanding where contamination can enter. Pre-LRM contamination occurs when benchmark data is included in the pretraining corpus, while post-LRM contamination can happen during fine-tuning or alignment. This distinction is crucial because it informs where detection and mitigation efforts should focus. By highlighting the fragility of existing detection methods, the authors challenge the community to rethink evaluation protocols for reasoning models.

Technical Contributions

  • Two-Stage Contamination Framework: The paper introduces a structured approach to categorize contamination points, enabling more precise analysis and detection.
  • Systematic Study: It provides the first comprehensive examination of contamination in LRMs, filling a significant gap in the literature.
  • Fragility Analysis: The authors demonstrate that current contamination detection techniques, which may work for standard LLMs, are insufficient for reasoning models due to their unique output characteristics and training paradigms.
  • Call for New Methods: The work motivates the development of contamination detection methods that are robust to the complexities of reasoning models.

Results

The abstract does not provide specific quantitative results, but the key finding is that existing contamination detection methods are fragile when applied to LRMs. This suggests that many reported performance numbers for reasoning models may be inflated due to undetected contamination. The paper likely includes experiments showing how detection methods fail, but these details are not in the abstract.

Significance

This paper has profound implications for AI evaluation. As reasoning models become more prevalent, the integrity of benchmarks is paramount. The fragility of contamination detection means that the AI community cannot trust current evaluations of LRMs, potentially leading to overestimation of capabilities and misguided research directions. This work serves as a wake-up call, urging researchers to develop more rigorous evaluation standards and contamination detection techniques. It also highlights the need for transparency in model training data and benchmark curation. Ultimately, this paper contributes to the broader goal of ensuring that AI progress is measured accurately and honestly.