Preprint
Machine Learning

Tabpfn-2.5: Advancing the state of the art in tabular foundation models

L'eo Grinsztajn, Klemens Floge, Oscar Key, Felix Birkel, Philipp Jund, Brendan Roof, Benjamin Jager, Dominik Safaric, Simone Alessi, A. Hayler, Mihir Manium, Rose Yu, F. Jablonski, Shi Bin Hoo, Anurag Garg, Jake Robertson, Magnus Buhler, Vladyslav Moroshan, Lennart Purucker, Clara Cornu, Lilly Charlotte Wehrhahn, Alessandro Bonetto, Bernhard Scholkopf, Sauraj Gambhir, N. Hollmann, Frank Hutter
November 1, 2025arXiv.org118 citations

118

Citations

17

Influential Citations

arXiv.org

Venue

2025

Year

Abstract

… We are actively developing new techniques including retrieval, fine-tuning, and novel architectures - and anticipate that systems based on Tabular Foundation Models (TFMs) will define …

Analysis

Why This Paper Matters

Tabular data remains the most common data format in industry, yet it has lagged behind other domains in the foundation model revolution. Tabpfn-2.5 addresses this gap by advancing the state of the art in tabular foundation models (TFMs). The paper signals a shift toward applying large-scale pretraining and adaptation techniques to structured data, which could unlock new levels of performance and efficiency for a wide range of applications.

The authors explicitly state that they are actively developing techniques such as retrieval, fine-tuning, and novel architectures. This indicates a maturing field where tabular models are moving beyond simple baselines to more sophisticated, adaptable systems. The anticipation that TFMs will define future systems underscores the potential for these models to become standard tools in data science pipelines.

Technical Contributions

The paper's key innovations include:

  • Retrieval: Integrating retrieval mechanisms to fetch relevant training examples or knowledge during inference, improving model predictions.
  • Fine-tuning: Adapting pretrained tabular models to specific datasets or domains, enabling better performance on downstream tasks.
  • Novel architectures: Exploring new model designs tailored to tabular data, potentially addressing issues like feature interactions and missing values.
  • Foundation model paradigm: Applying the pretrain-then-adapt approach to tabular data, which has proven successful in NLP and vision.

These contributions collectively aim to improve the accuracy, robustness, and usability of tabular models.

Results

While the abstract does not provide specific numerical results, the paper claims to advance the state of the art in tabular foundation models. Given the citation count of 118, the work has already gained attention in the research community. The lack of explicit metrics in the abstract suggests that detailed results are presented in the full paper, likely including comparisons against existing TFMs and traditional methods like gradient boosting.

Significance

The broader impact of Tabpfn-2.5 lies in its potential to democratize high-performance tabular modeling. By leveraging foundation model techniques, it could reduce the need for extensive feature engineering and model tuning, making advanced machine learning more accessible. This could accelerate progress in fields like finance, healthcare, and e-commerce, where tabular data is prevalent. Moreover, the exploration of retrieval and fine-tuning for tabular data opens new research directions, potentially inspiring further innovations in structured data learning.