Preprint
AI Safety & Alignment

PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods

Karel G.M. Moons(Utrecht University), Johanna AAG Damen(Utrecht University), T. K. Kaul(Utrecht University), Lotty Hooft(Utrecht University), Constanza L. Andaur Navarro(Utrecht University), Paula Dhiman(Nuffield Orthopaedic Centre), Andrew L. Beam(Harvard University), Ben Van Calster(KU Leuven), Leo Anthony Celi(Harvard University), Spiros Denaxas(British Heart Foundation), Alastair K. Denniston(University College Birmingham), Marzyeh Ghassemi(Massachusetts Institute of Technology), Georg Heinze(Medical University of Vienna), André Pascal Kengne(University of Cape Town), Lena Maier‐Hein(German Cancer Research Center), Xiaoxuan Liu(University Hospitals Birmingham NHS Foundation Trust), Patrícia Logullo(Nuffield Orthopaedic Centre), Melissa D. McCradden(Hospital for Sick Children), Nan Liu(Duke-NUS Medical School), Lauren Oakden‐Rayner(Australian Centre for Robotic Vision), Karandeep Singh(University of Michigan), Daniel Shu Wei Ting(Duke-NUS Medical School), Laure Wynants(Maastricht University), Bada Yang(Utrecht University), Johannes B. Reitsma(Utrecht University), Richard D Riley(University College Birmingham), Professor Gary S. Collins(Nuffield Orthopaedic Centre), Maarten van Smeden(Utrecht University)
March 24, 2025BMJ520 citations

520

Citations

19

Influential Citations

BMJ

Venue

2025

Year

Abstract

The Prediction model Risk Of Bias ASsessment Tool (PROBAST) is used to assess the quality, risk of bias, and applicability of prediction models or algorithms and of prediction model/algorithm studies. Since PROBAST’s introduction in 2019, much progress has been made in the methodology for prediction modelling and in the use of artificial intelligence, including machine learning, techniques. An update to PROBAST-2019 is thus needed. This article describes the development of PROBAST+AI. PROBAST+AI consists of two distinctive parts: model development and model evaluation. For model development, PROBAST+AI users assess quality and applicability using 16 targeted signalling questions. For model evaluation, PROBAST+AI users assess the risk of bias and applicability using 18 targeted signalling questions. Both parts contain four domains: participants and data sources, predictors, outcome, and analysis. Applicability of the prediction model is rated for the participants and data sources, predictors, and outcome domains. PROBAST+AI may replace the original PROBAST tool and allows all key stakeholders (eg, model developers, AI companies, researchers, editors, reviewers, healthcare professionals, guideline developers, and policy organisations) to examine the quality, risk of bias, and applicability of any type of prediction model in the healthcare sector, irrespective of whether regression modelling or AI techniques are used.

Analysis

Why This Paper Matters

PROBAST+AI addresses a critical gap in the evaluation of clinical prediction models. The original PROBAST, released in 2019, was designed primarily for regression-based models. With the rapid adoption of machine learning and deep learning in healthcare, there is an urgent need for a tool that can assess the risk of bias and applicability of AI-based models. This paper responds to that need by providing an updated framework that is explicitly inclusive of AI techniques.

The significance of this work lies in its potential to become a standard for regulatory and clinical evaluation. As AI models increasingly influence patient care, the ability to systematically assess their quality is essential. PROBAST+AI offers a structured approach that can be used by model developers, journals, and health technology assessment bodies, thereby promoting reproducibility and accountability.

Technical Contributions

  • Two-part structure: PROBAST+AI separates assessment for model development and model evaluation, recognizing that these stages have different risk profiles.
  • Expanded signaling questions: 16 questions for development and 18 for evaluation, covering key aspects such as data sources, predictor handling, outcome definition, and analysis methods.
  • Domain-based assessment: Retains the four domains (participants/data sources, predictors, outcome, analysis) and applicability ratings for three domains, ensuring continuity with the original tool.
  • AI-specific considerations: The tool is designed to handle complexities introduced by AI, such as feature selection, overfitting, and external validation.

Results

The paper does not present quantitative results or validation metrics. Instead, it describes the development and structure of PROBAST+AI. The abstract mentions that the tool consists of 16 and 18 signaling questions for development and evaluation, respectively, but no data on inter-rater reliability or usability are provided. This is a methodological paper, so the 'results' are the tool itself and its rationale.

Significance

PROBAST+AI has the potential to become the de facto standard for assessing prediction models in healthcare, regardless of whether they are based on regression or AI. By providing a common framework, it can facilitate comparisons across studies, improve the quality of published research, and support evidence-based adoption of AI in clinical practice. The tool's broad applicability to stakeholders—from developers to regulators—makes it a pivotal contribution to the field of clinical AI.

However, the lack of empirical validation is a limitation. Future work should test the tool's reliability and usability in real-world settings. Nevertheless, PROBAST+AI represents a major step forward in ensuring that AI-based prediction models are held to the same rigorous standards as traditional statistical models.