ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
520
Citations
19
Influential Citations
BMJ
Venue
2025
Year
The Prediction model Risk Of Bias ASsessment Tool (PROBAST) is used to assess the quality, risk of bias, and applicability of prediction models or algorithms and of prediction model/algorithm studies. Since PROBAST’s introduction in 2019, much progress has been made in the methodology for prediction modelling and in the use of artificial intelligence, including machine learning, techniques. An update to PROBAST-2019 is thus needed. This article describes the development of PROBAST+AI. PROBAST+AI consists of two distinctive parts: model development and model evaluation. For model development, PROBAST+AI users assess quality and applicability using 16 targeted signalling questions. For model evaluation, PROBAST+AI users assess the risk of bias and applicability using 18 targeted signalling questions. Both parts contain four domains: participants and data sources, predictors, outcome, and analysis. Applicability of the prediction model is rated for the participants and data sources, predictors, and outcome domains. PROBAST+AI may replace the original PROBAST tool and allows all key stakeholders (eg, model developers, AI companies, researchers, editors, reviewers, healthcare professionals, guideline developers, and policy organisations) to examine the quality, risk of bias, and applicability of any type of prediction model in the healthcare sector, irrespective of whether regression modelling or AI techniques are used.
PROBAST+AI addresses a critical gap in the evaluation of clinical prediction models. The original PROBAST, released in 2019, was designed primarily for regression-based models. With the rapid adoption of machine learning and deep learning in healthcare, there is an urgent need for a tool that can assess the risk of bias and applicability of AI-based models. This paper responds to that need by providing an updated framework that is explicitly inclusive of AI techniques.
The significance of this work lies in its potential to become a standard for regulatory and clinical evaluation. As AI models increasingly influence patient care, the ability to systematically assess their quality is essential. PROBAST+AI offers a structured approach that can be used by model developers, journals, and health technology assessment bodies, thereby promoting reproducibility and accountability.
The paper does not present quantitative results or validation metrics. Instead, it describes the development and structure of PROBAST+AI. The abstract mentions that the tool consists of 16 and 18 signaling questions for development and evaluation, respectively, but no data on inter-rater reliability or usability are provided. This is a methodological paper, so the 'results' are the tool itself and its rationale.
PROBAST+AI has the potential to become the de facto standard for assessing prediction models in healthcare, regardless of whether they are based on regression or AI. By providing a common framework, it can facilitate comparisons across studies, improve the quality of published research, and support evidence-based adoption of AI in clinical practice. The tool's broad applicability to stakeholders—from developers to regulators—makes it a pivotal contribution to the field of clinical AI.
However, the lack of empirical validation is a limitation. Future work should test the tool's reliability and usability in real-world settings. Nevertheless, PROBAST+AI represents a major step forward in ensuring that AI-based prediction models are held to the same rigorous standards as traditional statistical models.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba