Preprint
Machine Learning

Receiver operating characteristic (ROC) curves: review of methods with applications in diagnostic medicine

Nancy A. Obuchowski(Cleveland Clinic), Jennifer Bullen(Cleveland Clinic)
March 7, 2018Physics in Medicine and Biology551 citations

551

Citations

35

Influential Citations

Physics in Medicine and Biology

Venue

2018

Year

Abstract

Receiver operating characteristic (ROC) analysis is a tool used to describe the discrimination accuracy of a diagnostic test or prediction model. While sensitivity and specificity are the basic metrics of accuracy, they have many limitations when characterizing test accuracy, particularly when comparing the accuracies of competing tests. In this article we review the basic study design features of ROC studies, illustrate sample size calculations, present statistical methods for measuring and comparing accuracy, and highlight commonly used ROC software. We include descriptions of multi-reader ROC study design and analysis, address frequently seen problems of verification and location bias, discuss clustered data, and provide strategies for testing endpoints in ROC studies. The methods are illustrated with a study of transmission ultrasound for diagnosing breast lesions.

Analysis

Why This Paper Matters

This paper is a foundational review of ROC analysis, a cornerstone for evaluating diagnostic tests and prediction models. In an era of machine learning, where model performance is often summarized by AUC, this paper provides the rigorous statistical framework needed to properly design, analyze, and interpret ROC studies. It addresses practical issues like sample size, bias, and clustered data that are often overlooked in AI applications but critical for clinical validity.

The paper's emphasis on multi-reader studies and bias correction is particularly relevant as AI models are increasingly used as 'readers' in radiology and pathology. Understanding these statistical nuances is essential for AI practitioners to avoid overestimating model performance and to ensure reproducibility in clinical settings.

Technical Contributions

The paper's key contributions include:

  • Study Design Guidance: Detailed discussion of prospective vs. retrospective designs, and the importance of defining the target population and test protocol.
  • Sample Size Calculations: Methods for determining the number of subjects and readers needed to achieve desired power, with formulas and examples.
  • Statistical Methods: Overview of parametric and non-parametric approaches for estimating ROC curves and their AUCs, including methods for comparing correlated ROC curves.
  • Multi-Reader ROC: Explanation of study designs where multiple readers interpret tests, and the corresponding analysis methods (e.g., Obuchowski-Rockette, Dorfman-Berbaum-Metz).
  • Bias and Clustering: Strategies to handle verification bias (when not all subjects receive the gold standard) and location bias (when lesions are not independent), as well as methods for clustered data.

Results

As a review, the paper does not present new empirical results. However, it illustrates the methods using a study of transmission ultrasound for diagnosing breast lesions, demonstrating how to apply the techniques in practice. The paper also highlights commonly used software (e.g., SAS, R packages) for ROC analysis, which is valuable for practitioners.

Significance

This paper has had a significant impact on the diagnostic medicine community, as evidenced by its 551 citations. It serves as a standard reference for designing and analyzing ROC studies, ensuring that conclusions about test accuracy are statistically sound. For AI practitioners, this work underscores the importance of rigorous evaluation beyond simple AUC, including considerations of bias, clustering, and reader variability. As AI models are integrated into clinical workflows, adopting these statistical practices will be crucial for regulatory approval and clinical acceptance.