Journal Article
Machine Learning

Artificial intelligence for diagnosis and Gleason grading of prostate cancer: the PANDA challenge

Wouter Bulten, Kimmo Kartasalo, Po-Hsuan Cameron Chen, Peter Ström, Hans Pinckaers, Kunal Nagpal, Yuannan Cai, David F. Steiner, Hester van Boven, Robert Vink, Christina Hulsbergen-van de Kaa, Jeroen van der Laak, Mahul B. Amin, Andrew J. Evans, Theodorus van der Kwast, Robert Allan, Peter A. Humphrey, Henrik Grönberg, Hemamali Samaratunga, Brett Delahunt, Toyonori Tsuzuki, Tomi Häkkinen, Lars Egevad, Maggie Demkin, Sohier Dane, Fraser Tan, Masi Valkonen, Greg S. Corrado, Lily Peng, Craig H. Mermel, Pekka Ruusuvuori, Geert Litjens, Martin Eklund, , Américo Brilhante, Aslı Çakır, Xavier Farré, Katerina Geronatsiou, Vincent Molinié, Guilherme Pereira, Paromita Roy, Günter Saile, Paulo G. O. Salles, Ewout Schaafsma, Joëlle Tschui, Jorge Billoch-Lima, Emíio M. Pereira, Ming Zhou, Shujun He, Sejun Song, Qing Sun, Hiroshi Yoshihara, Taiki Yamaguchi, Kosaku Ono, Tao Shen, Jianyi Ji, Arnaud Roussel, Kairong Zhou, Tianrui Chai, Nina Weng, Dmitry Grechka, Maxim V. Shugaev, Raphael Kiminya, Vassili Kovalev, Dmitry Voynov, Valery Malyshev, Elizabeth Lapo, Manuel Campos, Noriaki Ota, Shinsuke Yamaoka, Yusuke Fujimoto, Kentaro Yoshioka, Joni Juvonen, Mikko Tukiainen, Antti Karlsson, Rui Guo, Chia-Lun Hsieh, Igor Zubarev, Habib S. T. Bukhar, Wenyuan Li, Jiayun Li, William Speier, Corey Arnold, Kyungdoc Kim, Byeonguk Bae, Yeong Won Kim, Hong-Seok Lee, Jeonghyuk Park
January 1, 2022Nature Medicine594 citations

594

Citations

38

Influential Citations

Nature Medicine

Venue

2022

Year

Abstract

Abstract Artificial intelligence (AI) has shown promise for diagnosing prostate cancer in biopsies. However, results have been limited to individual studies, lacking validation in multinational settings. Competitions have been shown to be accelerators for medical imaging innovations, but their impact is hindered by lack of reproducibility and independent validation. With this in mind, we organized the PANDA challenge—the largest histopathology competition to date, joined by 1,290 developers—to catalyze development of reproducible AI algorithms for Gleason grading using 10,616 digitized prostate biopsies. We validated that a diverse set of submitted algorithms reached pathologist-level performance on independent cross-continental cohorts, fully blinded to the algorithm developers. On United States and European external validation sets, the algorithms achieved agreements of 0.862 (quadratically weighted κ, 95% confidence interval (CI), 0.840–0.884) and 0.868 (95% CI, 0.835–0.900) with expert uropathologists. Successful generalization across different patient populations, laboratories and reference standards, achieved by a variety of algorithmic approaches, warrants evaluating AI-based Gleason grading in prospective clinical trials.

Analysis

Why This Paper Matters

The PANDA challenge represents a landmark effort in medical AI by addressing two critical gaps: reproducibility and multinational validation. Previous studies on AI for prostate cancer diagnosis were limited to single-institution datasets, making it unclear whether algorithms would generalize across different patient populations, laboratories, and staining protocols. By organizing the largest histopathology competition to date with 1,290 participants and 10,616 biopsies, the authors created a rigorous benchmark that forced algorithmic diversity and independent validation. The fact that multiple distinct approaches achieved pathologist-level performance on fully blinded cross-continental cohorts is a strong signal that AI-based Gleason grading is ready for prospective clinical trials.

Technical Contributions

  • Massive-scale competition infrastructure: The PANDA challenge provided a standardized dataset of 10,616 digitized biopsies from multiple sources, with clear evaluation metrics (quadratically weighted κ) and fully blinded external validation on US and European cohorts.
  • Algorithmic diversity: The competition attracted a wide range of approaches, from deep learning ensembles to hybrid models, all achieving high performance, demonstrating that the task is solvable by multiple methods.
  • Reproducibility focus: Unlike many prior competitions, PANDA emphasized reproducibility by requiring code submission and providing a public leaderboard with held-out test sets.
  • Generalization across domains: Algorithms were validated on independent cohorts from different continents, with different staining protocols and reference standards, showing robust generalization.

Results

On the US external validation set, algorithms achieved a quadratically weighted κ of 0.862 (95% CI 0.840–0.884) with expert uropathologists. On the European external validation set, the agreement was 0.868 (95% CI 0.835–0.900). These metrics are comparable to inter-pathologist agreement levels, indicating that the AI systems performed at a clinically acceptable level. The top-performing algorithms maintained high accuracy across different Gleason grade groups, with particularly strong performance on the clinically critical distinction between Gleason 3+4 and 4+3.

Significance

The PANDA challenge sets a new standard for medical imaging competitions by prioritizing reproducibility and multinational validation. Its success demonstrates that AI can match expert pathologists in a complex diagnostic task, paving the way for prospective clinical trials. The open-source code and public leaderboard provide a lasting resource for the community, accelerating further research into AI-assisted pathology. This work also highlights the value of large-scale competitions in driving algorithmic innovation and building trust in AI systems for clinical deployment.