ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
594
Citations
38
Influential Citations
Nature Medicine
Venue
2022
Year
Abstract Artificial intelligence (AI) has shown promise for diagnosing prostate cancer in biopsies. However, results have been limited to individual studies, lacking validation in multinational settings. Competitions have been shown to be accelerators for medical imaging innovations, but their impact is hindered by lack of reproducibility and independent validation. With this in mind, we organized the PANDA challenge—the largest histopathology competition to date, joined by 1,290 developers—to catalyze development of reproducible AI algorithms for Gleason grading using 10,616 digitized prostate biopsies. We validated that a diverse set of submitted algorithms reached pathologist-level performance on independent cross-continental cohorts, fully blinded to the algorithm developers. On United States and European external validation sets, the algorithms achieved agreements of 0.862 (quadratically weighted κ, 95% confidence interval (CI), 0.840–0.884) and 0.868 (95% CI, 0.835–0.900) with expert uropathologists. Successful generalization across different patient populations, laboratories and reference standards, achieved by a variety of algorithmic approaches, warrants evaluating AI-based Gleason grading in prospective clinical trials.
The PANDA challenge represents a landmark effort in medical AI by addressing two critical gaps: reproducibility and multinational validation. Previous studies on AI for prostate cancer diagnosis were limited to single-institution datasets, making it unclear whether algorithms would generalize across different patient populations, laboratories, and staining protocols. By organizing the largest histopathology competition to date with 1,290 participants and 10,616 biopsies, the authors created a rigorous benchmark that forced algorithmic diversity and independent validation. The fact that multiple distinct approaches achieved pathologist-level performance on fully blinded cross-continental cohorts is a strong signal that AI-based Gleason grading is ready for prospective clinical trials.
On the US external validation set, algorithms achieved a quadratically weighted κ of 0.862 (95% CI 0.840–0.884) with expert uropathologists. On the European external validation set, the agreement was 0.868 (95% CI 0.835–0.900). These metrics are comparable to inter-pathologist agreement levels, indicating that the AI systems performed at a clinically acceptable level. The top-performing algorithms maintained high accuracy across different Gleason grade groups, with particularly strong performance on the clinically critical distinction between Gleason 3+4 and 4+3.
The PANDA challenge sets a new standard for medical imaging competitions by prioritizing reproducibility and multinational validation. Its success demonstrates that AI can match expert pathologists in a complex diagnostic task, paving the way for prospective clinical trials. The open-source code and public leaderboard provide a lasting resource for the community, accelerating further research into AI-assisted pathology. This work also highlights the value of large-scale competitions in driving algorithmic innovation and building trust in AI systems for clinical deployment.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba