ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
6.6k
Citations
419
Influential Citations
IEEE Transactions on Medical Imaging
Venue
2014
Year
In this paper we report the set-up and results of the Multimodal Brain Tumor Image Segmentation Benchmark (BRATS) organized in conjunction with the MICCAI 2012 and 2013 conferences. Twenty state-of-the-art tumor segmentation algorithms were applied to a set of 65 multi-contrast MR scans of low- and high-grade glioma patients-manually annotated by up to four raters-and to 65 comparable scans generated using tumor image simulation software. Quantitative evaluations revealed considerable disagreement between the human raters in segmenting various tumor sub-regions (Dice scores in the range 74%-85%), illustrating the difficulty of this task. We found that different algorithms worked best for different sub-regions (reaching performance comparable to human inter-rater variability), but that no single algorithm ranked in the top for all sub-regions simultaneously. Fusing several good algorithms using a hierarchical majority vote yielded segmentations that consistently ranked above all individual algorithms, indicating remaining opportunities for further methodological improvements. The BRATS image data and manual annotations continue to be publicly available through an online evaluation system as an ongoing benchmarking resource.
The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS) addressed a critical need in medical image analysis: the lack of standardized, reproducible evaluation for brain tumor segmentation algorithms. Prior to BRATS, researchers used private datasets and varied metrics, making it impossible to compare methods fairly. By organizing a community challenge with 20 algorithms on a common dataset, this paper established a rigorous benchmark that has become the gold standard in the field. The finding that no single algorithm excels at all sub-regions—and that ensemble fusion outperforms individuals—highlighted the complexity of the task and spurred research into multi-method integration.
Moreover, the public release of 65 multi-contrast MR scans with manual annotations from multiple raters provided a lasting resource. The benchmark's design, including both real clinical data and simulated scans, allowed for controlled evaluation of algorithm robustness. This work catalyzed a wave of deep learning approaches in subsequent years, as BRATS became the primary testbed for innovations in segmentation architectures.
BRATS fundamentally changed the landscape of brain tumor segmentation research. It provided a common yardstick that enabled rapid progress, particularly with the advent of deep learning. The benchmark's emphasis on multimodal MRI (T1, T1c, T2, FLAIR) encouraged methods that leverage complementary contrast information. The finding that ensemble methods outperform individuals influenced later work on model ensembling and knowledge distillation. Beyond segmentation, BRATS inspired similar benchmarks in other medical imaging domains (e.g., lung, liver, cardiac). The public dataset and evaluation server remain active, with over 6,600 citations, making this paper one of the most influential in medical image analysis.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba