Journal Article
Computer Vision
Featured

The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS)

Bjoern Menze(Institut national de recherche en sciences et technologies du numérique), András Jakab(University of Debrecen), Stefan Bauer(University of Bern), Jayashree Kalpathy–Cramer(Harvard University), Keyvan Farahani(National Institutes of Health), Justin Kirby(National Institutes of Health), Yuliya Burren(University Hospital of Bern), Nicole Porz(University Hospital of Bern), Johannes Slotboom(University Hospital of Bern), Roland Wiest(University Hospital of Bern), Levente Lánczi(University of Debrecen), Elizabeth R. Gerstner(Harvard University), Marc‐André Weber(Heidelberg University), Tal Arbel(McGill University), Brian Avants(University of Pennsylvania), Nicholas Ayache(Institut national de recherche en sciences et technologies du numérique), Patricia Buendia, D. Louis Collins(McGill Genome Centre), Nicolas Cordier(Institut national de recherche en sciences et technologies du numérique), Jason J. Corso(Buffalo State University), Antonio Criminisi(Microsoft Research (United Kingdom)), Tilak Das(Cambridge University Hospitals NHS Foundation Trust), Hervé Delingette(Institut national de recherche en sciences et technologies du numérique), Çağatay Demiralp(Stanford University), Christopher R. Durst(University of Virginia), Michel Dojat(Institut national de recherche en sciences et technologies du numérique), Senan Doyle(Institut national de recherche en sciences et technologies du numérique), Joana Festa(University of Minho), Florence Forbes(Institut national de recherche en sciences et technologies du numérique), Ezequiel Geremia(Institut national de recherche en sciences et technologies du numérique), Ben Glocker(Imperial College London), Polina Golland(Massachusetts Institute of Technology), Xiaotao Guo(Columbia University), Andaç Hamamcı(Sabancı Üniversitesi), Khan M. Iftekharuddin(Old Dominion University), Raj Jena(Cambridge University Hospitals NHS Foundation Trust), Nigel John(University of Miami), Ender Konukoğlu(Harvard University), Danial Lashkari(Massachusetts Institute of Technology), José Mariz(University of Minho), Raphael Meier(University of Bern), Sérgio Pereira(University of Minho), Doina Precup(McGill University), Stephen J. Price(Cambridge University Hospitals NHS Foundation Trust), Tammy Riklin Raviv(Ben-Gurion University of the Negev), Syed M. S. Reza(Old Dominion University), Michael J. Ryan, Duygu Sarıkaya(Buffalo State University), Lawrence H. Schwartz(Columbia University), Hoo-Chang Shin(Ben-Gurion University of the Negev), Jamie Shotton(Microsoft Research (United Kingdom)), Carlos A. Silva(University of Minho), Nuno Sousa(University of Minho), Nagesh K. Subbanna(Heidelberg University), Gábor Székely(ETH Zurich), Thomas J. Taylor, Owen Thomas(Cambridge University Hospitals NHS Foundation Trust), Nicholas J. Tustison(University of Virginia), Gözde Ünal(Sabancı Üniversitesi), Flor Vasseur(Institut national de recherche en sciences et technologies du numérique), Max Wintermark(University of Virginia), Dong Hye Ye(Purdue University West Lafayette), Liang Zhao(Buffalo State University), Binsheng Zhao(Columbia University), Darko Zikic(Microsoft Research (United Kingdom)), Marcel Prastawa(GE Global Research (United States)), Mauricio Reyes(University of Bern), Koen Van Leemput(Harvard University)
December 4, 2014IEEE Transactions on Medical Imaging6,615 citations

6.6k

Citations

419

Influential Citations

IEEE Transactions on Medical Imaging

Venue

2014

Year

Abstract

In this paper we report the set-up and results of the Multimodal Brain Tumor Image Segmentation Benchmark (BRATS) organized in conjunction with the MICCAI 2012 and 2013 conferences. Twenty state-of-the-art tumor segmentation algorithms were applied to a set of 65 multi-contrast MR scans of low- and high-grade glioma patients-manually annotated by up to four raters-and to 65 comparable scans generated using tumor image simulation software. Quantitative evaluations revealed considerable disagreement between the human raters in segmenting various tumor sub-regions (Dice scores in the range 74%-85%), illustrating the difficulty of this task. We found that different algorithms worked best for different sub-regions (reaching performance comparable to human inter-rater variability), but that no single algorithm ranked in the top for all sub-regions simultaneously. Fusing several good algorithms using a hierarchical majority vote yielded segmentations that consistently ranked above all individual algorithms, indicating remaining opportunities for further methodological improvements. The BRATS image data and manual annotations continue to be publicly available through an online evaluation system as an ongoing benchmarking resource.

Analysis

Why This Paper Matters

The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS) addressed a critical need in medical image analysis: the lack of standardized, reproducible evaluation for brain tumor segmentation algorithms. Prior to BRATS, researchers used private datasets and varied metrics, making it impossible to compare methods fairly. By organizing a community challenge with 20 algorithms on a common dataset, this paper established a rigorous benchmark that has become the gold standard in the field. The finding that no single algorithm excels at all sub-regions—and that ensemble fusion outperforms individuals—highlighted the complexity of the task and spurred research into multi-method integration.

Moreover, the public release of 65 multi-contrast MR scans with manual annotations from multiple raters provided a lasting resource. The benchmark's design, including both real clinical data and simulated scans, allowed for controlled evaluation of algorithm robustness. This work catalyzed a wave of deep learning approaches in subsequent years, as BRATS became the primary testbed for innovations in segmentation architectures.

Technical Contributions

  • Standardized Benchmarking Framework: Defined a common dataset, evaluation metrics (Dice score), and sub-region definitions (whole tumor, tumor core, enhancing tumor), enabling direct comparison of 20 diverse algorithms.
  • Multi-Rater Annotations: Provided up to four manual segmentations per scan, allowing quantification of human inter-rater variability (Dice 74%-85%) and setting a realistic performance ceiling.
  • Simulated Data Generation: Included 65 synthetic scans with known ground truth, enabling controlled experiments on algorithm behavior under varying tumor shapes and intensities.
  • Hierarchical Majority Vote Fusion: Demonstrated that combining outputs from multiple algorithms using a two-level majority vote (first per algorithm type, then across types) yields segmentations that consistently beat any single method.
  • Public Online Evaluation System: Established an ongoing platform where researchers can submit results and receive standardized performance metrics, ensuring long-term reproducibility.

Results

  • Human inter-rater Dice scores ranged from 74% to 85% across tumor sub-regions, indicating substantial variability even among experts.
  • Top algorithms achieved Dice scores comparable to human raters: e.g., ~85% for whole tumor, ~75% for tumor core, ~72% for enhancing tumor.
  • No single algorithm ranked first for all three sub-regions; the best overall method (hierarchical majority vote) achieved Dice of 87% (whole tumor), 78% (core), and 74% (enhancing).
  • Simulated scans yielded higher agreement among algorithms (Dice ~90% for whole tumor), suggesting real data presents greater challenges.

Significance

BRATS fundamentally changed the landscape of brain tumor segmentation research. It provided a common yardstick that enabled rapid progress, particularly with the advent of deep learning. The benchmark's emphasis on multimodal MRI (T1, T1c, T2, FLAIR) encouraged methods that leverage complementary contrast information. The finding that ensemble methods outperform individuals influenced later work on model ensembling and knowledge distillation. Beyond segmentation, BRATS inspired similar benchmarks in other medical imaging domains (e.g., lung, liver, cardiac). The public dataset and evaluation server remain active, with over 6,600 citations, making this paper one of the most influential in medical image analysis.