Preprint
Machine Learning

Deep learning for healthcare: review, opportunities and challenges

Riccardo Miotto(Icahn School of Medicine at Mount Sinai), Fei Wang(Cornell University), Shuang Wang(University of California San Diego), Xiaoqian Jiang(University of California San Diego), Joel T. Dudley(Icahn School of Medicine at Mount Sinai)
April 5, 2017Briefings in Bioinformatics3,049 citations

3.0k

Citations

51

Influential Citations

Briefings in Bioinformatics

Venue

2017

Year

Abstract

Gaining knowledge and actionable insights from complex, high-dimensional and heterogeneous biomedical data remains a key challenge in transforming health care. Various types of data have been emerging in modern biomedical research, including electronic health records, imaging, -omics, sensor data and text, which are complex, heterogeneous, poorly annotated and generally unstructured. Traditional data mining and statistical learning approaches typically need to first perform feature engineering to obtain effective and more robust features from those data, and then build prediction or clustering models on top of them. There are lots of challenges on both steps in a scenario of complicated data and lacking of sufficient domain knowledge. The latest advances in deep learning technologies provide new effective paradigms to obtain end-to-end learning models from complex data. In this article, we review the recent literature on applying deep learning technologies to advance the health care domain. Based on the analyzed work, we suggest that deep learning approaches could be the vehicle for translating big biomedical data into improved human health. However, we also note limitations and needs for improved methods development and applications, especially in terms of ease-of-understanding for domain experts and citizen scientists. We discuss such challenges and suggest developing holistic and meaningful interpretable architectures to bridge deep learning models and human interpretability.

Analysis

Why This Paper Matters

This 2017 review by Miotto et al. arrived at a pivotal moment when deep learning was revolutionizing fields like computer vision and NLP, but its application to healthcare was still nascent. The paper systematically surveys the landscape of deep learning across diverse biomedical data types—electronic health records, medical imaging, genomics, sensor data, and clinical text—providing a foundational reference for researchers and practitioners. By highlighting both the transformative potential and the critical gaps, particularly around interpretability, it set the agenda for a generation of work in AI for healthcare.

The paper's significance lies in its comprehensive scope and its balanced perspective. It does not merely celebrate deep learning's successes but also candidly discusses the challenges of applying black-box models in high-stakes medical contexts. This nuanced view helped steer the community toward developing more transparent and trustworthy AI systems, a direction that remains central today.

Technical Contributions

  • End-to-end learning paradigm: The paper emphasizes how deep learning eliminates the need for manual feature engineering, which is especially valuable for high-dimensional, heterogeneous biomedical data.
  • Categorization by data modality: It organizes the literature by data type (EHR, imaging, -omics, sensor, text), making it easy for practitioners to find relevant approaches.
  • Identification of key challenges: The review explicitly calls out the lack of interpretability as a major barrier, advocating for holistic, meaningful architectures that bridge model outputs and human understanding.
  • Future directions: It suggests developing methods that are not only accurate but also understandable by domain experts and citizen scientists, laying groundwork for explainable AI in healthcare.

Results

The paper does not present new experimental results but synthesizes findings from over 100 studies. It reports that deep learning models have achieved state-of-the-art performance in tasks such as disease prediction from EHRs, lesion detection in medical images, and genomic variant classification. However, it notes that most studies were retrospective and lacked prospective clinical validation. The review concludes that while deep learning holds great promise, significant work remains to make models interpretable, robust, and clinically deployable.

Significance

This review has been cited over 3,000 times, reflecting its outsized influence on the field. It helped legitimize deep learning as a serious tool for healthcare AI and directed attention to the critical issue of interpretability. The paper's call for interpretable architectures has inspired numerous subsequent works on explainable AI, attention mechanisms, and model-agnostic explanation methods. Its balanced perspective continues to inform how researchers and practitioners approach the translation of AI into clinical practice, emphasizing that technical performance alone is insufficient without trust and understanding.