Journal Article
Computer Vision

A tutorial survey of architectures, algorithms, and applications for deep learning

Li Deng(Microsoft (United States))
January 1, 2014APSIPA Transactions on Signal and Information Processing773 citations

773

Citations

30

Influential Citations

APSIPA Transactions on Signal and Information Processing

Venue

2014

Year

Abstract

In this invited paper, my overview material on the same topic as presented in the plenary overview session of APSIPA-2011 and the tutorial material presented in the same conference [1] are expanded and updated to include more recent developments in deep learning. The previous and the updated materials cover both theory and applications, and analyze its future directions. The goal of this tutorial survey is to introduce the emerging area of deep learning or hierarchical learning to the APSIPA community. Deep learning refers to a class of machine learning techniques, developed largely since 2006, where many stages of non-linear information processing in hierarchical architectures are exploited for pattern classification and for feature learning. In the more recent literature, it is also connected to representation learning, which involves a hierarchy of features or concepts where higher-level concepts are defined from lower-level ones and where the same lower-level concepts help to define higher-level ones. In this tutorial survey, a brief history of deep learning research is discussed first. Then, a classificatory scheme is developed to analyze and summarize major work reported in the recent deep learning literature. Using this scheme, I provide a taxonomy-oriented survey on the existing deep architectures and algorithms in the literature, and categorize them into three classes: generative, discriminative, and hybrid. Three representative deep architectures – deep autoencoders, deep stacking networks with their generalization to the temporal domain (recurrent networks), and deep neural networks (pretrained with deep belief networks) – one in each of the three classes, are presented in more detail. Next, selected applications of deep learning are reviewed in broad areas of signal and information processing including audio/speech, image/vision, multimodality, language modeling, natural language processing, and information retrieval. Finally, future directions of deep learning are discussed and analyzed.

Analysis

Why This Paper Matters

This tutorial survey, authored by Li Deng, is a seminal reference that introduced the APSIPA community to the emerging field of deep learning. Published in 2014, it captures the foundational developments that shaped modern AI, including deep autoencoders, deep stacking networks, and deep neural networks pretrained with deep belief networks. The paper's timing is crucial—it bridges the period between the 2006 resurgence of deep learning and the deep learning revolution that followed, making it a key resource for understanding the origins of current techniques.

The paper's significance lies in its comprehensive taxonomy, which categorizes deep architectures into generative, discriminative, and hybrid classes. This classification helps researchers navigate the complex landscape of deep learning and understand the relationships between different approaches. By providing a structured overview, the paper enables practitioners to identify suitable architectures for their specific problems and to appreciate the theoretical underpinnings of deep learning.

Technical Contributions

The paper makes several key technical contributions:

  • Taxonomy of Deep Architectures: It proposes a three-class scheme—generative, discriminative, and hybrid—to organize deep learning methods, clarifying the distinctions and connections between them.
  • Detailed Analysis of Representative Architectures: It provides in-depth descriptions of three archetypal models: deep autoencoders (generative), deep stacking networks and their recurrent extensions (discriminative), and deep neural networks pretrained with deep belief networks (hybrid).
  • Historical Context: It traces the evolution of deep learning from early neural networks to the 2006 breakthroughs, offering insights into the motivations and challenges that drove the field.
  • Application Survey: It systematically reviews applications across audio/speech, image/vision, multimodality, language modeling, NLP, and information retrieval, demonstrating the versatility of deep learning.
  • Future Directions: It discusses open problems and potential research avenues, such as scaling to larger models and improving unsupervised learning.

Results

As a survey paper, it does not present new experimental results. Instead, it synthesizes findings from numerous studies, highlighting the superior performance of deep learning methods in various tasks. For example, it notes that deep neural networks pretrained with deep belief networks achieved significant error reductions in speech recognition and image classification compared to shallow models. The paper also emphasizes the importance of feature learning and representation learning, which have become central to modern deep learning success.

Significance

The broader impact of this paper is substantial. It has been cited 773 times, indicating its influence on both academic research and practical applications. By providing a clear and accessible introduction to deep learning, it has helped democratize knowledge and encouraged wider adoption of these techniques. The taxonomy and architectural insights remain relevant today, as many modern models (e.g., transformers) can be viewed through the lens of hybrid or discriminative architectures. The paper's emphasis on representation learning foreshadowed the deep learning revolution's focus on learned features, which is now a cornerstone of AI. For practitioners, this survey offers a valuable historical perspective and a foundational understanding that aids in navigating the rapidly evolving deep learning landscape.