ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
773
Citations
30
Influential Citations
APSIPA Transactions on Signal and Information Processing
Venue
2014
Year
In this invited paper, my overview material on the same topic as presented in the plenary overview session of APSIPA-2011 and the tutorial material presented in the same conference [1] are expanded and updated to include more recent developments in deep learning. The previous and the updated materials cover both theory and applications, and analyze its future directions. The goal of this tutorial survey is to introduce the emerging area of deep learning or hierarchical learning to the APSIPA community. Deep learning refers to a class of machine learning techniques, developed largely since 2006, where many stages of non-linear information processing in hierarchical architectures are exploited for pattern classification and for feature learning. In the more recent literature, it is also connected to representation learning, which involves a hierarchy of features or concepts where higher-level concepts are defined from lower-level ones and where the same lower-level concepts help to define higher-level ones. In this tutorial survey, a brief history of deep learning research is discussed first. Then, a classificatory scheme is developed to analyze and summarize major work reported in the recent deep learning literature. Using this scheme, I provide a taxonomy-oriented survey on the existing deep architectures and algorithms in the literature, and categorize them into three classes: generative, discriminative, and hybrid. Three representative deep architectures – deep autoencoders, deep stacking networks with their generalization to the temporal domain (recurrent networks), and deep neural networks (pretrained with deep belief networks) – one in each of the three classes, are presented in more detail. Next, selected applications of deep learning are reviewed in broad areas of signal and information processing including audio/speech, image/vision, multimodality, language modeling, natural language processing, and information retrieval. Finally, future directions of deep learning are discussed and analyzed.
This tutorial survey, authored by Li Deng, is a seminal reference that introduced the APSIPA community to the emerging field of deep learning. Published in 2014, it captures the foundational developments that shaped modern AI, including deep autoencoders, deep stacking networks, and deep neural networks pretrained with deep belief networks. The paper's timing is crucial—it bridges the period between the 2006 resurgence of deep learning and the deep learning revolution that followed, making it a key resource for understanding the origins of current techniques.
The paper's significance lies in its comprehensive taxonomy, which categorizes deep architectures into generative, discriminative, and hybrid classes. This classification helps researchers navigate the complex landscape of deep learning and understand the relationships between different approaches. By providing a structured overview, the paper enables practitioners to identify suitable architectures for their specific problems and to appreciate the theoretical underpinnings of deep learning.
The paper makes several key technical contributions:
As a survey paper, it does not present new experimental results. Instead, it synthesizes findings from numerous studies, highlighting the superior performance of deep learning methods in various tasks. For example, it notes that deep neural networks pretrained with deep belief networks achieved significant error reductions in speech recognition and image classification compared to shallow models. The paper also emphasizes the importance of feature learning and representation learning, which have become central to modern deep learning success.
The broader impact of this paper is substantial. It has been cited 773 times, indicating its influence on both academic research and practical applications. By providing a clear and accessible introduction to deep learning, it has helped democratize knowledge and encouraged wider adoption of these techniques. The taxonomy and architectural insights remain relevant today, as many modern models (e.g., transformers) can be viewed through the lens of hybrid or discriminative architectures. The paper's emphasis on representation learning foreshadowed the deep learning revolution's focus on learned features, which is now a cornerstone of AI. For practitioners, this survey offers a valuable historical perspective and a foundational understanding that aids in navigating the rapidly evolving deep learning landscape.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba