ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
722
Citations
24
Influential Citations
Durham Research Online (Durham University)
Venue
2021
Year
Deep generative models are a class of techniques that train deep neural networks to model the distribution of training samples. Research has fragmented into various interconnected approaches, each of which make trade-offs including run-time, diversity, and architectural restrictions. In particular, this compendium covers energy-based models, variational autoencoders, generative adversarial networks, autoregressive models, normalizing flows, in addition to numerous hybrid approaches. These techniques are compared and contrasted, explaining the premises behind each and how they are interrelated, while reviewing current state-of-the-art advances and implementations.
Deep generative models have become a cornerstone of modern AI, enabling applications from image synthesis to data augmentation. However, the field has fragmented into multiple competing paradigms, each with its own strengths and weaknesses. This paper addresses a critical need for a comprehensive, accessible overview that clarifies the landscape. By systematically comparing VAEs, GANs, normalizing flows, energy-based models, and autoregressive models, it provides a roadmap for both newcomers and experienced researchers.
The paper's significance lies in its emphasis on the interconnections between these approaches. Rather than presenting them as isolated techniques, it highlights how they relate to each other, often through hybrid models. This holistic perspective is essential for advancing the field, as it encourages cross-pollination of ideas and helps researchers identify promising directions for innovation.
As a review paper, it does not introduce new experimental results. Instead, it synthesizes existing knowledge, offering qualitative comparisons. For example, it notes that GANs typically produce sharp samples but suffer from mode collapse, while VAEs offer stable training but produce blurrier outputs. Normalizing flows provide exact likelihood estimation but are computationally expensive. Autoregressive models achieve high-quality samples but are slow to generate. Energy-based models are flexible but challenging to train. These insights are drawn from the literature, not from new experiments.
The paper's impact is primarily educational and organizational. It serves as a comprehensive reference that can help researchers quickly understand the strengths and limitations of each approach, potentially guiding them toward more effective model design. By highlighting the interconnections between methods, it may inspire novel hybrid models that leverage the best of multiple paradigms. The paper's citation count (722) indicates its widespread use as a foundational resource in the deep generative modeling community. Its publication at Durham University adds to the academic credibility, and its open-access nature ensures broad accessibility. Overall, this review contributes to the maturation of the field by consolidating knowledge and fostering a more integrated understanding of deep generative models.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba