ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2021
Year
… Though foundation models are based on standard deep … widespread deployment of foundation models, we currently lack … of the critical research on foundation models will require deep …
This paper is a seminal work that introduced the term "foundation models" to the AI community, capturing the paradigm shift towards large-scale pre-trained models like BERT, GPT-3, and CLIP. It matters because it systematically articulates both the immense potential and the profound risks of these models, which have become central to modern AI. The paper serves as a wake-up call for researchers and practitioners to consider not just performance gains but also ethical, environmental, and societal consequences.
The significance lies in its balanced perspective: while foundation models enable unprecedented capabilities in natural language processing, computer vision, and multimodal tasks, they also inherit and amplify biases from training data, pose challenges for interpretability, and require massive computational resources. This dual framing has shaped subsequent research agendas and policy discussions.
As a position paper, it does not present experimental results. Instead, it synthesizes observations from existing models (e.g., GPT-3, BERT) to illustrate points. For example, it notes that GPT-3 exhibits biases related to race and gender, and that training a single large model can emit as much carbon as several cars over their lifetimes. These qualitative findings have been corroborated by later empirical studies.
The paper has had a lasting impact on the AI field by coining a term that unified research on large pre-trained models. It has been cited extensively and has influenced both academic research and industry practices. It prompted initiatives like the Stanford Center for Research on Foundation Models (CRFM) and shaped discussions around responsible AI development. The paper's balanced view encourages practitioners to innovate while remaining vigilant about the broader implications of their work.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba