ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2019
Year
Demonstrates that language models begin to learn various language processing tasks without any explicit supervision.
This paper is a landmark in natural language processing because it demonstrates that a single language model, trained purely on text prediction, can acquire a broad range of linguistic abilities without any explicit supervision. Prior to GPT-2, most NLP systems required task-specific architectures and labeled data. By showing that scaling up model size and data leads to emergent zero-shot capabilities, the authors challenged the field to rethink the role of unsupervised learning.
The paper also sparked widespread debate about the risks of releasing large language models, leading to a staged release strategy. This made GPT-2 a touchstone for discussions on AI safety, ethics, and responsible publication.
GPT-2 fundamentally shifted the NLP landscape by proving that unsupervised pretraining at scale can produce general-purpose linguistic knowledge. It directly inspired subsequent models like GPT-3, ChatGPT, and many others. The paper also forced the community to confront the dual-use nature of AI, leading to new norms around model release and safety evaluation. Its findings continue to influence research on in-context learning, scaling laws, and the capabilities of large language models.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba