ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
2.5k
Citations
100
Influential Citations
Minds and Machines
Venue
2020
Year
Abstract In this commentary, we discuss the nature of reversible and irreversible questions, that is, questions that may enable one to identify the nature of the source of their answers. We then introduce GPT-3, a third-generation, autoregressive language model that uses deep learning to produce human-like texts, and use the previous distinction to analyse it. We expand the analysis to present three tests based on mathematical, semantic (that is, the Turing Test), and ethical questions and show that GPT-3 is not designed to pass any of them. This is a reminder that GPT-3 does not do what it is not supposed to do, and that any interpretation of GPT-3 as the beginning of the emergence of a general form of artificial intelligence is merely uninformed science fiction. We conclude by outlining some of the significant consequences of the industrialisation of automatic and cheap production of good, semantic artefacts.
This paper provides a timely and critical perspective on GPT-3, one of the most influential large language models. At a time when media and some researchers were speculating about GPT-3 as a step toward artificial general intelligence (AGI), Floridi and Chiriatti offer a sobering analysis grounded in philosophical distinctions. Their work matters because it reframes the conversation around what GPT-3 actually does—producing human-like text without understanding—and what it does not do. This is crucial for AI practitioners who need to set realistic expectations for deployment and avoid overinterpreting model outputs.
The paper also introduces a novel conceptual tool: the distinction between reversible and irreversible questions. This framework helps determine whether an answer comes from a genuine understanding or mere pattern matching. For practitioners, this is a practical heuristic for evaluating when to trust model outputs.
The paper does not present new experimental results but synthesizes existing observations. Key claims: GPT-3 fails all three proposed tests, confirming it is not designed for general intelligence. The authors emphasize that GPT-3's success in generating coherent text does not imply understanding or reasoning. No quantitative metrics are provided; the analysis is conceptual.
This paper has had a significant impact on the AI community, accumulating over 2500 citations. It serves as a cautionary note against AGI hype and provides a framework for evaluating language models critically. For practitioners, it underscores the importance of task-specific evaluation and the dangers of anthropomorphizing AI outputs. The ethical implications of mass-produced semantic artefacts remain highly relevant as GPT-3 and its successors are deployed in real-world applications.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba