Conference Paper
Large Language Models

GPT-3: Its Nature, Scope, Limits, and Consequences

Luciano Floridi(Turing Institute), Massimo Chiriatti
November 1, 2020Minds and Machines2,511 citations

2.5k

Citations

100

Influential Citations

Minds and Machines

Venue

2020

Year

Abstract

Abstract In this commentary, we discuss the nature of reversible and irreversible questions, that is, questions that may enable one to identify the nature of the source of their answers. We then introduce GPT-3, a third-generation, autoregressive language model that uses deep learning to produce human-like texts, and use the previous distinction to analyse it. We expand the analysis to present three tests based on mathematical, semantic (that is, the Turing Test), and ethical questions and show that GPT-3 is not designed to pass any of them. This is a reminder that GPT-3 does not do what it is not supposed to do, and that any interpretation of GPT-3 as the beginning of the emergence of a general form of artificial intelligence is merely uninformed science fiction. We conclude by outlining some of the significant consequences of the industrialisation of automatic and cheap production of good, semantic artefacts.

Analysis

Why This Paper Matters

This paper provides a timely and critical perspective on GPT-3, one of the most influential large language models. At a time when media and some researchers were speculating about GPT-3 as a step toward artificial general intelligence (AGI), Floridi and Chiriatti offer a sobering analysis grounded in philosophical distinctions. Their work matters because it reframes the conversation around what GPT-3 actually does—producing human-like text without understanding—and what it does not do. This is crucial for AI practitioners who need to set realistic expectations for deployment and avoid overinterpreting model outputs.

The paper also introduces a novel conceptual tool: the distinction between reversible and irreversible questions. This framework helps determine whether an answer comes from a genuine understanding or mere pattern matching. For practitioners, this is a practical heuristic for evaluating when to trust model outputs.

Technical Contributions

  • Reversible vs. Irreversible Questions: A question is reversible if the answer reveals the nature of the source (e.g., asking a human vs. a machine). GPT-3's answers are often indistinguishable from human ones, making questions irreversible.
  • Three Tests: The authors propose three tests that GPT-3 fails:
    • Mathematical: GPT-3 cannot reliably solve novel math problems requiring reasoning.
    • Semantic (Turing Test): GPT-3 can mimic conversation but lacks genuine understanding.
    • Ethical: GPT-3 cannot make consistent ethical judgments.
  • Industrialization of Semantic Artefacts: The paper highlights how GPT-3 enables cheap, automatic production of text, which has profound implications for misinformation, content creation, and labor.

Results

The paper does not present new experimental results but synthesizes existing observations. Key claims: GPT-3 fails all three proposed tests, confirming it is not designed for general intelligence. The authors emphasize that GPT-3's success in generating coherent text does not imply understanding or reasoning. No quantitative metrics are provided; the analysis is conceptual.

Significance

This paper has had a significant impact on the AI community, accumulating over 2500 citations. It serves as a cautionary note against AGI hype and provides a framework for evaluating language models critically. For practitioners, it underscores the importance of task-specific evaluation and the dangers of anthropomorphizing AI outputs. The ethical implications of mass-produced semantic artefacts remain highly relevant as GPT-3 and its successors are deployed in real-world applications.