Preprint
Machine Learning

There Will Be a Scientific Theory of Deep Learning

James B. Simon, D. Kunin, Alexander Atanasov, Enric Boix-Adserà, Blake Bordelon, Jeremy Cohen, Nikhil Ghosh, Florentin Guth, Arthur Jacot, M. Kamb, Dhruva Karkada, Eric J. Michaud, Berkan Ottlik, Joseph Turnbull
April 23, 2026arXiv.org15 citations

15

Citations

0

Influential Citations

arXiv.org

Venue

2026

Year

Abstract

In this paper, we make the case that a scientific theory of deep learning is emerging. By this we mean a theory which characterizes important properties and statistics of the training process, hidden representations, final weights, and performance of neural networks. We pull together major strands of ongoing research in deep learning theory and identify five growing bodies of work that point toward such a theory: (a) solvable idealized settings that provide intuition for learning dynamics in realistic systems; (b) tractable limits that reveal insights into fundamental learning phenomena; (c) simple mathematical laws that capture important macroscopic observables; (d) theories of hyperparameters that disentangle them from the rest of the training process, leaving simpler systems behind; and (e) universal behaviors shared across systems and settings which clarify which phenomena call for explanation. Taken together, these bodies of work share certain broad traits: they are concerned with the dynamics of the training process; they primarily seek to describe coarse aggregate statistics; and they emphasize falsifiable quantitative predictions. We argue that the emerging theory is best thought of as a mechanics of the learning process, and suggest the name learning mechanics. We discuss the relationship between this mechanics perspective and other approaches for building a theory of deep learning, including the statistical and information-theoretic perspectives. In particular, we anticipate a symbiotic relationship between learning mechanics and mechanistic interpretability. We also review and address common arguments that fundamental theory will not be possible or is not important. We conclude with a portrait of important open directions in learning mechanics and advice for beginners. We host further introductory materials, perspectives, and open questions at learningmechanics.pub.

Analysis

Why This Paper Matters

This paper is significant because it addresses a long-standing critique of deep learning: that it lacks a rigorous scientific foundation. By synthesizing five distinct but converging research directions, the authors make a compelling case that a theory of deep learning is not only possible but already emerging. This matters for practitioners because a scientific theory promises to move deep learning from an art guided by intuition and heuristics to a principled engineering discipline. For researchers, the paper provides a roadmap and a shared vocabulary—'learning mechanics'—that could accelerate progress by clarifying which questions are most important and which phenomena are universal.

The paper also tackles head-on the skepticism that deep learning is too complex for fundamental theory. By pointing to successful examples in physics (e.g., thermodynamics emerging from statistical mechanics) and by emphasizing falsifiable quantitative predictions, the authors argue that a coarse-grained, statistical theory is both achievable and useful. This is a refreshingly optimistic and constructive perspective that could help attract more researchers to theoretical work.

Technical Contributions

The paper's main technical contribution is its taxonomy of five bodies of work that constitute the emerging theory:

  • Solvable idealized settings: Simplified models (e.g., linear networks, kernel methods) that capture essential dynamics of real networks.
  • Tractable limits: Asymptotic regimes (e.g., infinite width, infinite depth) where analysis becomes feasible while still revealing fundamental phenomena.
  • Simple mathematical laws: Empirical scaling laws and other regularities that describe macroscopic observables like loss, accuracy, and feature learning.
  • Theories of hyperparameters: Frameworks that disentangle hyperparameter effects from the core learning process, simplifying the system under study.
  • Universal behaviors: Phenomena observed across architectures, datasets, and tasks, indicating which aspects of learning are fundamental and which are incidental.

The paper also introduces the concept of 'learning mechanics' as a distinct perspective, contrasting it with statistical learning theory (which focuses on generalization bounds) and information-theoretic approaches (which focus on compression and mutual information). It further discusses a symbiotic relationship with mechanistic interpretability, where learning mechanics provides the 'how' of training and mechanistic interpretability provides the 'what' of learned representations.

Results

As a position paper, there are no new experimental results. The paper's 'results' are its synthesis and argumentation. It cites 15 references, drawing on a broad range of theoretical and empirical work to support its claims. The paper's main deliverable is a conceptual framework that organizes existing knowledge and identifies gaps.

Significance

The broader impact of this paper could be substantial. If the learning mechanics perspective gains traction, it could reshape how deep learning is taught, researched, and applied. It offers a unifying language for theorists and practitioners, potentially reducing fragmentation in the field. The paper's emphasis on falsifiable predictions and coarse-grained observables aligns with the needs of engineering: practitioners care about reliable predictions of training behavior, not just post-hoc explanations. By explicitly addressing common objections and providing a beginner-friendly resource (learningmechanics.pub), the authors lower the barrier to entry for new researchers. This paper is likely to become a touchstone for discussions about the scientific foundations of deep learning.