ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
… Despite the fact that experimental neural scaling laws have substantially guided empirical … Our theory exhibits a remarkable match with experimentally measured neural scaling laws …
Neural scaling laws have been empirically observed across many domains, showing that model performance improves predictably with compute, data, and parameters. However, a theoretical derivation from first principles has been lacking. This paper addresses that gap by deriving scaling laws from the statistics of natural language itself, offering a principled explanation for why these laws emerge. This is significant because it moves scaling laws from empirical observation to a theoretically grounded phenomenon, which could lead to more reliable predictions and better resource allocation in model development.
The remarkable match between theory and experiment suggests that the underlying assumptions about language statistics capture essential aspects of the data distribution. This could unify various empirical findings and provide a framework for predicting scaling behavior in new settings, potentially saving significant computational resources by informing model design before training.
The paper reports a remarkable match between the derived theory and experimentally measured neural scaling laws. While specific numerical metrics are not provided in the abstract, the qualitative agreement is emphasized. This suggests that the theory accurately captures the functional form of scaling laws, including the power-law behavior observed in practice. The match likely covers various model sizes and data scales, indicating robustness.
This work has the potential to transform how the AI community understands and utilizes scaling laws. By providing a theoretical foundation, it enables more principled predictions of model performance, which can guide decisions on data collection, model architecture, and compute allocation. It also opens avenues for further research into the statistical properties of language and their implications for learning. Ultimately, this could lead to more efficient and effective AI systems, reducing the empirical trial-and-error that currently dominates large-scale model development.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba