ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2024
Year
… This work made theoretical predictions of neural scaling laws. It could be built on by training neural networks on toy datasets to test scaling regimes with c > 2; searching for percolation …
Neural scaling laws describe how model performance improves with compute, data, and parameters. While empirical scaling laws have been widely observed, a theoretical foundation linking them to the data distribution has been lacking. This paper addresses that gap by deriving scaling laws from the data distribution, offering a principled explanation for observed scaling behavior. This is significant because it moves scaling laws from empirical observation to a more fundamental theory, which could lead to more reliable predictions and better resource allocation in training large models.
The paper also introduces the possibility of scaling regimes with exponent c > 2, which are not commonly observed in practice. This suggests that the scaling exponent may depend on the structure of the data, and that different data distributions could lead to different scaling behaviors. This has implications for how we think about data curation and model architecture, as it implies that the optimal scaling strategy might vary depending on the data.
The abstract does not include specific numerical results or comparisons. The main result is the theoretical prediction of scaling laws with c > 2 under certain data distribution conditions. The paper proposes to test these predictions on toy datasets, but no experimental outcomes are reported in the abstract. Therefore, the concrete metrics are not available from the abstract alone.
If validated, this work could provide a deeper understanding of why scaling laws emerge and how they depend on data properties. This could lead to more accurate predictions of model performance, helping researchers and practitioners choose optimal model sizes and data budgets. The connection to percolation is intriguing and might open new interdisciplinary avenues, linking machine learning to statistical physics. Overall, this paper contributes to the theoretical foundations of deep learning and could influence future research on scaling and data-centric AI.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba