ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
12k
Citations
504
Influential Citations
Sociological Methods & Research
Venue
2004
Year
The model selection literature has been generally poor at reflecting the deep foundations of the Akaike information criterion (AIC) and at making appropriate comparisons to the Bayesian information criterion (BIC). There is a clear philosophy, a sound criterion based in information theory, and a rigorous statistical foundation for AIC. AIC can be justified as Bayesian using a “savvy” prior on models that is a function of sample size and the number of model parameters. Furthermore, BIC can be derived as a non-Bayesian result. Therefore, arguments about using AIC versus BIC for model selection cannot be from a Bayes versus frequentist perspective. The philosophical context of what is assumed about reality, approximating models, and the intent of model-based inference should determine whether AIC or BIC is used. Various facets of such multimodel inference are presented here, particularly methods of model averaging.
This paper is a landmark in the model selection literature, addressing widespread misunderstandings about the Akaike information criterion (AIC) and Bayesian information criterion (BIC). Burnham and Anderson clarify that the choice between AIC and BIC is not a matter of Bayesian versus frequentist philosophy, but rather depends on the underlying assumptions about the true data-generating process and the goals of inference. By grounding AIC in information theory and showing its Bayesian justification, the paper elevates AIC from a heuristic to a principled criterion.
The paper's advocacy for multimodel inference—especially model averaging—has had a transformative impact. Instead of selecting a single 'best' model, practitioners are encouraged to consider a set of plausible models and average their predictions, thereby accounting for model uncertainty. This approach has been widely adopted in ecology, economics, and other fields where model uncertainty is substantial.
The paper does not present experimental results or quantitative comparisons. Its contributions are entirely theoretical and conceptual. However, its impact is measurable: with over 11,670 citations, it is one of the most cited papers in the model selection literature. The ideas have been validated through widespread application in fields such as ecology, biology, economics, and machine learning, where multimodel inference has become standard practice.
Burnham and Anderson's work fundamentally changed how researchers approach model selection. By emphasizing multimodel inference and model averaging, it moved the field away from the risky practice of selecting a single model and then making inferences as if that model were true. The paper's clear philosophical and mathematical arguments have made it a cornerstone reference for anyone dealing with model uncertainty. In the AI and machine learning community, the principles of multimodel inference align with ensemble methods and Bayesian model averaging, influencing modern practices in deep learning and probabilistic modeling.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba