ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
1.8k
Citations
80
Influential Citations
Electronics
Venue
2019
Year
Machine learning systems are becoming increasingly ubiquitous. These systems’s adoption has been expanding, accelerating the shift towards a more algorithmic society, meaning that algorithmically informed decisions have greater potential for significant social impact. However, most of these accurate decision support systems remain complex black boxes, meaning their internal logic and inner workings are hidden to the user and even experts cannot fully understand the rationale behind their predictions. Moreover, new regulations and highly regulated domains have made the audit and verifiability of decisions mandatory, increasing the demand for the ability to question, understand, and trust machine learning systems, for which interpretability is indispensable. The research community has recognized this interpretability problem and focused on developing both interpretable models and explanation methods over the past few years. However, the emergence of these methods shows there is no consensus on how to assess the explanation quality. Which are the most suitable metrics to assess the quality of an explanation? The aim of this article is to provide a review of the current state of the research field on machine learning interpretability while focusing on the societal impact and on the developed methods and metrics. Furthermore, a complete literature review is presented in order to identify future directions of work on this field.
This survey addresses a critical challenge in modern machine learning: the opacity of complex models. As AI systems increasingly influence high-stakes decisions in healthcare, finance, and criminal justice, the inability to understand their reasoning poses risks of bias, unfairness, and lack of accountability. The paper highlights that even experts cannot fully explain black-box models, making interpretability essential for trust and regulatory compliance (e.g., GDPR). By focusing on both methods and metrics, it provides a structured overview of the field, helping practitioners navigate the growing landscape of explanation techniques.
The paper categorizes interpretability methods into two main types: intrinsic (interpretable models like linear regression, decision trees) and post-hoc (explanations for black-box models, such as LIME, SHAP, and feature importance). It also reviews metrics for evaluating explanation quality, including fidelity, comprehensibility, and stability. Key innovations include:
The survey does not present new experimental results but synthesizes findings from over 100 references. It notes that while many explanation methods exist, there is no agreement on how to measure their quality. For example, fidelity (how well an explanation matches the model) and comprehensibility (how easily humans understand it) are often in conflict. The paper also observes that most methods are evaluated on toy datasets or specific domains, limiting generalizability.
This paper has become a highly cited reference (1781 citations) in the interpretability community. It underscores the societal need for transparent AI and provides a roadmap for future work, such as developing user-centric metrics and integrating interpretability into the model development lifecycle. For practitioners, it serves as a starting point for selecting appropriate explanation methods and understanding their limitations.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba