ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
839
Citations
41
Influential Citations
Computational Visual Media
Venue
2019
Year
Detecting and segmenting salient objects from natural scenes, often referred to as salient object detection, has attracted great interest in computer vision. While many models have been proposed and several applications have emerged, a deep understanding of achievements and issues remains lacking. We aim to provide a comprehensive review of recent progress in salient object detection and situate this field among other closely related areas such as generic scene segmentation, object proposal generation, and saliency for fixation prediction. Covering 228 publications, we survey i) roots, key concepts, and tasks, ii) core techniques and main modeling trends, and iii) datasets and evaluation metrics for salient object detection. We also discuss open problems such as evaluation metrics and dataset bias in model performance, and suggest future research directions.
Salient object detection (SOD) is a fundamental problem in computer vision with wide-ranging applications, from image compression to content-aware editing and visual tracking. This survey, published in 2019, arrives at a critical juncture when deep learning has revolutionized the field, yet a consolidated understanding of the rapid progress was missing. By reviewing 228 publications, the authors provide a much-needed structured overview, helping researchers navigate the crowded landscape of models, datasets, and metrics.
The paper's significance extends beyond mere cataloging. It explicitly situates SOD among related tasks like generic scene segmentation, object proposal generation, and saliency for fixation prediction, clarifying the subtle distinctions and connections. This contextualization is invaluable for interdisciplinary researchers and for avoiding conceptual confusion that often arises when these terms are used interchangeably. The survey also identifies persistent issues such as dataset bias and the inadequacy of existing evaluation metrics, which are crucial for the field's healthy development.
The survey's main technical contributions are organizational and analytical:
As a survey, the paper does not present new experimental results. Instead, it synthesizes findings from the literature, noting that deep learning-based methods have significantly outperformed traditional approaches on standard benchmarks. It also points out that performance varies across datasets due to bias, and that common metrics like F-measure and MAE have limitations. The survey's value lies in its meta-analysis, offering a bird's-eye view of the field's progress and remaining challenges.
The impact of this survey is substantial. It has become a widely cited reference (839 citations) for researchers entering the field and for practitioners selecting appropriate methods. By clarifying terminology and consolidating knowledge, it helps reduce redundancy and fosters more targeted research. Its discussion of open problems has likely influenced subsequent work on more robust evaluation metrics and unbiased datasets. Moreover, by linking SOD to broader vision tasks, it encourages cross-pollination of ideas, potentially benefiting areas like object detection and segmentation. For the AI community, this survey exemplifies the importance of periodic synthesis to guide a rapidly evolving field.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba