ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
1.3k
Citations
0
Influential Citations
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
Venue
2024
Year
Data for Sub-Task 1 of the Advertisement in Retrieval-Augmented Generation task at Touché 2025. The dataset contains segments retrieved from the segmented version of MS MARCO V2.1. The queries used in retrieval are taken from the Webis Generated Native Ads 2024 dataset.
This paper addresses a growing intersection of retrieval-augmented generation (RAG) and advertising. As RAG systems become more prevalent in real-world applications, their use in generating contextual advertisements is a natural extension. The dataset provided here fills a gap by offering a benchmark for evaluating RAG-based ad generation, which is crucial for both academic research and industry deployment.
The Touché 2025 shared task on Advertisement in RAG is a timely initiative, and this dataset serves as its foundation. By leveraging MS MARCO V2.1, a well-known large-scale retrieval corpus, and the Webis Generated Native Ads 2024 dataset, the authors ensure that the data is both realistic and challenging. This allows participants to focus on algorithmic improvements rather than data collection.
The key technical contribution is the creation of a segmented version of MS MARCO V2.1, which enables passage-level retrieval instead of document-level. This is important because ads often need to be matched to specific content segments, not entire documents. The dataset also integrates queries from the Webis Generated Native Ads 2024 dataset, which are designed to simulate real ad-related search queries.
Another contribution is the clear specification of Sub-Task 1, which likely involves retrieving relevant segments for a given ad query. This provides a standardized evaluation protocol, making results comparable across different systems. The dataset is publicly available, promoting reproducibility and further research.
As a dataset paper, there are no experimental results or metrics. The primary outcome is the dataset itself, which will be used in the Touché 2025 shared task. The paper does not include baseline evaluations or comparisons, so concrete performance numbers are absent. However, the dataset's construction from established resources (MS MARCO and Webis) suggests it is of high quality and relevance.
The broader impact of this work lies in enabling research on RAG for advertising. This could lead to more effective and personalized ad generation, which is valuable for businesses and users alike. The dataset also encourages the development of evaluation metrics for RAG systems in non-traditional domains, potentially influencing future benchmarks.
Moreover, by making the dataset available, the authors contribute to the open research ecosystem, allowing others to build upon their work. This is particularly important in the fast-evolving field of RAG, where data resources are often scarce. The paper sets the stage for a competitive shared task that could yield innovative solutions and insights.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba