Journal Article
Machine Learning

Analysis of Points of Interests Recommended for Leisure Walk Descriptions

Bajaj, Payal(Leipzig University), Campos, Daniel(Friedrich Schiller University Jena), Craswell, Nick(University of Kassel), Deng, Li(Friedrich Schiller University Jena), Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Gia Nguyen, Mir Rosenberg, Song, Xia, Alina Stoica, Saurabh Tiwary, Tong Wang
October 10, 2024DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)1,292 citations

1.3k

Citations

0

Influential Citations

DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)

Venue

2024

Year

Abstract

Data for Sub-Task 1 of the Advertisement in Retrieval-Augmented Generation task at Touché 2025. The dataset contains segments retrieved from the segmented version of MS MARCO V2.1. The queries used in retrieval are taken from the Webis Generated Native Ads 2024 dataset.

Analysis

Why This Paper Matters

This paper addresses a growing intersection of retrieval-augmented generation (RAG) and advertising. As RAG systems become more prevalent in real-world applications, their use in generating contextual advertisements is a natural extension. The dataset provided here fills a gap by offering a benchmark for evaluating RAG-based ad generation, which is crucial for both academic research and industry deployment.

The Touché 2025 shared task on Advertisement in RAG is a timely initiative, and this dataset serves as its foundation. By leveraging MS MARCO V2.1, a well-known large-scale retrieval corpus, and the Webis Generated Native Ads 2024 dataset, the authors ensure that the data is both realistic and challenging. This allows participants to focus on algorithmic improvements rather than data collection.

Technical Contributions

The key technical contribution is the creation of a segmented version of MS MARCO V2.1, which enables passage-level retrieval instead of document-level. This is important because ads often need to be matched to specific content segments, not entire documents. The dataset also integrates queries from the Webis Generated Native Ads 2024 dataset, which are designed to simulate real ad-related search queries.

Another contribution is the clear specification of Sub-Task 1, which likely involves retrieving relevant segments for a given ad query. This provides a standardized evaluation protocol, making results comparable across different systems. The dataset is publicly available, promoting reproducibility and further research.

Results

As a dataset paper, there are no experimental results or metrics. The primary outcome is the dataset itself, which will be used in the Touché 2025 shared task. The paper does not include baseline evaluations or comparisons, so concrete performance numbers are absent. However, the dataset's construction from established resources (MS MARCO and Webis) suggests it is of high quality and relevance.

Significance

The broader impact of this work lies in enabling research on RAG for advertising. This could lead to more effective and personalized ad generation, which is valuable for businesses and users alike. The dataset also encourages the development of evaluation metrics for RAG systems in non-traditional domains, potentially influencing future benchmarks.

Moreover, by making the dataset available, the authors contribute to the open research ecosystem, allowing others to build upon their work. This is particularly important in the fast-evolving field of RAG, where data resources are often scarce. The paper sets the stage for a competitive shared task that could yield innovative solutions and insights.