Conference Paper
Large Language Models

On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation

Lin Long, Rui Wang, Rui Xiao, Junbo Zhao, Xiao Ding, Gang Chen, Haobo Wang
June 14, 2024Annual Meeting of the Association for Computational Linguistics344 citations

344

Citations

21

Influential Citations

Annual Meeting of the Association for Computational Linguistics

Venue

2024

Year

Abstract

Within the evolving landscape of deep learning, the dilemma of data quantity and quality has been a long-standing problem. The recent advent of Large Language Models (LLMs) offers a data-centric solution to alleviate the limitations of real-world data with synthetic data generation. However, current investigations into this field lack a unified framework and mostly stay on the surface. Therefore, this paper provides an organization of relevant studies based on a generic workflow of synthetic data generation. By doing so, we highlight the gaps within existing research and outline prospective avenues for future study. This work aims to shepherd the academic and industrial communities towards deeper, more methodical inquiries into the capabilities and applications of LLMs-driven synthetic data generation.

Analysis

Why This Paper Matters

This paper addresses a critical bottleneck in deep learning: the scarcity and quality limitations of real-world data. By leveraging Large Language Models (LLMs) for synthetic data generation, the field gains a scalable, data-centric solution. However, prior work lacked a unified framework, making it difficult to compare methods or identify systematic gaps. This paper fills that void by proposing a generic workflow that covers generation, curation, and evaluation, thus providing a structured lens for both researchers and practitioners.

The timing is significant: with LLMs becoming more capable and accessible, synthetic data is increasingly used in domains like NLP, code generation, and multimodal learning. Without a coherent organizational scheme, efforts remain fragmented. This paper's taxonomy helps consolidate knowledge and directs future work toward the most pressing challenges, such as data diversity, fidelity, and bias mitigation.

Technical Contributions

  • Unified Workflow: The paper defines a three-stage pipeline—generation, curation, evaluation—that encapsulates the end-to-end process of using LLMs for synthetic data.
  • Gap Analysis: By mapping existing studies onto this workflow, the authors identify underexplored areas, such as automated curation methods and robust evaluation metrics.
  • Future Directions: The paper outlines specific research avenues, including improving controllability of generation, reducing artifacts, and developing domain-specific benchmarks.
  • Cross-Domain Applicability: The framework is designed to be generic, applicable to text, code, and potentially other modalities.

Results

As a survey and position paper, no experimental results are reported. The main output is a structured organization of the literature and a set of identified research gaps. The paper does not provide quantitative metrics or comparisons between synthetic data generation methods.

Significance

This paper serves as a roadmap for the growing field of LLM-driven synthetic data. By providing a common vocabulary and workflow, it enables more systematic comparisons and collaborations. For industry practitioners, it offers a checklist for building robust synthetic data pipelines. For academics, it highlights open problems that could drive future research. The framework's generality means it can adapt as LLMs evolve, ensuring long-term relevance.