Journal Article
Large Language Models

A Survey on ChatGPT: AI–Generated Contents, Challenges, and Solutions

Yuntao Wang(Xi'an Jiaotong University), Yanghe Pan(Xi'an Jiaotong University), Miao Yan(Xi'an Jiaotong University), Zhou Su(Xi'an Jiaotong University), Tom H. Luan(Xi'an Jiaotong University)
January 1, 2023IEEE Open Journal of the Computer Society333 citations

333

Citations

10

Influential Citations

IEEE Open Journal of the Computer Society

Venue

2023

Year

Abstract

With the widespread use of large artificial intelligence (AI) models such as ChatGPT, AI-generated content (AIGC) has garnered increasing attention and is leading a paradigm shift in content creation and knowledge representation. AIGC uses generative large AI algorithms to assist or replace humans in creating massive, high-quality, and human-like content at a faster pace and lower cost, based on user-provided prompts. Despite the recent significant progress in AIGC, security, privacy, ethical, and legal challenges still need to be addressed. This paper presents an in-depth survey of working principles, security and privacy threats, state-of-the-art solutions, and future challenges of the AIGC paradigm. Specifically, we first explore the enabling technologies, general architecture of AIGC, and discuss its working modes and key characteristics. Then, we investigate the taxonomy of security and privacy threats to AIGC and highlight the ethical and societal implications of GPT and AIGC technologies. Furthermore, we review the state-of-the-art AIGC watermarking approaches for regulatable AIGC paradigms regarding the AIGC model and its produced content. Finally, we identify future challenges and open research directions related to AIGC.

Analysis

Why This Paper Matters

This survey arrives at a critical juncture where large AI models like ChatGPT are being rapidly deployed across industries, yet their security, privacy, and ethical implications remain underexplored. By systematically cataloging threats and solutions, the paper serves as a foundational reference for both researchers and practitioners. It highlights the dual-use nature of AIGC—while enabling efficient content creation, it also introduces risks such as misinformation, data leakage, and model misuse. The emphasis on watermarking as a regulatory mechanism is particularly timely given ongoing debates about AI content provenance and accountability.

The paper's comprehensive taxonomy of threats—ranging from adversarial attacks on generative models to privacy violations in training data—provides a structured lens for understanding the attack surface of AIGC systems. This is essential for developing robust defenses and for informing policy decisions. The inclusion of ethical and societal dimensions, such as bias amplification and job displacement, broadens the discussion beyond technical fixes to include human-centric concerns.

Technical Contributions

  • Enabling Technologies and Architecture: The paper delineates the general architecture of AIGC systems, including data collection, model training, and content generation pipelines. It categorizes working modes (e.g., text-to-text, text-to-image) and key characteristics like scalability and prompt sensitivity.
  • Security and Privacy Threat Taxonomy: A structured classification of threats is provided, covering model inversion, membership inference, adversarial examples, and backdoor attacks. Privacy threats include data extraction and unintended memorization.
  • Watermarking Approaches: The survey reviews both model-level watermarking (embedding identifiers in model weights) and content-level watermarking (e.g., invisible patterns in generated text or images). It compares robustness, capacity, and detectability trade-offs.
  • Ethical and Societal Implications: The paper discusses bias, fairness, transparency, and accountability issues, linking them to specific AIGC use cases.

Results

As a survey, the paper does not present new experimental results. Instead, it aggregates findings from prior works, noting that watermarking techniques achieve detection rates above 90% in controlled settings but degrade under content modification attacks. The threat taxonomy is validated by citing real-world incidents (e.g., ChatGPT data leaks). No quantitative comparisons between methods are provided.

Significance

This survey fills a gap by providing a holistic view of AIGC challenges beyond model performance. It is likely to influence both academic research directions—such as developing more robust watermarking and privacy-preserving training—and industry practices around content moderation and compliance. The paper's structured threat taxonomy can serve as a checklist for security audits of AIGC systems. Its discussion of ethical implications also contributes to ongoing regulatory efforts, such as the EU AI Act and watermarking mandates. By framing AIGC as a dual-use technology, the paper encourages balanced innovation that prioritizes safety and trustworthiness.