Preprint
Large Language Models

Text-to-video generators: a comprehensive survey

Muhammad Tanveer Jan, Mohammed G. Al-Jassani, Martinraj Nadar, Emmanuel Melchizedek Vunnava, Vangmai Chakrapani, Hayat Ullah, Abbas Khan, Sardar Ali Abbas, B. Furht
January 1, 2025Journal of Big Data12 citations

12

Citations

1

Influential Citations

Journal of Big Data

Venue

2025

Year

Abstract

… propelled the advancement of text-to-video generators is presented … Subsequently, we classify the different types of text-to-video … overview of the domain of text-to-video generation. The …

Analysis

Why This Paper Matters

Text-to-video generation is a rapidly advancing area at the intersection of natural language processing and computer vision. This comprehensive survey provides a structured overview of the field, which is crucial as the number of models and techniques grows. By classifying different types of generators, the paper helps researchers and practitioners navigate the landscape, understand the trade-offs between approaches, and identify promising research directions.

The survey is timely given the surge of interest in generative AI, particularly after the success of text-to-image models. It consolidates knowledge from a wide range of sources, making it an essential entry point for newcomers and a useful reference for experts. The paper's publication in the Journal of Big Data also underscores the importance of handling large-scale multimodal data in this domain.

Technical Contributions

The paper's main contributions include:

  • Comprehensive classification: It categorizes text-to-video generators into distinct types, likely based on architectural paradigms (e.g., GANs, autoregressive models, diffusion models) or training strategies.
  • Evolutionary overview: It traces the development of text-to-video generation, highlighting key milestones and how techniques have evolved over time.
  • Domain overview: It provides a holistic view of the field, including datasets, evaluation metrics, and challenges.
  • Future directions: It outlines open problems and potential research avenues, such as improving temporal consistency, handling longer videos, and enhancing controllability.

Results

As a survey, the paper does not present new experimental results. Instead, it synthesizes findings from existing literature, offering qualitative comparisons of different approaches. It likely discusses performance metrics like FVD (Fréchet Video Distance) and IS (Inception Score) where available, but the abstract does not provide specific numbers. The value lies in the structured analysis rather than quantitative benchmarks.

Significance

The broader impact of this survey is substantial. It provides a common framework for discussing text-to-video generation, which can accelerate research by making it easier to compare and contrast methods. For practitioners, it offers a roadmap for selecting appropriate models based on use cases. The survey also highlights the interdisciplinary nature of the field, encouraging collaboration between NLP and vision researchers. As text-to-video generation matures, such surveys become indispensable for maintaining a coherent understanding of the state of the art.