Video generation from text
Unknown
This paper introduces text-to-video generation, a conditional video generation approach that produces videos from textual descriptions.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper introduces text-to-video generation, a conditional video generation approach that produces videos from textual descriptions.
Unknown
This paper proposes a recipe for scaling text-to-video generation by leveraging widely accessible text-free videos, studying the scaling trend and enabling compositional video synthesis.
Muhammad Tanveer Jan, Mohammed G. Al-Jassani, Martinraj Nadar, et al.
A comprehensive survey of text-to-video generators, classifying approaches and providing an overview of the domain.
Unknown
This paper introduces Snap Video, a scaled spatiotemporal transformer model for text-to-video synthesis that generates high-quality videos from text prompts.
Wenyi Hong, Ming Ding, Wendi Zheng, et al.
CogVideo is the first large-scale pretrained transformer for text-to-video generation, leveraging a pretrained text-to-image model to achieve state-of-the-art results.
Unknown
This paper introduces Make-a-video, a method for text-to-video generation that leverages text-to-image models and unsupervised video data, achieving state-of-the-art results without paired text-video data.
Unknown
CogVideoX is a large-scale text-to-video diffusion transformer model that generates 10-second continuous videos aligned with text.
Unknown
Show-1 introduces a hybrid text-to-video generation framework that combines pixel-based and latent diffusion models to leverage their complementary strengths.
Unknown
Controlvideo introduces a training-free framework for controllable text-to-video generation by leveraging pre-trained text-to-image models.
Unknown
Openvid-1m introduces a large-scale, high-quality dataset for text-to-video generation, filtered by aesthetics scores to retain the top 20% of videos.
Unknown
ModelScopeT2V is a text-to-video synthesis model that evolves from Stable Diffusion by incorporating spatio-temporal modeling to generate coherent videos from text prompts.
Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, et al.
A controlled study using a procedural testbed reveals that data distribution balance and caption quality critically impact text-to-video model generalization and training efficiency.