ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… To enhance text-to-video generation, we filter out videos with low aesthetics scores using the LAION Aesthetics Predictor. This results in a subset SA with the top 20% highest-scoring …
Text-to-video generation is a rapidly evolving field, yet the quality of training data often lags behind the sophistication of models. Openvid-1m addresses this gap by introducing a large-scale dataset specifically curated for high aesthetic quality. By filtering videos using the LAION Aesthetics Predictor, the authors ensure that models trained on this data are exposed to visually appealing content, which is crucial for generating videos that meet user expectations.
The dataset's scale—one million video-text pairs—is significant. Many existing datasets are either small or contain noisy, low-quality samples. Openvid-1m aims to combine scale with quality, offering a resource that can be used for both training and benchmarking. This is particularly important as the community moves toward more realistic and engaging video generation.
Moreover, the methodology of using an automated aesthetic filter is a practical approach that can be replicated and adapted. It demonstrates a cost-effective way to curate large datasets without manual annotation, which is often impractical at this scale.
The abstract does not include quantitative results, as the paper focuses on dataset construction. However, the implied outcome is that models trained on Openvid-1m will produce higher-quality videos compared to those trained on unfiltered data. The top-20% filtering likely improves metrics such as FID or user preference scores, though specific numbers are not provided in the abstract.
Openvid-1m has the potential to become a foundational resource in text-to-video research, similar to how ImageNet and LAION-5B have influenced image generation. By providing a high-quality, large-scale dataset, it lowers the barrier for researchers to train competitive models and encourages standardization in evaluation. The aesthetic filtering approach also highlights the importance of data quality over sheer quantity, a lesson that extends beyond video generation to other multimodal tasks.
Future work could explore extending the dataset to more diverse domains or incorporating additional quality metrics. The dataset's release is likely to spur innovation in video generation, making it a significant contribution to the AI community.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba