
Google DeepMind: Video Generators Hold Key to Computer Vision
Researchers at Google DeepMind have developed GenCeption, a model that repurposes a pre-trained video generator for classic computer vision tasks like depth estimation and segmentation. The system builds on an open-source video model from Alibaba and delivers results in a single forward pass, guided by text prompts. It trained on just a small set of synthetic videos and needs far less data than competing approaches. In benchmarks, GenCeption matches established specialized models and transfers its abilities to real-world footage and untrained categories like animals. The authors say this supports the contested idea that video generators can serve as the basis for universal world models in computer vision.






