高清无码

高清无码

Video Generation Models are General-Purpose Vision Learners

发布时间:2026-08-25

时   间:10:00-11:00, Aug 7, 2026 (Fri)

地   点://meeting.tencent.com/dm/z9BkdA8pj32P (//meeting.tencent.com/dm/z9BkdA8pj32P)

内容:

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision?
In this talk, we contend that large-scale text-to-video generation serves as a pivotal pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions.
Empirical results demonstrate that achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g., DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2).
Furthermore, we demonstrate the video generative pretrained backbone outperforms alternative pretraining paradigms, exhibits preliminary data and model scaling properties, along with exceptional data efficiency, and triggers intriguing emergent behaviors. These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world.

个人简介:

Letian Wang is a final-year Ph.D. student at the University of Toronto and the Vector Institute, advised by Prof. Steven Waslander. He has conducted research with Google DeepMind, NVIDIA Research, Carnegie Mellon University, and UC Berkeley. He is the author of two books, and his work has been recognized with the Qualcomm Fellowship, the RA-L Best Paper Award Honorable Mention, and the CARLA Autonomous Driving Challenge championship. His research aims to develop multimodal AI systems that can generalize and adapt to the open world, much like humans do. More recently, he has focused on scalable perception and generalizable decision-making by leveraging foundation models and learning paradigms that scale effectively with data.

返回列表
演讲人 Letian Wang 时间 10:00-11:00, Aug 7, 2026 (Fri)
地点 EN
TOP