对着麦克风说话,AI实时生成你的数字分身视频
你对着麦克风说一句,屏幕里的数字人立刻照做——不是录好的,是实时生成的。Vidu S1 把视频生成从「等几分钟出片」变成了「边说边演」:你上传一张真人、动漫或宠物的照片,选个声线,然后像直播一样说话,AI 就以最高每秒 42 帧的速度生成 540p 视频,嘴型、表情、动作都跟着你的语音走,而且可以无限延长,不会模糊或变形。它跑在普通消费级显卡上,不是只有大公司才玩得起的。目前有在线 demo 可玩。这不是你明天就能用来做电影的工具,但它第一次让「对着空气说话就能生成自己的数字分身」变得实时、流畅、人人可试。
📄 原文摘要(英文)
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions. Vidu S1 supports infinite-length real-time video generation without blurring, drift, or visual distortion. Built with TurboDiffusion and TurboServe, Vidu S1 outputs 540p real-time videos at up to 42 FPS on regular consumer GPUs. Users can upload custom images of real people, anime, and pets, and choose different voice tones for personalized experiences. Experiments show that Vidu S1 achieves the best performance across all test metrics while fully meeting real-time inference requirements. A playable online demo is available at https://vidu.com/vidu-stream.