AI Pulse
📄 论文解读

机器人学动作,不再需要先预测未来

过去让 AI 学操控,主流做法是先让它预测下一帧画面,再从中提取动作。这篇论文反着来:直接让模型在当前画面上做「去噪」,一边清理视觉噪声一边输出动作,把生成模型的看家本领和操控任务焊在一起。实验里,预测未来和不预测未来效果差不多,但只学干净画面、不走去噪过程,鲁棒性明显变差——说明真正有用的不是「未来」,而是「去噪这条路径」本身。在 LIBERO-Plus 基准上,新方法把成功率从 81.6% 提到 87.7%,训练数据量减半,单步耗时从 2.85 秒降到 1.63 秒。它不是你明天就能装进机器人的东西,但它给了一个更省算力、更快的思路:与其让机器人先想象未来,不如让它把当下看清。

📄 原文摘要(英文)

Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, while restricting training to the clean endpoint substantially reduces robustness, suggesting that a separate future target is not essential for generative adaptation, while the denoising trajectory remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784 to 392) and reducing step time from 2.85 s to 1.63 s, a 1.8x speedup. With the pure text-to-image Z-Image backbone, NowWAM further reaches 87.8%, showing that strong control adaptation is not tied to video generation or image-editing backbones.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新