AI Pulse
📄 论文解读

机器人能听懂“刚才那个东西”了

机器人一直听不懂“刚才你拿的那个杯子”——你指的是事件里的角色,不是名字,而且它可能已经被挡住看不见了。这套系统把“事件记忆”和“当前画面”拼起来:先靠历史推断你要的是哪个物体,再结合当前几何结构主动换视角,把被挡住的物体找出来再抓。零训练、直接用预训练模型,真实机器人上手,一开始就看得见的抓取成功率76%,被挡住的也有77%,而最强基线分别只有40%和55%。它不是你明天能用上的,但“机器人能理解事件指代”这件事,往前迈了一步。

📄 原文摘要(英文)

A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests may specify a grasp target by the role it played in a past event rather than by its name or appearance. Moreover, the target may no longer be visible when the robot is asked to act. We present BeyondSCe, a zero-shot robotic grasping system for this event-referential setting. Given the event history and the current scene, the system identifies the requested object or part and localizes it for grasping. If the target is occluded, it combines an event prior recovered from the history with current scene geometry to select camera viewpoints likely to reveal the target. The system uses pretrained models without additional task-specific training. In real-robot experiments with a single wrist-mounted RGB-D camera, it achieves grasp success rates of 76% and 77% for initially visible and occluded targets, respectively, compared with 40% and 55% for the strongest baseline in each condition. On four additional scenes with heavy occlusion, it increases grasp success rates from 75% to 95% while reducing the mean number of views from 3.35 to 2.20, compared with an active-perception baseline given the target's ground-truth 3D bounding box.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新