AI Pulse
📄 论文解读

机器人学坏容易,学“进步”难

机器人学一个任务,通常只靠“成功/失败”的最终信号。但现实是:它可能中途走偏、原地打转,甚至倒退——这些信息全被浪费了。这篇综述把散落的研究串起来,统一了“进度奖励”的框架:先看模型接收什么信息(比如摄像头画面还是关节角度),再看它如何判断“有进展”(比如距离目标更近、动作更流畅),最后看用什么数据训练和验证。它不是你明天就能用的工具,但如果你关心机器人怎么学会“自己判断做对了”,这是目前最清晰的路线图。

📄 原文摘要(英文)

Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three connected steps. We first study the interface of a progress model. This defines the problem from the outside by asking what information the model receives and what form of progress signal it produces. We then move inside the model and study the methods used to construct this signal. This reveals the different assumptions and mechanisms behind progress estimation and reward generation. Finally, we examine the data and benchmarks that support these methods. This shows how progress supervision is obtained and what different evaluations actually measure. Together, these three perspectives connect what a progress model is, how it is built, and how its quality is validated. We further summarize the main limitations of current approaches and discuss future research directions.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新