AI Pulse
📄 论文解读

AI 程序员开始考机器人工程师执照

AI 写代码已经卷到物理世界了,但以前的考试只考它能不能写出一个能动的机器人;这篇把考试范围拉大,考它能不能像真正的机器人工程师一样干活:不光写控制程序,还要训练策略、做感知估计、设计机械结构,而且要在资源有限、反馈来自真实传感器的情况下,把一堆零件拼起来、修好、调优。研究者给 AI 程序员出了一套四门课的资格考,每门用不同的标准打分,最后合成一个总指数。结果不意外:AI 在单个任务上能打,但一进入要整合、诊断、迭代的完整工程流程就露馅。它不是你明天能用上的东西,但它划了一条线:机器人这行,AI 离出师还远。

📄 原文摘要(英文)

Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from multimodal feedback. To evaluate these broader capabilities, we introduce RLE-Bench, a benchmark of robot-learning tasks spanning four representative robotics development workflows: interactive control, policy learning, perception and estimation, and mechanical design. We use diverse task-specific metrics to evaluate the artifacts submitted by the coding agents, from the success rate the agents achieved to the policy agents trained, the harness agent built, and the mechanical structures the agent designed. We aggregate these metrics into an overall RLE Index and report workflow-specific capability profiles, enabling systematic comparison of coding agents' capabilities across multiple capability dimensions. Beyond performance ranks, we also conduct in-depth case studies examining agent behavior on representative tasks, highlighting both current capabilities and limitations, and pointing to the opportunities robotics tasks have to offer for future agent training.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新