AI Pulse
📄 论文解读

机器人有了通用大脑:一个模型搞定看、想、动

机器人通常一个任务训练一个模型,换个环境就废。RynnBrain 1.1 用一个统一框架,把感知、空间推理、定位、规划全塞进一个模型,从20亿到1220亿参数都有。它新增了接触点预测和原生3D定位,让机器人知道该抓哪里、怎么放。在真实机器人上,它比Qwen等通用模型成功率更高,而且一个模型能同时控制不同机器人(人形、双臂等),多任务一起学反而比单独学更好。这不是你明天能用的产品,但它指向一个方向:机器人不再需要为每个动作单独编程,一个通用大脑就够了。

📄 原文摘要(英文)

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新