让AI学会自己纠错:一个数学小技巧
训练AI做动作(比如机器人操作)时,通常先让它模仿人类示范,再用强化学习让它自己改进。但现有方法在每一步都要算复杂的数学量,成本高。研究者发现,预训练模型的数学结构有个规律:平均来看,速度的雅可比矩阵几乎是对角的。利用这一点,他们推导出一个简单的标量公式,把最终动作的价值信息直接传回每一步,省掉了大量计算。在四个最难的任务上,成功率比最强基线高18到35个百分点;在真实双臂机器人上,三个任务都超过监督微调。这不是你明天就能用的工具,但它让AI自我改进变得更便宜、更可行,是前沿方向上的一个实在进展。
📄 原文摘要(英文)
Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.