AI Pulse
📄 论文解读

AI技能进化卡在单轮反馈,这篇用多轮追问打破僵局

现在的AI智能体技能,要么人写,要么一次生成,用完就完,不会从自己犯的错里学习。最近有人想补上这个闭环,但反馈只来自单轮问答——第一轮把能补的洞补完,进化就停了,那些只在多轮对话里才露出来的毛病永远看不见。SkillEvo把多轮用户模拟从「考试」变成「出题人」:追问一层层逼出缺陷,每修一轮,又产生新的反馈,进化不停。同时它不再用单一分数当闸门,而是加了个独立治理层,主动修复事实错误和结构臃肿,防止进化方向跑偏。在6类云服务、9个生产级技能、98个技能文件上,比自我反思式进化高23分,比单轮问答驱动高15.4分。这不是你明天能用的东西,但它指了个方向:AI技能要持续变强,关键不在改代码的能力,而在反馈能不能一直给得出新东西。

📄 原文摘要(英文)

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新