AI Pulse
📄 论文解读

AI 分析时间序列时,自带工具反而帮倒忙

给 AI 配一套专家工具,本意是让它分析时间序列更准;结果这套工具在部分任务上反而把准确率拉低了,而且 AI 自己改一轮答案,改了 147 个、改坏了 56 个,总分却几乎没动。问题出在:工具好不好用,得看具体问题,而人是一次性配齐、按平均分评判的。研究者让 AI 自己诊断失败、按缺口造工具、再逐对验收,从空库开始,在十个任务和三个模型上全部变准,而且小模型长出的工具装进大模型照样有用。它不是你明天能用上的,但提醒你一件事:给 AI 加装备,别只看平均分,得看它在哪些问题上拖后腿。

📄 原文摘要(英文)

Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test. Silent Harm: one round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point. Both follow from the same gap: whether a tool helps is decided question by question at runtime, while tools are supplied in advance and judged by a single average. To address this, we propose TimeEvo, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones. Code is available at https://github.com/Muyiiiii/TimeEvo.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新