AI Pulse
📄 论文解读

AI排行榜骗了你:高分不代表好用

你看到的AI智能体排行榜,可能全是假的。这篇论文用14个平行实验、7个已有基准,证明了一个反直觉的事实:在测试集上得分最高的模型,换到真实场景中排名可能暴跌。研究者发现,当前排行榜只看平均分,但平均分无法预测模型在“没见过”的任务上的表现——就像高考状元不一定能解决实际问题。他们提出用“预测有效性”代替平均分:即模型在测试集上的排名,与它在真实场景中排名的相关性。这就像用模拟考成绩预测高考排名,而不是只看一次月考分数。论文还设计了一套12维测量框架,覆盖部署时真正重要的维度(如多模态、推理模式、基础设施优化等),并给出了3个可验证的“出分布”标准。虽然证据还不够充分,但方向明确:别再迷信排行榜,要看模型在真实场景中的稳定性。

📄 原文摘要(英文)

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新