AI Pulse
📄 论文解读

AI 评测自己出题,把科学家考倒了

现在的 AI 评测有个尴尬:题出得太慢,AI 进步太快,等题出好,AI 已经会了。这篇让 AI 自己出题考自己:先定一个「考什么」的概念(领域、数据类型、推理方式),再写一份「怎么考」的配方(问题、环境、标准答案),AI 试着解题,评测系统看它怎么解、错在哪,然后改题——专堵 AI 的捷径,逼它去看原始数据、解读中间结果、整合证据。从已有题库出发,在计算生物学和材料科学上,AI 自己出的题把 AI 的准确率压低了 22 和 25 个百分点,而且题的质量评分还更高。它不是你明天能用上的东西,但它指向一个趋势:以后衡量 AI 能力,可能不再靠人出题,而是靠 AI 和 AI 互相较劲。

📄 原文摘要(英文)

As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain expertise, making it difficult to keep evaluation aligned with advances in agent capabilities. We address this challenge by investigating whether scientific-agent benchmarks can be automatically generated and iteratively adapted as agent capabilities evolve. We introduce AutoSciBench, a framework that represents each task as a high-level concept specifying the scientific domain, data modality, and required reasoning approach, together with a low-level recipe specifying how the question, environment, and ground-truth answer are constructed and verified. Agents attempt to solve each task, producing solver trajectories and corresponding judge feedback which AutoSciBench uses to revise the recipe or concept, closing observed shortcuts and shifting tasks toward raw-data re-examination, interpretation of intermediate results, and evidence integration. Experience distilled from completed refinement trajectories further guides new concept generation, allowing lessons from earlier task refinement to inform subsequent benchmark construction. Starting from existing benchmarks, we evaluate AutoSciBench across computational biology, materials science, and clinical imaging. Generated benchmarks reduce average solver accuracy by 22.4 and 25.5 percentage points relative to the human-curated benchmarks in computational biology and materials science, respectively, while generated tasks receive higher average quality ratings across all three domains, suggesting that scientific-agent evaluation can adapt as agent capabilities advance.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新