AI Pulse
📄 论文解读

AI 在真实企业数据上干活,最强模型也只及格一半

现在的 AI 评测都在考「写 SQL 查一张表」,但真实企业里,一个业务问题要跨几十张表、做统计、再动手改数据。这篇造了一个纽约外卖平台的完整模拟世界:8,100 万订单、235 张表、75 亿行,还藏了「真相」不让 AI 直接看到——它得自己翻仓库找线索,再执行封号、调预算、补发工资这类操作,最后按后果打分。14 个最强模型里,最好的也只在一半任务上拿到 60 分以上。它不是你明天能用上的工具,但它第一次把 AI 从「答题」推到了「在真实数据环境里干活」的考场上。

📄 原文摘要(英文)

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新