AI Pulse
📄 论文解读

AI 做生意一年,只赚到人类零头

让 AI 当电商老板,给它 26 个工具、98,843 条真实商品数据,模拟经营一整年。结果最强的大模型只赚到人类平均净资产的 27.3%。这不是它不会聊天,而是它撑不起长期连贯的决策:今天进的货影响明天定价,供应商消息来得快、订单反馈来得慢,AI 容易顾此失彼。它不是你明天能雇来管店的,但这份差距划出了 AI 从「会答题」到「会做事」的真实距离。

📄 原文摘要(英文)

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新