AI Pulse
📄 论文解读

让AI开网店,最强模型也亏不过人类策略

让AI当老板,它行吗?这篇论文搭了个模拟跨境小店的沙盘,让AI从进货、定价到应对延迟到货和监管,全程自己拍板,最后看谁赚得多。结果:15个前沿模型里,最会赚钱的也比不过人类设计的简单策略,而最差的跟最好的差了9倍身家。更关键的是,研究者把AI的盈亏一笔笔拆回具体决策,发现有的AI是薄利多销型,有的是死磕高利润,还有的靠客服挽回差评——它们不是不会做生意,是各有各的生意经。这不是你明天能用的工具,但它第一次让『AI的商业头脑』有了可量化的考场。

📄 原文摘要(英文)

Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新