AI Pulse
📄 论文解读

AI 在对话中跟不上你变卦

你刚说想买红色跑车,聊两句又改成蓝色 SUV——AI 助手大概率已经懵了。研究者把静态任务改造成多轮对话,让用户意图在对话中逐步揭示、修改甚至转向,结果发现:在单轮任务上表现最好的模型,到了意图动态变化的环境里,性能大幅跳水。这不是你明天能用上的功能,但它点出了一个关键盲区:今天的 AI 擅长一次猜对,却不擅长跟着你变。

📄 原文摘要(英文)

As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新