AI审稿人被措辞骗了:同一篇论文换个说法,评分就变
AI 审稿人有个隐蔽的毛病:同一篇论文,换个措辞重写一遍,它给出的评分可能就变了。这意味着作者与其打磨科学,不如打磨修辞——AI 在奖励“说得好”,而不是“做得好”。研究者造了一个受控数据集:把 1260 个论文版本按“内容不变、只换说法”的方式生成,拿去测 30 种审稿配置,发现很多模型对改写不敏感,但代价是对不同论文的评分全挤在一起、失去区分度——这叫“假稳健”。更麻烦的是,人类偏好和修辞稳健性给审稿人排出的名次不一样:你觉得它靠谱,它可能只是对措辞迟钝。他们提出的解法是把论文拆成“科学核心”(提取出的结构化内容)和“全文”两条路分别打分再平均,让修辞的影响被稀释。这不是你明天能用的工具,但它指出了一个真问题:当 AI 开始审稿,作者的第一竞争力可能不再是科学,而是话术。
📄 原文摘要(英文)
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.