AI Pulse
📄 论文解读

AI审稿人会被修辞骗:同一篇论文,换个说法分就变了

AI 审稿人不是铁面无私的。研究者把 120 篇真实论文改写成 4200 个版本,只动修辞不动内容——比如把同样的发现说成「突破」还是「初步证据」——结果 AI 审稿人的评分跟着变,而且规律很怪:原本低分的论文换个说法容易涨分,原本高分的反而容易跌,中间分数段最敏感。更反直觉的是,越复杂的改写流程(多轮、带指导)并不比简单改一遍更有效,效果主要取决于改写者本身。这不是说 AI 审稿没用,而是提醒:当 AI 开始参与论文评审,措辞本身就成了变量。它不是你明天能用上的工具,但如果你在投稿,知道「AI 审稿人可能吃修辞这一套」这件事本身,就是信息。

📄 原文摘要(英文)

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新