AI审稿人被措辞骗了:同一篇论文换个说法,评分就变
AI 审稿人有个隐蔽的毛病:同一篇论文,只是换了个更漂亮的写法,它给出的评分就可能不一样——这意味着作者靠润色、而不是靠科研本身,就能刷高分。研究者造了一个 1260 个版本的受控数据集,让 30 种审稿配置去评,发现很多模型看似对措辞不敏感,其实是把分数压平了,谁都差不多,等于没审。更麻烦的是,人类对齐度高的模型,修辞鲁棒性未必好,两者是两码事。他们提出的解法 SciCore 把论文拆成「全文判断」和「只读科学内核(方法、结果、结论)」两条路,取平均,在 GPT-5.5 上做到了既稳定又能区分论文好坏。这不是你明天能用的工具,但它指出了一个真问题:如果 AI 审稿被措辞左右,那它评的到底是科学,还是文笔?
📄 原文摘要(英文)
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.