AI Pulse
📄 论文解读

AI打分总往中间挤,不是笨,是偏见

AI 做分类或打分时,明明该给出明确答案,却总爱往中间选项靠。研究者发现,这类模型在 36 个有序评分任务上,只用了真实答案范围的三分之二左右;选项越多,压缩越狠,14 个档位时利用率掉到 26%—75%。这不是模型笨,也不是数据不平衡,而是它学来的习惯——用少量训练数据微调后,利用率能从 47% 拉回 86%。换句话说,你让 AI 评 1 到 10 分,它心里其实只有 1 到 7。

📄 原文摘要(英文)

Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8\% of all predictions and 51.3\% of errors to Neutral despite 74.95\% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76\% of the effective gold support, versus 87--102\% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from K=2 to 14; utilization falls for every model and reaches 26--75\% at K=14, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47\% to 86\% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新