给AI打分的人也有偏见,这篇能现场改掉
给大模型当裁判的奖励模型,也会看人下菜:回答越长、语气越自信,它就给高分,哪怕内容并不更对。以前想治这个毛病,要么重训整个模型(费数据费算力),要么只能针对某一种已知偏见(比如“偏爱长回答”)做固定修正。这篇提出 BiasReducer,只改奖励模型最末端那层线性头,而且能按新数据集现场挑毛病:先用稀疏自编码器找出模型到底对哪些表面特征敏感,再学会对每个特征该往哪个方向、调多少分,最后面对新数据时按影响力排序、只改相关的。在五个奖励模型上,它让三个基准平均提升 8.3、18.0、6.9 个百分点,超过两种需要重训的方案,而且这种修正还能传导到下游,让模型少说废话、少拍马屁,同时评判质量不掉。它不是你明天能直接装上的工具,但“只动最后一层、按数据现场选偏见”这个思路,让修 AI 裁判的成本从重训降到了微调级别。
📄 原文摘要(英文)
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.