AI Pulse
📄 论文解读

给AI打分前,先想清楚该打什么分

现在的视觉奖励模型,给AI生成的图片打分时,基本是“看题就打分”——不管题目是什么,直接拿条件和你生成的图,硬算出一个分数。这篇论文反着来:先想清楚“这一题到底该看什么”,再照着这个标准去评。研究者把打分拆成三步:先为每个具体任务生成一份“评分细则”,再按细则逐项评估,最后给出细粒度分数。他们还发现,常见的“两两对比”训练法会让分数两极分化,于是设计了一种新训练方式,既保留细粒度打分,又提升区分度。在图像生成和编辑的评测上,它超过了所有开源奖励模型,逼近闭源商业模型;拿它当强化学习的奖励信号,也能稳定提升多种生成模型的表现。它不是你明天就能直接用的工具,但方向值得注意:AI评估AI,正在从“凭感觉打分”走向“先定标准再打分”。

📄 原文摘要(英文)

Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新