给AI审美打分,不如给它一套评分标准
给AI模型做“审美训练”一直很难:人类觉得好看的标准太模糊,单一打分器容易被钻空子。这篇的做法是把奖励拆成两半——一半学人类偏好数据,另一半用明确的评分规则盯住“是否忠实于你的指令”,再把两者组合起来,而不是简单加权平均。结果:在公开竞技榜上,训练后的模型Elo分比基础版高69分,其中一个模型直接超过所有开源对手。它不是你明天能用上的东西,但提示了一个方向:想让AI更听话,别只教它“好看”,还得教它“守规矩”。
📄 原文摘要(英文)
Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals. Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking. A key challenge is how to combine these heterogeneous reward signals. We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction. In the Arena text-to-image leaderboard (https://arena.ai/), our RL-trained Flux2dev achieves an Elo rating 69 points above the base model, and our post-trained Ideogram-4 surpasses every open-source model on the leaderboard, reaching an Elo of 1223.5. (Claims of state-of-the-art performance are based on the Arena leaderboard snapshot as of September 4, 2026.) Our results suggest that effective rewards for frontier generative-model training require broad coverage of user intent and robustness to exploitation under optimization. To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.