AI Pulse
📄 论文解读

会下棋还会讲棋的AI,棋力逼近特级大师

现在的棋类AI是沉默的专家:能赢你,但说不清为什么。这篇把「会下棋的引擎」和「会说话的模型」拼在一起,造出一个4B参数的小模型Queen,棋力从1782分一路练到2697分——跨过特级大师门槛,还超过了所有前沿大模型,而参数只有它们的千分之一。它的训练方式有点意思:先让模型看自己最看好的几步棋之后会怎样,再把「为什么这么走」总结成话,反复蒸馏回自己脑子里,七轮迭代涨了900多分。它不是你明天能用上的东西,但「让专家开口说话」这条路,对机器人、电脑操作这类有高手但没解说的领域,是个值得盯的方向。

📄 原文摘要(英文)

Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新