AI用希腊语思考,准确率没变,但你看不懂了
让大模型用希腊语思考,准确率几乎不变——但这不是重点。研究者把三个前沿MoE模型(阿里、OpenAI、英伟达,各36-40亿活跃参数)微调成用希腊语推理。基准测试上几乎没变化,而且这个基准本身就有问题:只改随机种子就能让分数波动7.7分,比所有数据和训练方法的效果都大。真正的变化在准确率看不到的地方:基础模型从不用希腊语思考——1000条推理轨迹里一条都没有,哪怕问题本身就是希腊语,所以模型用用户看不懂、无法审计、无法纠正的形式给出正确答案。经过监督微调(SFT)后,每个发布的检查点都在约98%的样本上用问题语言推理,一个模型家族还省了3倍token,语法评分提升,通用能力几乎不变:没忘东西,多了流畅。但SFT修不了自己的缺陷:四分之一答案跳过要求的格式,答案泄露到推理通道里,明确要求“用英语思考”时服从率不到一半。用可验证奖励的强化学习(训练前预注册)直接修好了前两个(格式回退从24%降到2.5%,泄露从3.5%降到0.0%),第三个提升了9.1个百分点,而希腊语思考习惯在只优化准确率的梯度下纹丝不动。这不是你明天能用上的技术,但它揭示了一个关键盲区:我们衡量AI的“准确率”可能完全错过模型实际在做什么——它可能用你听不懂的语言思考,而你只看到答案对了。
📄 原文摘要(英文)
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.