AI Pulse
📄 论文解读

大模型换种语言就变笨?差距能算出来

同一个模型,用英语考是学霸,换成马拉地语就掉成学渣——这不是玄学,是可以算出来的。研究者造了3万道15种语言的数学题,发现决定模型跨语言表现的因素按影响力排序:模型大小、语言资源丰富度、推理能力、语言亲缘距离。最直观的换算:32B模型用马拉地语答题,效果约等于10B模型用英语。而且模型越大、推理越强,低资源语言和高资源语言的差距越小;但语言亲缘距离远,这些招都不太管用。这套框架能解释92%的语言间差异,预测模型在没见过的语言上的表现,误差在6个百分点内。对开发者来说,这意味着不用把每种语言都测一遍,先看模型大小和语言资源等级就能估个大概。

📄 原文摘要(英文)

We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size (β= 1.77), language resource level (β= 0.77), reasoning (β= 0.67) and typological distance (β= -0.25). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages (β= -0.27 and β= -0.20, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新