韩语AI为何总差一截?新谜题基准揪出真凶
韩语AI总被说比英语差,但差在哪?新基准NOLLI用15类谜题、7500道题,把语言差距拆开看:纯翻译成韩语,模型表现几乎不掉;但一涉及韩文字母的拆解组合,差距能拉到68.7个百分点。更意外的是,谜题变难时,题目变长并不等于更难——7类谜题在“简单”和“难”档上难度没区别。这不是你明天能用的工具,但它告诉我们:韩语AI的短板不在语言本身,而在处理韩文书写系统的精细操作上。
📄 原文摘要(英文)
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.