韩语AI为何总差一截?新基准把差距拆开了
韩语AI总被说“差一截”,但差在哪没人说清。新基准NOLLI把英语和韩语的任务做成一一对应的谜题,发现一个反直觉的事实:光是把英文题翻译成韩文,顶尖模型的准确率几乎不掉(差距在10个百分点内)。真正的坑在“写字”上——韩文是音节下拆成字母再拼的,模型一碰这种多步拼写就崩,韩语密码题比英语版差最多68.7个百分点。这不是明天能用的工具,但它告诉你:韩语AI的短板不在“懂韩语”,而在“写韩语”。
📄 原文摘要(英文)
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.