AI Pulse
📄 论文解读

翻译模型被人类出题人集体考倒

机器翻译的基准测试正在失效:模型太强,旧题全被做穿,自动评分又能被钻空子,连人工评分都不可复现。于是研究者换了个思路——不再出“正常句子”,而是专门收集人类精心设计、能考倒最强翻译模型的刁钻例子,涵盖文字、图片、音频、视频,每个例子还配了人工写的“判卷规则”,专门盯着模型具体会在哪里翻车。这不是你明天能用的工具,但它把“AI 翻译到底行不行”这个问题的答案,从“分数很高”变成了“在哪些具体地方不行”。

📄 原文摘要(英文)

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新