AI Pulse
📄 论文解读

字幕翻译的AI终于开始追剧了

字幕翻译一直是个被低估的难题:一集剧里的梗、人物语气、专有名词,得跨几十集保持统一,而现在的AI基本是一句一句翻,翻完就忘。这篇提出一个会自我进化的多智能体系统:它先看一部分剧集,把人物、术语、风格记成长期记忆,再让多个AI角色分工——有查术语的、有检查字幕格式的、有翻上下文的——最后还有个裁判AI给翻译打分、写评语,系统根据评语自动调整每个角色的工作方式,不用重新训练模型。在15种语言方向上都拿了最高分,比最强对手平均少扣6.9%的分。它不是你明天就能用上的工具,但说明AI正在从『翻得准』走向『翻得懂』——懂一部剧的来龙去脉。

📄 原文摘要(英文)

Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On the public MuSC benchmark, SMART obtains the best model result across all four language pairs and also achieves the best human-evaluation result, with an overall score of 4.50/5.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新