AI 训练缺好题?这篇把终端任务拆成三件套再拼回去
训练一个能操作终端的 AI 不难,难的是给它足够多、足够靠谱的练习题。每道题得同时包含:人话指令、初始环境、标准答案、自动判分器——四样东西如果各写各的,很容易出现「指令说删文件,环境里根本没那个文件」的乌龙。FACET 的做法是:先从一个真实技能场景出发,把四样东西绑在同一个容器环境里生成,谁出问题就修谁,不重写整个题。最终产出的题有密集的可执行检查点,拿这些题去微调小模型,在 Terminal-Bench 2.1 上效果明显好于其他造题方式。它不是你明天能用上的,但它是那种「让 AI 自己学会用电脑」的基础设施级进步。
📄 原文摘要(英文)
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. We present FACET (Fine-grained Agentic Construction of Executable Tasks), a framework that addresses both information preservation and cross-artifact consistency. FACET reconstructs related agent skills into coherent, information-rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution-based validation and targeted repair correct artifact-specific failures without unnecessarily regenerating valid components. FACET produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data-efficient supervision. Fine-tuning models across multiple scales consistently improves performance on Terminal-Bench 2.1, while analyses of alternative generation schemes support the importance of environment-grounded construction for task validity and solution-verifier alignment. These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis.