AI Pulse
📄 论文解读

给AI换副新骨架,比换大脑更管用

大模型的能力不只看模型本身,还看外面那层“骨架”——记忆怎么管、计划怎么定、工具怎么调。这篇论文训练了一个专门的AI,能根据任务现场给任何现成大模型配一副新骨架,不用改模型本身。结果:一个中端模型配上它,在两项测试上超过了目前最强的GPT-5.6;另一个本来就很强的模型,最多涨了20分。也就是说,换骨架这件事本身可以训练、可以迁移、还能越用越强,跟堆模型参数是两条独立的路。它不是你明天能用上的东西,但它暗示了一个方向:未来AI的竞争,可能不只是比谁的大脑大,还比谁更会给自己搭脚手架。

📄 原文摘要(英文)

Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新