AI Pulse
📄 论文解读

AI画图跑得对,但画得不对

AI 写代码画图,代码能跑通,画出来的东西却可能完全不是你要的——构图、外观、运动全不对。研究者把这叫「程序到视觉的落差」,并造了个测试框架,让 11 个顶级 AI 模型按指令画图和生成视频,再逐项检查结果。结果最强模型 GPT-6-Astra 虽然任务全跑通,但只有 96% 的图片和 76.9% 的视频真正达标。更意外的是,通用能力评分相近的模型,视觉表现可能差很远——也就是说,你平时看 AI 多聪明,跟它能不能画对,是两码事。这套框架把「画图过程」拆成可追踪的步骤,方便定位问题出在哪一步。它不是你明天就能用的工具,但给「AI 画图到底靠不靠谱」提供了一份更诚实的体检报告。

📄 原文摘要(英文)

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新