AI Pulse
📄 论文解读

改图终于能同时听懂你的话和你的手势了

以前改图要么靠文字描述(“把猫往左移”),但位置总不准;要么靠鼠标拖拽(“把这里拖到那里”),但AI可能误解你的意图。这篇把两者合在一起:你既说“把杯子放桌上”,又用鼠标点一下杯子和桌子的位置,AI同时理解语义和空间,改出来的图既符合你的话,位置也精确。研究者从视频中提取了2.3万组“文字+拖拽点+原图+结果”的数据来训练,效果比纯文字或纯拖拽都好。做设计、修图、电商产品图的人,这是你明天就能用上的那种——不用再反复调位置了。

📄 原文摘要(英文)

Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spatial control of the editing results. In contrast, visual prompts such as drag and point can provide precise spatial guidance, but are limited by the inherent ambiguity in semantic intent. To unify the strength of textual and visual prompts, we present Text-Vision Co-Instructed Image Editing, which jointly models textual instructions as semantic intent and sparse visual instructions as spatial guidance, aiming to achieve precise and intent-faithful image manipulation. To this end, we first construct a textual-visual instruction paired dataset with more than 23K samples derived from dynamic videos, enabling aligned supervision for cross-modal instruction. We then propose TV-Edit, a Textual-Visual instruction unified Editing framework to contextualize drag or point-based visual instructions with image-text semantics and lift them into semantic-aware control representations for pretrained editing backbones. By integrating semantic intent and spatial constraints, TV-Edit leads to more precise spatial control, less instruction ambiguity, and stronger structural consistency than text-only or drag-based alternatives. Finally, we establish TV-Edit-Bench, a deliberately designed benchmark to evaluate semantic faithfulness, spatial alignment, and visual consistency with ground-truth references and controlled textual-visual variations for reliable assessment. Our experiments across multiple editing backbones demonstrate that TV-Edit consistently yields more precise and intent-faithful edits, significantly outperforming state-of-the-art instruction-based and drag-based baselines.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新