搜图终于能说「要这件但换个颜色」了
你搜「找这件衬衫,但要粉色的」,现有AI要么把整句话揉成一团数字,要么自由发挥漏掉「粉色」这个关键约束。这篇把搜图拆成六类语义槽——哪些东西必须保留、哪些属性要改、哪些背景可以忽略——然后像打分表一样逐条核对。在五个评测集上,不经过训练就超过了所有现有方法,而且还能把这套规则教给一个小模型,保留九成以上能力。它不是你明天就能用的产品,但指明了方向:搜图不该靠猜,而该靠清单。
📄 原文摘要(英文)
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.