AI Pulse
📄 论文解读

给AI画图加个“安全词”:不碰模型,只改提示词

现在的AI画图工具,你输入“一个穿比基尼的女人”它可能直接给你生成露骨图,哪怕这句话本身没毛病。研究者发现,问题出在模型学到的数据分布上——有些词组合在一起,就会触发它“画歪”。他们做了个叫DiSCO的插件,不碰模型内部,只在你输入的提示词后面悄悄加一串“安全后缀”,然后让模型自己生成几张图对比,挑出最不露骨的那版。在多个模型上测试,能把有害图生成率降低37.7%,而且画出来的图还是你要的那个意思。它不是你明天就能装上的工具,但给那些想给AI画图加护栏的公司指了条路:不用改模型,改提示词就行。

📄 原文摘要(英文)

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the benign adversarial problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新