[Submitted on 22 Feb 2026 (v1), last revised 25 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Image-generative models are widely deployed across industries. Recent studies show that they can be exploited to produce unacceptable content. Existing mitigation strategies rely on prompt filtering and safety-aware training, both of which can be bypassed and often degrade generative quality. In this work, we propose ReVision, a training-free, prompt-based, post-hoc safety framework for image-generation pipeline. ReVision acts as a post-generation safeguard by analyzing generated images and selectively editing unsafe concepts without altering the underlying generator. Prior post-hoc editing methods often rely on imprecise spatial localization, limiting deployability, in multi-concept scenes. To address this limitation, ReVision introduces a VLM-assisted spatial gating mechanism for instance-consistent localization, enabling integrity-preserving edits. We introduce an 800-image benchmark spanning single- and multi-unsafe-concept images, each composed alongside benign concepts in shared scenes. On this benchmark, ReVision improves CLIP alignment toward safe prompts by +0.121, reduces multi-concept background LPIPS from 0.166 to 0.058, and eliminates NudeNet detections (70.51 -> 0). Across external benchmarks, ReVision outperforms prior methods, and a human study shows it reduces recognizability of unacceptable content from 96% to 10%.

Submission history

From: Gurjot Singh [view email]
[v1] Sun, 22 Feb 2026 12:30:01 UTC (28,727 KB)
[v2] Mon, 2 Mar 2026 07:13:22 UTC (28,727 KB)
[v3] Sat, 25 Jul 2026 04:40:07 UTC (5,031 KB)