[Submitted on 29 Nov 2024 (v1), last revised 24 Jul 2026 (this version, v4)]
Abstract:We describe Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector. Although CLIP is highly versatile, adapting it for face forgery detection is non-trivial as forgery-related knowledge is entangled with a wide range of unrelated knowledge. Existing methods treat CLIP merely as a feature extractor, lacking task-specific adaptation, which limits their effectiveness. To address this, we introduce an adapter to learn face forgery traces -- the blending boundaries unique to forged faces, guided by task-specific objectives. Then we enhance the CLIP visual tokens with a dedicated interaction strategy that communicates knowledge across CLIP and the adapter. Since the adapter is alongside CLIP, its versatility is highly retained, naturally ensuring strong generalizability in face forgery detection. {With only $\textbf{5.7M}$ trainable parameters, our method achieves superior performance across six standard datasets.} Additionally, we describe Forensics Adapter++, an extended method that incorporates textual modality via a newly proposed forgery-aware prompt learning strategy. This extension leads to a further $\textbf{1.3\%}$ performance boost over the original Forensics Adapter. We believe the proposed methods can serve as a baseline for future CLIP-based face forgery detection methods. The code has been released at this https URL.
Submission history
From: Yuezun Li [view email]
[v1]
Fri, 29 Nov 2024 14:02:11 UTC (16,334 KB)
[v2]
Mon, 24 Mar 2025 09:41:55 UTC (16,337 KB)
[v3]
Fri, 23 May 2025 16:14:40 UTC (12,169 KB)
[v4]
Fri, 24 Jul 2026 08:17:19 UTC (6,415 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.