Return
Towards few-shot deepfake detection with an enhanced CLIP model
Y
Z
Y
G
DOI:10.1016/j.neunet.2026.109027.png)
Abstract
En 中文
The proliferation of deepfakes threatens privacy, public trust, and democratic integrity, necessitating robust detection methods effective in practical settings. Current deepfake detectors focus on generalizing to unseen forgery methods through supervised learning, requiring extensive training data. As forgery generation techniques diversify and often bear little relation to one another, cross-technique generalization becomes increasingly difficult. Recent advances in few-shot learning show promising yet limited potential for deepfake detection, as their performance has not fully translated from conventional image classification tasks to this specific domain. To overcome these challenges, we propose a three-pronged methodology called Instance-level Few-shot Prompt Learning (IFPL) under the CLIP framework. In the text branch, we replace manual templates with a multi-scale adaptive context and enhance feature extraction by incorporating Instance-level facial embeddings. In the visual branch, we introduce learnable visual perturbation blocks per input sample to guide the encoder toward forgery-specific artifacts. Finally, we introduce a nonparametric prototype-based instance cache module to provide external guidance to CLIP for robust decisions. Experiments demonstrate that IFPL consistently surpasses state-of-the-art methods under few-shot conditions, establishing a novel solution for fine-grained visual tasks with minimal samples.
Keywords:
Deepfake detections
Few-shot learning
CLIP model
Prompt learning
Journal
IF:
6.3
Papers:
7.7K
Citations:
3.0W
