The paper introduces Separable Prompt Learning (SePL), a method that enhances face forgery detection by leveraging CLIP’s textual encoder through two separate learnable prompts. SePL incorporates a cross-modality alignment strategy and specific objectives to distill forgery knowledge from CLIP. Experiments show that SePL outperforms existing approaches in cross-dataset and cross-method evaluations.
By Enrui Yang, Baoyuan Wu, Yuezun Li
The paper introduces IMFD, an end‑to‑end multi‑face forgery detector that uses instruction‑based Large Vision‑Language Models (LVLMs). IMFD jointly localizes faces and predicts forgery labels in a single stage, explicitly incorporating predicted face bounding boxes into the textual instruction to improve grounding and detection. Experiments on converted multi‑face forgery datasets show that IMFD outperforms several state‑of‑the‑art methods.
By Dasom Choi, Sangjun Moon, Hyeongchan Im, Jaeeon Park, Jingun Kwon, Hidetaka Kamigaito, Taro Watanabe, Manabu Okumura
FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.
By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury
arXiv:2608. 06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety.
By Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
arXiv:2608.17351v2 Announce Type: replace
Abstract: Open-world face anti-spoofing must address both covariate and semantic shifts: source and target domains differ in imaging conditions, while target...
By Fangling Jiang, Qi Li, Bing Liu, Weining Wang, Quilin Huang, Zhenan Sun, Ming-Hsuan Yang
arXiv:2609.01511v1 Announce Type: new
Abstract: Face forgery detectors often achieve strong results on controlled benchmarks, but their reliability under realistic image degradations remains limited....
By Lucas Cunha, Lucas Sotomaior, Lucas Gasperin, Beatriz Caldas, Eduardo Pianovski, Rayson Laroca
arXiv:2605. 09089v2 Announce Type: replace-cross Abstract: Digital onboarding and eKYC systems used by banks, fintech platforms, telecom providers, and other third-party services commonly verify users by comparing an uploaded identity document with a selfie or live facial capture.
By Abhishek Kumar, Riya Tapwal, Carsten Maple, Mark Hooper
arXiv:2609.36145v1 Announce Type: new
Abstract: Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective...
By Kaiqing Lin, Songze Li, Shen Chen, Yunfei Guo, Xiaoye Qiu, Haodong Li, Taiping Yao, Bo Wang, Youchang Xiao, Bin Li, Shouhong Ding
The paper introduces Band-Attention Modulation Network (BAM‑Net), a face forgery detection framework that learns fine‑grained, adaptive modulation of frequency bands in the Discrete Cosine Transform spectrogram. BAM‑Net dynamically reweights anti‑diagonal frequency bands to enhance forgery‑related spectral cues while suppressing irrelevant information, then fuses this modulated frequency data with spatial features using a lightweight backbone with distance‑decayed attention. Experiments on FaceForensics++, Celeb‑DF, and DFDC show that BAM‑Net achieves state‑of‑the‑art performance and strong generalization across datasets, compression levels, and manipulation types.
By Zhida Zhang, Wenkui Yang, Xinlei Ma, Qihang Fan, Jie Cao
arXiv:2607. 16273v1 Announce Type: cross Abstract: In forensic environments, automated identification of perpetrators is difficult due to pose changes, changes in light, occlusion, and lack of labeled data.
By Savitha N J, Lata B T
The paper introduces VeriFi, a watermarking framework that protects face images from AI‑generated manipulation. It embeds a compact semantic latent watermark to preserve content, localizes pixel‑level edits without explicit payloads, and simulates realistic deepfake attacks to improve robustness. Experiments on CelebA‑HQ and FFHQ show that VeriFi outperforms existing methods in robustness, localization accuracy, and recovery quality.
By Peipeng Yu, Jinfeng Xie, Chengfu Ou, Xiaoyu Zhou, Jianwei Fei, Yunshu Dai, Zhihua Xia, Chip Hong Chang
arXiv:2608.23984v1 Announce Type: new
Abstract: Recent advances in single-image 3D Gaussian head reconstruction have enabled highly realistic and freely renderable digital heads from a single portrai...
By Yujie Gao, Zijian Yu, Yan Hong, Jun Lan, Jianfu Zhang