arXiv Machine Learning By Xinjie Cui, Yuezun Li, Delong Zhu, Jiaran Zhou, Junyu Dong, Siwei Lyu

Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection

Read the original on arXiv Machine Learning →

arXiv:2411. 19715v4 Announce Type: replace-cross Abstract: We describe Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 17

Generalizable Face Forgery Detection via Separable Prompt Learning

The paper introduces Separable Prompt Learning (SePL), a method that enhances face forgery detection by leveraging CLIP’s textual encoder through two separate learnable prompts. SePL incorporates a cross-modality alignment strategy and specific objectives to distill forgery knowledge from CLIP. Experiments show that SePL outperforms existing approaches in cross-dataset and cross-method evaluations.

By Enrui Yang, Baoyuan Wu, Yuezun Li
arXiv Computer Vision
Sep 18

IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models

The paper introduces IMFD, an end‑to‑end multi‑face forgery detector that uses instruction‑based Large Vision‑Language Models (LVLMs). IMFD jointly localizes faces and predicts forgery labels in a single stage, explicitly incorporating predicted face bounding boxes into the textual instruction to improve grounding and detection. Experiments on converted multi‑face forgery datasets show that IMFD outperforms several state‑of‑the‑art methods.

By Dasom Choi, Sangjun Moon, Hyeongchan Im, Jaeeon Park, Jingun Kwon, Hidetaka Kamigaito, Taro Watanabe, Manabu Okumura
arXiv AI
Sep 18

FORGE: Forensic Reasoning with Grounded Evidence

FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.

By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury