Specialist-Generalist Fusion with Outcome-Supervised Rationales for Deepfake Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.
Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization proposes a new framework that improves deepfake detection and interpretability. The approach introduces Feature-robust Augmentation—diversified degradation-aware strategies combined with supervised contrastive learning and a mean-teacher architecture—to maintain accuracy on low-quality images. For explanations, it employs evidence-grounded preference optimization, guiding the model to focus on genuine manipulation traces by learning from chosen-rejected explanation pairs that omit evidence or inject irrelevant details. The method achieved first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge and is publicly available on GitHub.
arXiv:2608. 06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety.
arXiv:2610.08639v1 Announce Type: new Abstract: As generated images become increasingly realistic, reliable forgery detection is essential for maintaining trust in visual information. However, existi...
The MSU team presents a modular approach to the Explainable Deepfake Detection Challenge, combining multiple DINOv3 backbones with Mesorch manipulation-localization features for real/fake classification. They incorporate a Grounding-DINO-based pseudo-mask pipeline to generate artifact evidence maps and a local contrastive objective to separate artifact from authenticity cues. For explanations, class-conditional Qwen3-VL models produce complex descriptions, which are then simplified by a GRPO-optimized text model, achieving high detection and explanation scores on the XPlainVerse dataset.
As generated images become increasingly realistic, reliable forgery detection is essential for maintaining trust in visual information. However, existing methods primarily rely on task-specific superv...