arXiv Computer Vision

Certification of Real Images through Calibrated Content Authentication

The paper evaluates twenty deepfake detectors against ten recent generators, finding that accuracy has dropped from 99.5% to 76% and that adversarial perturbations can reduce detector performance to below 2%. It proposes a new detection paradigm that certifies authenticity only when a faithful reconstruction by a known generator is impossible, and shows that this calibrated detector can limit false certifications to at most 1% while remaining robust to bounded‑perturbation attacks. The study also highlights the erosion of post‑hoc verifiability, noting that newer generators can reproduce a larger fraction of previously unreplicable images.

Hugging Face Trending Papers
5d ago

Certification of Real Images through Calibrated Content Authentication

The paper examines the reliability of deepfake detectors, noting a decline in accuracy from 99.5% to 76% over four years and a drastic drop below 2% when adversarial perturbations are applied. It argues that content alone cannot determine provenance because generators can perfectly reproduce authentic images, so a new detection approach is proposed that certifies authenticity only if a generator cannot faithfully reconstruct the content. The authors demonstrate that their calibrated detector can limit false certifications to 1% while maintaining robustness against bounded‑perturbation attacks, though it struggles with arbitrary transformations.

arXiv Machine Learning
Sep 11

CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration

The paper introduces CertDW, a certified dataset watermark and ownership verification method that remains reliable even under malicious perturbations. By leveraging conformal prediction, it defines two statistical measures—principal probability (PP) and watermark robustness (WR)—to evaluate model stability on benign versus watermarked samples. The authors derive certification conditions linking WR to a PP-based threshold and provide a high‑probability bound on false positives, enabling robust ownership verification when a suspicious model’s WR exceeds the PP values of benign models.

By Ting Qiao, Yiming Li, Jianbin Li, Yingjia Wang, Leyi Qi, Junfeng Guo, Ruili Feng, Dacheng Tao
arXiv AI
Aug 14

SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

arXiv:2608. 12876v1 Announce Type: cross Abstract: Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving.

By Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan
arXiv AI
Aug 24

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization proposes a new framework that improves deepfake detection and interpretability. The approach introduces Feature-robust Augmentation—diversified degradation-aware strategies combined with supervised contrastive learning and a mean-teacher architecture—to maintain accuracy on low-quality images. For explanations, it employs evidence-grounded preference optimization, guiding the model to focus on genuine manipulation traces by learning from chosen-rejected explanation pairs that omit evidence or inject irrelevant details. The method achieved first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge and is publicly available on GitHub.

By Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu