arXiv Machine Learning

CHROMA: Detecting AI-Generated Images through Inter-Channel Color-Space Correlations

arXiv:2606. 08864v1 Announce Type: cross Abstract: The rapid adoption of diffusion and large-scale generative models has made it increasingly challenging to distinguish synthetic imagery from real photographs.

arXiv AI
Jun 2

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

arXiv:2606. 00101v1 Announce Type: cross Abstract: With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security.

By Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng
arXiv Computer Vision
Aug 31

FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization

FUSED is a new framework that jointly detects and localizes AI-generated inpainting by combining low-level forensic cues with high-level semantic features through a sparsely-gated Mixture-of-Experts architecture. It predicts both an image-level manipulation score and a pixel-level mask of the inpainted region. On the OpenSDID cross-generator benchmark, FUSED outperforms existing methods, especially on unseen generators, and transfers effectively to the AutoSplice and CocoGlide benchmarks, doubling localization performance.

By Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska
arXiv Computer Vision
Sep 16

Unifying Semantic Priors and High-Frequency Traces: Enhancing V-JEPA with Mixture-of-Experts for Robust Synthetic Image Forensics

The paper introduces MoE-JEPA, a dual‑stream deepfake detection model that combines a V‑JEPA backbone with a Residual Mixture‑of‑Experts mechanism and a noise stream branch. It further incorporates a Gated Attention Multiple Instance Learning module to refine spatial semantic understanding. On the SID‑Set benchmark, MoE‑JEPA achieves a new state‑of‑the‑art accuracy of 95.54%, outperforming much larger models.

By Simone Teglia, Irene Amerini
arXiv Computer Vision
4d ago

RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection

The paper introduces RED (Reconstruction Evolution Dynamics), a new framework for detecting AI-generated images that leverages the evolution of intermediate reconstruction stages rather than relying solely on static representations or endpoint discrepancies. RED uses a frozen multiscale VQ‑VAE and a frozen CLIP encoder to capture a reconstruction trajectory, then learns image‑adaptive stage weights from token negative log‑likelihoods provided by a frozen VAR model. Experiments on six benchmarks show RED achieves the highest average accuracy (92.5%) and precision (97.5%) among evaluated methods, and it remains robust to common image degradations.

By Wenpeng Mu, Junshan Jin, Tanfeng Sun, Xinghao Jiang, Qiang Xu
arXiv Computer Vision
Sep 11

A Multi-View and Confusion-Guided Ensemble Framework for Robust Synthetic Image Attribution

The paper introduces a multi‑view, confusion‑guided ensemble framework for synthetic image attribution, combining FFT‑ConvNeXt, DINOv2, CLIP, and Xception to capture frequency, semantic, and forensic cues. Extensive data augmentation simulates realistic post‑processing, while a binary expert classifier and class‑adaptive confidence calibration address ambiguities between similar diffusion models. The approach achieved 99.53% on the public leaderboard and 99.20% on the private leaderboard for the ICANN 2026 DLMMDD Workshop challenge.

By Zuomin Qu
arXiv Computer Vision
6d ago

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.

By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv AI
Sep 25

EIB-Net: Entropy-Guided Information Bottleneck for Generalizable AI-Generated Image Detection

EIB-Net is an Entropy‑Guided Information Bottleneck Network designed to detect AI‑generated images across diverse generative models. It introduces an Image Entropy metric to automatically select the most informative, low‑entropy patch and applies a Variational Information Bottleneck to learn compact, generalizable features. Experiments on DIFF, DiffusionForensics, and GenImage benchmarks show state‑of‑the‑art performance, achieving 85.7% accuracy with only 2% of training data and maintaining robust cross‑generator generalization.

By Zhida Zhang, Xinlei Ma, Jie Cao