Hugging Face Trending Papers

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud. Although prior work has explored vision-language model (VLM)-based synthetic image detection, these evaluations typically consider images in isolation.

arXiv AI
Sep 18

FORGE: Forensic Reasoning with Grounded Evidence

FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.

By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury
arXiv AI
4d ago

A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

arXiv:2609.37783v1 Announce Type: cross Abstract: Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not...

By Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin, Wentao Zhang, Jacquelyn Burkell, Yuntian Deng, Karen Eltis, Maura R. Grossman, Vered Shwartz, Ebrahim Bagheri
arXiv Computer Vision
Sep 16

Unifying Semantic Priors and High-Frequency Traces: Enhancing V-JEPA with Mixture-of-Experts for Robust Synthetic Image Forensics

The paper introduces MoE-JEPA, a dual‑stream deepfake detection model that combines a V‑JEPA backbone with a Residual Mixture‑of‑Experts mechanism and a noise stream branch. It further incorporates a Gated Attention Multiple Instance Learning module to refine spatial semantic understanding. On the SID‑Set benchmark, MoE‑JEPA achieves a new state‑of‑the‑art accuracy of 95.54%, outperforming much larger models.

By Simone Teglia, Irene Amerini
Hugging Face Trending Papers
Aug 19

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by tiny, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal interpretations. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these small perturbations can dramatically alter the models’ textual outputs.

arXiv AI
Jun 2

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

arXiv:2606. 00101v1 Announce Type: cross Abstract: With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security.

By Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng
arXiv AI
Aug 20

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by small, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal alignment. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these perturbations can significantly alter the models’ textual outputs.

By Ilan Zini, Boussad Addad, Katarzyna Kapusta