arXiv AI

DetectZoo: A Unified Toolkit for AI-Generated Content Detection Across Text, Audio, and Image Modalities

arXiv:2606. 04205v1 Announce Type: cross Abstract: The growing popularity and capacity of generative models have eroded the distinction between human and machine-generated content, motivating a growing body of work on detection across text, images, and audio.

arXiv AI
Jun 29

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook

arXiv:2411. 19537v2 Announce Type: replace-cross Abstract: We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content.

By Florinel-Alin Croitoru, Andrei-Iulian Hiji, Vlad Hondru, Nicolae Catalin Ristea, Paul Irofti, Marius Popescu, Cristian Rusu, Radu Tudor Ionescu, Fahad Shahbaz Khan, Mubarak Shah
arXiv AI
Aug 11

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.

By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
arXiv AI
2d ago

MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation

arXiv:2505.17613v2 Announce Type: replace Abstract: Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation...

By Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu
arXiv AI
2d ago

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.

By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
arXiv AI
Jun 2

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

arXiv:2606. 00101v1 Announce Type: cross Abstract: With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security.

By Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng
arXiv Machine Learning
Jun 16

MVEB: Massive Video Embedding Benchmark

arXiv:2606. 14958v1 Announce Type: cross Abstract: We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric question answering.

By Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao, Deep Shah, Manan Dey, Shriya Sudhakar, Zacharie Bugaud, Wissam Siblini, Ayush Sunil Munot, Yashwanth Devavarapu, Rakshitha Ireddi, Michelle Yang, M\'arton Kardos, Niklas Muennighoff, Kenneth Enevoldsen