Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
arXiv:2607. 25543v1 Announce Type: cross Abstract: Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes.
arXiv:2607. 25543v1 Announce Type: cross Abstract: Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes.
arXiv:2606. 00101v1 Announce Type: cross Abstract: With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security.
arXiv:2609.12668v1 Announce Type: new Abstract: Recent deepfake detection studies increasingly suggest remote photoplethysmography (rPPG) signals as an authenticity cue. However, existing benchmarks...
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
arXiv:2608. 06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety.
arXiv:2609.25775v1 Announce Type: new Abstract: Recent advances in generative video models have enabled the synthesis of visually realistic content, posing significant challenges to synthetic video d...
ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.
arXiv:2509.06422v2 Announce Type: replace Abstract: Video camouflaged object detection (VCOD) is challenging due to dynamic environments. Existing methods face two main issues: (1) SAM-based methods...
arXiv:2602.14633v3 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categori...
arXiv:2609.07670v1 Announce Type: cross Abstract: The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, dee...
The paper introduces a video forgery detection system that fuses spatial and frequency-domain features using a ResNet‑LSTM backbone with a Convolutional Block Attention Module (CBAM) and a Discrete Cosine Transform (DCT) module. Experiments on multiple benchmark datasets show that this hybrid architecture outperforms existing methods in distinguishing authentic from manipulated videos. Ablation and comparative studies confirm the individual contributions of each component, highlighting the model’s effectiveness across diverse forgery scenarios.
RoES is a Rotational Equivariant Selective-frequency fusion network that dynamically separates low- and high-frequency components of infrared-visible images. It uses a trainable rotation-enhanced updater to decouple frequencies, then fuses them with a dual-branch module: a rotation-equivariant Mamba for low-frequency structural dependencies and a polar spectral attention Dual-Fourier block for high-frequency detail refinement. Experiments show RoES outperforms existing methods in fusion quality and downstream object detection, offering a robust multimodal fusion solution.