arXiv Computer Vision By Taehoon Kim, Jongwook Choi, Heejae Jo, Byungmin Park, Jongwon Choi

Preserving Knowledge across Space and Time for Continual Video Deepfake Detection

Read the original on arXiv Computer Vision →

The paper introduces Modality‑Specific Frequency Distillation (MSFD), a continual learning framework for video deepfake detection that separates spatial, temporal, and spatiotemporal features in the frequency domain. By preserving each modality independently and applying a cross‑modality decorrelation loss, MSFD adapts to new forgery patterns while maintaining performance across diverse continual deepfake video scenarios. Experiments demonstrate that this approach outperforms state‑of‑the‑art methods in both adaptation and retention.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jul 16

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs

arXiv:2605. 16366v2 Announce Type: replace-cross Abstract: Video MLLMs face a persistent tension between spatial fidelity and temporal coverage: preserving fine-grained visual details requires many spatial tokens, while capturing short-lived events requires dense temporal sampling.

By Yigui Feng (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Qinglin Wang (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Yang Liu (The Shien-Ming Wu School of Intelligent Engineering, South China University of Technology, Guangzhou, Guangdong, China), Jie Liu (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China)
Hugging Face Trending Papers
Jul 8

ASFR-Net: Adversarial Alignment and Spatio-Frequency Refinement Network for Heterogeneous Remote Sensing Image Change Detection

The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significant modal disparities caused by distinct imaging mechanisms. These intrinsic inconsistencies are prone to introducing pseudo-changes, thereby constraining detection accuracy.

arXiv AI
Sep 2

Taming Modality Entanglement in Continual Audio-Visual Segmentation

The paper introduces a new Continual Audio‑Visual Segmentation (CAVS) task that enables continuous segmentation of new classes guided by audio. It identifies two key challenges—multi‑modal semantic drift and co‑occurrence confusion—and proposes a Collision‑based Multi‑modal Rehearsal (CMR) framework with Multi‑modal Sample Selection (MSS) and Collision‑based Sample Rehearsal (CSR) strategies to address them. Experiments on three audio‑visual incremental scenarios show that CMR outperforms single‑modal continual learning methods.

By Yuyang Hong, Qi Yang, Tao Zhang, Zili Wang, Zhaojin Fu, Kun Ding, Bin Fan, Shiming Xiang
arXiv AI
Jun 29

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook

arXiv:2411. 19537v2 Announce Type: replace-cross Abstract: We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content.

By Florinel-Alin Croitoru, Andrei-Iulian Hiji, Vlad Hondru, Nicolae Catalin Ristea, Paul Irofti, Marius Popescu, Cristian Rusu, Radu Tudor Ionescu, Fahad Shahbaz Khan, Mubarak Shah