arXiv:2608.21837v1 Announce Type: new
Abstract: Bitstream-corrupted Harsh Visual Understanding (BcHVU) aims to understand harshly degraded videos originally decoded from a severely corrupted bitstrea...
By Chaoran Huang, Fangcheng Li, Tianyi Liu, Wenyang Liu, Kejun Wu
arXiv:2610.08414v1 Announce Type: new
Abstract: Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded...
By Zhen Yu, Wenyang Liu, Kejun Wu, Chengwang Xiao, Renjie Qiao, Chengtao Cai
arXiv:2608.29212v1 Announce Type: cross
Abstract: Existing video watermarking systems are symmetric: the party that can verify a mark holds the extractor weights or generator secret and can therefore...
By Guang Yang, Fengchen Liu
arXiv:2609.39623v1 Announce Type: new
Abstract: The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images wh...
By Yoonseo Kim, Seungwoo Baek, Junyoung Park
COVER is a new video watermarking method that targets codec compression as its primary design goal. It embeds the watermark payload in the latent space of a frozen generative video autoencoder and recovers it by re‑encoding the received video into the same latent space. Using a differentiable codec surrogate bank, COVER achieves high bit accuracy across multiple codecs while keeping marked videos visually close to the originals.
By Yuxin Cao, Hao Yang, Ziqi Ding, Jie Hao, Wei Song
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.
By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li
arXiv:2606. 08063v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real-world visual corruptions.
By Jiaqi Tang, Jianmin Chen, Youyang Zhai, Wei Wei, Runtao Liu, Mengjie Zhao, Xiangyu Wu, Qingfa Xiao, Qifeng Chen
arXiv:2605.12006v2 Announce Type: replace
Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deploymen...
By Sohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler, Christos Sakaridis, Suha Kwak
arXiv:2606. 02120v1 Announce Type: cross Abstract: In this report, we address the problem of determining whether a user performs an action incorrectly from egocentric video data.
By Boyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang, Ruochen Cui, Qingming Huang
The paper proposes a unified forensics framework that extends traditional binary image manipulation detection to a multiclass setting—distinguishing real, fully synthetic, and tampered images. It adds a segmentation branch for pixel‑level localization of tampered regions, achieving higher classification accuracy and IoU scores compared to recent benchmarks. The authors provide the implementation on GitHub for reproducibility.
By Annalisa Gallina, Marco Fiorucci, Marco Brigo, Federica Battisti, Lamberto Ballan
arXiv:2609.31558v1 Announce Type: new
Abstract: Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities....
By Ahmed Abdelnaby, Mohamed Elmahallawy