arXiv Computer Vision
Sep 1

ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

ClearText-Video (CTVid) is a large-scale, scene-text-aware benchmark that examines text-centric video understanding under varying quality conditions. It comprises 4,639 real-world egocentric videos, over 550,000 frames, 1.6 million human-verified scene-text annotations, and more than 220,000 spatial/temporal question–answer pairs in Chinese and English. For each high-quality video, CTVid provides matched degraded- and restored-quality variants, enabling studies of Text-Centric Video Restoration and Multi-Quality VideoQA, and revealing that visual enhancement does not always improve textual fidelity or downstream reasoning.

By Jinlong Li, Jiaming Ding, Dingfu Lu, Malcolm Hsiu, Chuang Ke, Kangning Yang, Bochen Guan, Lan Fu, Jie Cai, Huiming Sun, Zibo Meng
arXiv Computer Vision
Aug 25

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

arXiv:2607.14935v2 Announce Type: replace Abstract: Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world application...

By Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
arXiv AI
Jun 9

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

arXiv:2606. 08063v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real-world visual corruptions.

By Jiaqi Tang, Jianmin Chen, Youyang Zhai, Wei Wei, Runtao Liu, Mengjie Zhao, Xiangyu Wu, Qingfa Xiao, Qifeng Chen