arXiv:2607. 24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.
By Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin
arXiv:2607. 25266v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible.
By Ghazal Kaviani, Ghassan AlRegib
arXiv:2609.38900v1 Announce Type: new
Abstract: Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons,...
By Yinying Li, Yuqian Fu, Yulin Dai, Jingyu Gong, Tianwen Qian, Xiaoling Wang
arXiv:2609.37426v1 Announce Type: cross
Abstract: Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture....
By Arka Mukherjee, Kaleen Shrestha, Larissa Zhu, Maja Matari\'c
arXiv:2610.01192v1 Announce Type: new
Abstract: Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between...
By Yi Chen, MingMing Yu, Rui-Qi Wang, Boran Wang, Xiaohang Cao, Chu Tang, Jingmin Chen, Jie Gu
arXiv:2609.40195v1 Announce Type: cross
Abstract: Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spann...
By Guangzhi Xiong, Xinyuan Zhang, Xiao Yang, Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun, Lucas Liang, Guangqiang Dong, Ejaz Ahmed, Ahmed A Aly, Anuj Kumar, Raffay Hamid, Aidong Zhang, Xin Luna Dong