RT-NeuS is a neuro‑symbolic framework for long‑form video question answering that retains the accuracy and formal guarantees of temporal‑logic‑guided methods while dramatically reducing inference latency. It achieves this by using coarse‑to‑fine adaptive sampling to focus on query‑relevant frames and batched proposition detection with KV‑cache reuse, enabling all propositions to be evaluated in a single forward pass. Experiments on LongVideoBench, Video‑MME, and MLVU show up to a 13× speed‑up on an NVIDIA H200 GPU while matching or surpassing prior neuro‑symbolic accuracy.
By Shawn Liang, Sahil Shah, Chengwei Zhou, S P Sharan, Harsh Goel, Arnab Sanyal, Sandeep Chinchali, Gourav Datta
arXiv:2602. 01801v2 Announce Type: replace-cross Abstract: Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines.
By Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari
arXiv:2610.01192v1 Announce Type: new
Abstract: Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between...
By Yi Chen, MingMing Yu, Rui-Qi Wang, Boran Wang, Xiaohang Cao, Chu Tang, Jingmin Chen, Jie Gu
StreamTTT is a streaming vision-language model that balances real-time perception with long-term memory by writing long-range history into fast weights outside the attention context, while keeping a short sliding key-value cache for recent evidence. The model is trained on both offline long-video QA and a new real-time QA corpus, and it outperforms SimpleStream-4B on OVO-Bench by 1.4 points in real-time perception and 3.7 points in backward tracing. StreamTTT-4B also competes with the larger SimpleStream-8B on the StreamingBench Real-Time Visual Understanding subset.
By Joya Chen, Zeyun Zhong, Mike Zheng Shou
arXiv:2606. 17798v1 Announce Type: cross Abstract: Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-horizon contextual memory.
By Zhenyu Yang, Kairui Zhang, Bing Wang, Shengsheng Qian, Changsheng Xu
ShallowStream is a framework for streaming video understanding that uses the shallow layers of a multimodal large language model (MLLM) to encode frames and build a lightweight index. During streaming, it maintains an always‑on index via the KV cache of shallow layers, and at query time it scores context frames using shallow‑layer attention and selects diverse evidence for answering. The approach matches the performance of leading streaming methods while cutting per‑frame prefill latency and 10‑second end‑to‑end latency by up to 52.1× and 11.9×, respectively.
By Jitai Hao, Ke Yang, Qiang Huang, Jun Yu
arXiv:2606. 06991v1 Announce Type: cross Abstract: Online Video Large Language Models (Video-LLMs) have advanced toward seamless human-AI interaction through frame-by-frame processing and proactive responding.
By Zhenyu Yang, Kairui Zhang, Shengsheng Qian, Weiming Dong, Changsheng Xu
arXiv:2609.23601v1 Announce Type: new
Abstract: Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens,...
By Siru Zhong, Qiongyan Wang, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
VisCache introduces a two-stage, plug‑and‑play framework for pruning visual key‑value caches in Vision Large Language Models without retraining. The first stage filters out temporally redundant keyframes, while the second stage, PruneKV, applies a parabolic layer‑wise budget and asymmetric update to selectively prune keys and fuse values, preserving essential context. Experiments show up to 2.35× speedup and significant memory savings with only 19–28% of the original cache retained, outperforming existing baselines.
By Lyuke Wang, Zhuo Li, Guangxu Zhu
arXiv:2609.00291v1 Announce Type: new
Abstract: Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily...
By Ce Zhang, Jing Bi, Jinxi He, Jianshu Zhang, Jingyang Lin, Yunzhong Xiao, Minghao Fu, Yaqi Xie, Zhentao Xie, Weicong Chen, Katia Sycara, Ming Zhou
arXiv:2608. 03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence.
By Ke Li, Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen
arXiv:2606. 07639v1 Announce Type: cross Abstract: Video understanding is shifting from the offline paradigm -- taking a fully recorded video as input and producing a single answer after it ends -- toward real-time interaction, in which the model perceives new frames while still replying, revises its answer as new evidence appears, and remains silent when there is nothing to say.
By Pengyu Wang, Chenkun Tan, Shaojun Zhou, Wei Huang, Qirui Zhou, Zhan Huang, Zhen Ye, Jijun Cheng, Xiaomeng Qian, Yanxin Chen, Xingyang He, Huazheng Zeng, Chenghao Wang, Pengfei Wang, Hongkai Wang, Shanqing Gao, Yixian Tian, Chenghao Liu, Xinghao Wang, Botian Jiang, Xipeng Qiu