arXiv Computer Vision
Sep 24

The Past Frames the Future: Memory for Autoregressive Video Generation

arXiv:2609.28466v1 Announce Type: new Abstract: Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environme...

By Harold Haodong Chen, Rongjin Guo, Disen Lan, Wen-Jie Shu, Hongfei Zhang, Hanzhe Hu, Shengtao Yao, Zixin Zhang, Guibin Zhang, Zhefan Rao, Jinxiu Liu, Yexin Liu, Rui Peng, Yuhao Liu, Bin Ren, Shuai Yang, Yukang Chen, Salman Khan, Ying-Cong Chen, Ser-Nam Lim, Rynson W. H. Lau, Nicu Sebe, Yu Cheng, Ming-Hsuan Yang, Qifeng Chen
arXiv AI
Aug 19

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

SemComp-Bench introduces a new video generation task called Semantic Task Completion, where a model must produce a video that achieves a specified outcome while maintaining semantic alignment with a reference image. The benchmark includes the SemComp-Data dataset, spanning six domains, and a four-stage curation pipeline that transforms raw videos into standardized instances. Evaluation is performed via a vision‑language model that answers structured binary questions, yielding Outcome Achievement (OA) and Generation Reliability (GR) scores.

By Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
arXiv Computer Vision
2d ago

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer is a streaming video model that jointly learns to record evidence and respond to tasks through a shared proactive generation process. Its Proactive Hierarchical Caption Memory creates time‑grounded local‑detail captions and event summaries, while Proactive State Transition Learning reduces waiting states by supervising all output anchors. The authors also built a large OneStreamer‑1M dataset and show that a 4B model outperforms baselines on eight streaming video benchmarks, with ablations confirming the benefits of generated captions and PSTL.

By Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si, Dingyu Yao, Changlian Ma, Haoran Chen, Xinyu Chen, Yansong Shi, Junhao Zhou, Yifei Li, Jun Zhang, Chuanyu Qin, Chenxu Yang, Xinlei Yu, Kun Ouyang, Yuchen Shao, Qianshan Wei, Changhai Zhou, Jun Gao, Jiaqi Wang, Limin Wang