arXiv:2608.23329v1 Announce Type: cross
Abstract: Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video a...
By Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu, Bohan Hou, Xuanyu Zheng, Changyi Liu, Tianke Zhang, Haonan Fan, Kaiyu Jiang, Yingxin Li, Jiankang Chen, Xu Wang, Bin Wen, Tingting Gao, Han Li, Jianhua Yin, Yinwei Wei, Xuemeng Song
arXiv:2609.09300v1 Announce Type: new
Abstract: Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are dif...
By Zhenxin Qin, Peng Shi, Cong Han, Yinlong Qian, Zequn Jie, Lin Ma
LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. By confining bidirectional spatio‑temporal modeling to a fixed‑size window and using bounded temporal and global appearance memories, it emits clean video chunks with low latency. A progressive distillation pipeline further refines the model, achieving superior quality with 26× lower latency and 11× higher throughput compared to prior methods.
arXiv:2606. 20559v1 Announce Type: cross Abstract: Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action.
By Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le, Srijan Das
arXiv:2604. 06010v2 Announce Type: replace Abstract: Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed.
By Yukun Wang, Ruihuang Li, Jiale Tao, Shiyuan Yang, Liyi Chen, Zhantao Yang, Handz, Yulan Guo, Shuai Shao, Qinglin Lu
arXiv:2609.13024v1 Announce Type: new
Abstract: As a key model compression technique, knowledge distillation aims to transfer knowledge from a high-capacity teacher model to a lightweight student mod...
By Yanjiang Shi, Peng Zhao, Nan Qi, Guiqin Wang
arXiv:2608. 05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.
By Yushe Cao, Shikun Feng, Fei Shen, Haikuo Peng, Jianqiang Xia, Yiheng Zhu, Dianxi Shi, Chun Yu
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Video...
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both...
The paper introduces Split-then-Merge (StM), a new framework for generative video composition that improves control and tackles data scarcity. StM divides a large set of unlabeled videos into dynamic foreground and background layers, then self‑composes them to learn how subjects interact with varied scenes. The method employs a transformation‑aware training pipeline with multi‑layer fusion, augmentation, and an identity‑preservation loss, achieving superior performance over state‑of‑the‑art methods in both quantitative and qualitative evaluations.
By Ozgur Kara, Yujia Chen, Ming-Hsuan Yang, James M. Rehg, Wen-Sheng Chu, Du Tran
OmniVBench introduces a comprehensive benchmark and a large-scale dataset for omni reference-to-video (R2V) generation, addressing gaps in existing evaluations that focus only on limited reference types and holistic consistency. The benchmark expands evaluation across 7 task families and 18 fine-grained tasks, covering content, motion, style, structure, narrative, and multi-reference settings, and employs a factor‑grounded evaluation with 12,172 checklist items to assess preservation, disentanglement, and routing of reference factors. The accompanying Omni‑R2V Dataset provides 340K training samples derived from professional video footage, along with task‑specific pipelines for scalable data construction, enabling broader research and revealing performance gaps in current R2V models.
By Wenxue Li, Peiyan Guan, Haoyang Jiang, Junxian Cai, Hualuo Liu, Chunjie Zhang, Chong Guan, Songlian Li, Taiyi Wu, Yongjian Yu, Xiaotong Zhao, Alan Zhao, Eric Liu, Xi Chen, Yu Liu, Lei Zhu
LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.
By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing