arXiv:2606. 07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces.
By Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung, Jaejin Lee, Taesup Kim
arXiv:2607. 27036v1 Announce Type: cross Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time.
By Taiye Chen, Qi Zhang, Yisen Wang
CoRe introduces a co‑evolving reward framework to mitigate latent reward hacking in video diffusion models. By continuously refitting the latent‑reward model on the generator’s current samples and anchoring it to real‑video preferences, CoRe prevents the generator from drifting outside the reward model’s training support. Experiments on Wan2.1‑T2V‑1.3B demonstrate that CoRe improves generation quality over pretrained models and prior alignment methods while avoiding quality collapse.
By Zhaolong Su, Yujin Han, Feng Wang, Jameson Dong, Hins Hu, Difan Zou
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2604.16067v2 Announce Type: replace-cross
Abstract: Fine-tuning pre-trained Vision-Language Models (VLMs) for robotic manipulation introduces a fundamental stability-plasticity dilemma: continu...
By Guransh Singh
arXiv:2506. 01274v2 Announce Type: replace-cross Abstract: Recent progress in Large Multi-modal Models (LMMs) has enabled effective vision-language reasoning, yet the ability to video understanding remains constrained by suboptimal frame selection strategies, albeit with the rapid development of video-specialized LMMs.
By Hosu Lee, Junho Kim, Hyunjun Kim, Yong Man Ro
arXiv:2609.24788v1 Announce Type: new
Abstract: In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-qualit...
By Yujia Hu, Jiajun Li, Zihao He, Songhua Liu
The paper introduces Latent-Centroid Steering (LCS), a single-pass classifier-free guidance method for vision‑language autonomous driving models. LCS replaces instance‑level residuals with class‑level latent shifts, projecting conditional representations toward precomputed command‑specific centroids to enhance command adherence. Experiments on Bench2Drive and nuScenes show that LCS cuts inference latency by about 50% while improving driving performance.
By Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao
arXiv:2608.30194v1 Announce Type: new
Abstract: Diffusion models have recently advanced text-to-video (T2V) generation, yet they still struggle with fine-grained compositional alignment, such as attr...
By Yujiang Pu, Yu Kong
arXiv:2609.37250v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
By Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
arXiv:2608. 07585v1 Announce Type: cross Abstract: Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams.
By Zijian Wang, Junnan Zhu, Rongzhen Li, Xiao Liu, Guohui Xiang, Quan Lu, Lijia Liu, Yining Wang, Jiang Zhong, Kaiwen Wei
arXiv:2605.08712v2 Announce Type: replace
Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dime...
By Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang, Youqi Tao, Junfeng Duan, Daguang Xu, Qi Dou, Xin Jin, Wenjun Zeng, Hao Zhao, Yueming Jin