Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.
arXiv:2605. 23045v2 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years.
By Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
arXiv:2603. 12478v2 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven.
By Rujie Wu, Haozhe Zhao, Hai Ci, Yizhou Wang
arXiv:2601. 22108v2 Announce Type: replace-cross Abstract: Continued pretraining is optimized with fixed self-supervised tasks but selected by downstream performance, creating a coarse feedback loop in which practitioners evaluate checkpoints, change data mixtures or objectives, and restart runs, while individual updates remain blind to target capabilities.
By Shuqi Ke, Giulia Fanti
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
arXiv:2510. 09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage.
By Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, Yao Lu, Song Han
arXiv:2606. 15956v1 Announce Type: cross Abstract: Progress in AI has largely been driven by methods that assume less.
By Ninad Daithankar, Alexi Gladstone, Yann LeCun, Heng Ji
arXiv:2607. 09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models.
By Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
Instruction-Based Video Editing by Repurposing an Image Editing Model demonstrates that a strong image‑editing model can be adapted to edit videos by operating on video‑VAE latents. The authors tile latent frames into a large virtual image, reuse the editor’s positional encoding, and bridge latent spaces with lightweight projections, fine‑tuning on Ditto‑1M editing triplets. Their experiments show that per‑frame video latents are close enough to the image domain that mature image‑editing priors transfer with minimal adaptation.
By Yunpeng Bai, Yossi Gandelsman, Micha\"el Gharbi, Qixing Huang
arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.
By Jakob Geusen, Ender Konukoglu
arXiv:2608.25729v1 Announce Type: new
Abstract: Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a...
By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny
arXiv:2605. 18324v2 Announce Type: replace-cross Abstract: Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders.
By Jaskirat Singh, Boyang Zheng, Zongze Wu, Richard Zhang, Eli Shechtman, Saining Xie