arXiv AI By Ao Shen, Yongheng Zhang, Yinghui Li, Manning Wang, Di Yin, Xing Sun

Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

Read the original on arXiv AI →

arXiv:2608. 16316v1 Announce Type: cross Abstract: Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 28

Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models

Video-OPSD introduces a post‑training framework for Video Large Language Models that leverages privileged visual evidence to enhance on‑policy self‑distillation. The method constructs a self‑teacher conditioned only on annotated evidence frames, while the student processes the full video, allowing the teacher to provide more focused supervision. Additionally, an evidence‑guided token optimization scheme weights distillation based on each token’s reliance on privileged evidence, improving perceptually grounded reasoning. Experiments demonstrate consistent gains over standard OPSD and comparable performance to GRPO with less training time.

By Ziyue Wang, Shiqi Huang, Weiwen Xu, Bihan Wen, Xudong Jiang
arXiv Computer Vision
3d ago

VR-JEPA: Learning Contrastive-State Latent Guidance for Generation-based Video Reasoning

arXiv:2609.40129v1 Announce Type: new Abstract: Reasoning through video generation offers a promising path toward visual intelligence by modeling latent visual states and their dynamics. However, cur...

By Zehua Ma, Kun Xiang, Yunshuang Nie, Quanlin Chen, Haoyuan Li, Xiuwei Chen, Jiang Ji, Haijun Wu, Zhenyu Xie, Michael Kampffmeyer, Hanhui Li, Xiaodan Liang