arXiv Machine Learning

Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models

arXiv:2606. 15099v1 Announce Type: cross Abstract: Existing Vision-Language-Action (VLA) models predominantly rely on explicit Chain-of-Thought (CoT) reasoning to bridge perception and action.

arXiv Computer Vision
Aug 27

Latent Chain-of-Thought World Modeling for End-to-End Driving

Latent-CoT-Drive (LCDrive) is a vision‑language‑action model for autonomous driving that replaces natural‑language chain‑of‑thought reasoning with a latent language capturing possible outcomes of driving actions. The model interleaves action‑proposal tokens, aligned with the model’s output actions, and world‑model tokens grounded in a learned latent world model to reason about future outcomes. After a supervised cold‑start using ground‑truth future rollouts, LCDrive is further refined with closed‑loop reinforcement learning, achieving faster inference, higher‑quality trajectories, and greater gains from interactive RL than both non‑reasoning and text‑reasoning baselines on a large‑scale end‑to‑end driving benchmark.

By Shuhan Tan, Kashyap Chitta, Yuxiao Chen, Ran Tian, Yurong You, Yan Wang, Wenjie Luo, Yulong Cao, Philipp Krahenbuhl, Marco Pavone, Boris Ivanovic
arXiv Machine Learning
Aug 21

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

arXiv:2608. 19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage.

By Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi
Hugging Face Trending Papers
Aug 20

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one in each stage.

arXiv AI
Jul 28

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

arXiv:2512. 24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models.

By Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo