arXiv AI

Hierarchical Latent Prediction for Language Models

arXiv:2608. 05806v1 Announce Type: cross Abstract: While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optimal for long-horizon reasoning and planning.

arXiv AI
Jul 29

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

arXiv:2607. 25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens.

By Yutong Chen, Shouqian Shi, Xinran Liu, Haochen Wang, Jiaying Wang, Tianxing Xu, Yuanxi Wang, Zirui Ding
arXiv Machine Learning
Jun 5

Latent Reasoning with Normalizing Flows

arXiv:2606. 06447v1 Announce Type: cross Abstract: Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation.

By Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang, Haoqiang Kang, Lianhui Qin, Yizhe Zhang, Jiatao Gu
arXiv Machine Learning
Sep 1

A Model with No Head and Many Thoughts

arXiv:2608.31069v1 Announce Type: new Abstract: Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and fo...

By Nikita Koriagin, Yaroslav Aksenov, George Bredis, Gleb Gerasimov, Nikita Balagansky, Daniil Gavrilov
Hugging Face Trending Papers
Jun 4

Latent Reasoning with Normalizing Flows

Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning step must be verbalized before the model can proceed, even when the underlying update is semantic, uncertain, or only partially formed.

arXiv Computation and Language
Sep 11

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

NCP-ArchPreview is a latent‑space language model that extends standard next‑token prediction (NTP) with a Next Concept Prediction (NCP) objective, allowing the model to predict discrete concepts spanning multiple tokens. The architecture builds a product‑quantized concept vocabulary from hidden states, uses a dedicated Concept Module to forecast future concepts, and feeds these predictions back to guide token‑level generation, all trained jointly end‑to‑end. Trained on 5.73 T tokens with 8.9 B parameters, it achieves the final pretraining loss of OLMo‑3‑7B using only 51.3 % of the tokens, outperforms OLMo‑3‑7B on downstream tasks (including a 5.99‑point GSM8K gain), and demonstrates that the learned latent space enables lightweight domain adaptation and improved drafting performance.

By NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng, Xu Cheng, Beiya Dai, Yufan Feng, Kewen Ge, Ruijun Ge, Jiayi Huang, Yang Jiao, Dahua Lin, Zhouhan Lin, Yifan Liu, Yuliang Liu, Biqing Qi, Mowen Ruan, Junzhe Shen, Yunchong Song, Hao Sun, Zhongbo Tian, Yixuan Wang, Rubin Wei, Jiaxin Xiong, Kangyu Yang, Qian Yao, Qi Zhang, Bowen Zhou