arXiv AI By Yuyan Zhou, Jiarui Yu, Hande Dong, Zhezheng Hao, Hong Wang, Jianqing Zhang, Qiang Lin

LEPO: Latent Reasoning Policy Optimization for Large Language Models

Read the original on arXiv AI →

arXiv:2604. 17892v4 Announce Type: replace-cross Abstract: Recently, latent reasoning has been introduced into large language models (LLMs) to leverage rich information within a continuous space.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.