arXiv Machine Learning

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference

arXiv:2607. 08724v1 Announce Type: new Abstract: Human decision-making is highly flexible -- some actions are taken immediately; others require longer deliberation.

arXiv AI
Sep 1

AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

AgenticRag‑R1 is a reinforcement‑learning framework that integrates reasoning, retrieval, and memory through a stack and fine‑grained action space. It uses hierarchical action‑aware rewards and an information‑aware trajectory rejection strategy to support long‑horizon learning. Experiments on multi‑hop, open‑domain, and agentic reasoning benchmarks show that AgenticRag‑R1 outperforms strong baselines and produces robust, interpretable, memory‑aware reasoning behaviors.

By Xinke Jiang, Yue Fang, Zhibang Yang, Jiaran Gao, Zhixin Zhang, Tao Feng, Rihong Qiu, Wentao Zhang, Hongxin Ding, Ruizhe Zhang, Yongxin Xu, Yuheng Huang, Xu Chu, Junfeng Zhao, Yasha Wang
arXiv Computation and Language
3d ago

LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation

LatentHarness unifies memory access and latent reasoning by treating them as sequential latent actions—THINK, RECALL, and EXIT—within a language model. It is trained via counterfactual policy distillation, which evaluates the impact of each action on the emitted token and learns when to recall evidence versus continue reasoning. On six long‑context reasoning benchmarks, a 1.4B‑parameter LatentHarness model outperforms the strongest baselines by 2.8% and 10.0% relative, while running 5.9× faster than the leading long‑context baseline.

By Xiaoqiang Wang, Suyuchen Wang, Bang Liu
arXiv Computer Vision
Sep 17

Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation

The paper introduces Latent Motion Reasoning (LMR), a two‑stage approach that separates text‑to‑motion generation into a planning phase and an execution phase. LMR uses a Dual‑Granularity Tokenizer to create a compressed, semantically rich reasoning latent for global trajectory planning and a high‑frequency execution latent for detailed motion fidelity. Experiments on T2M‑GPT and MotionStreamer show that this architecture improves both semantic alignment and physical plausibility compared to direct translation methods.

By Yijie Qian, Juncheng Wang, Yuxiang Feng, Chao Xu, Wang Lu, Yang Liu, Baigui Sun, Yiqiang Chen, Yong Liu, Shujun Wang
arXiv AI
Jul 29

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

arXiv:2607. 25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens.

By Yutong Chen, Shouqian Shi, Xinran Liu, Haochen Wang, Jiaying Wang, Tianxing Xu, Yuanxi Wang, Zirui Ding
arXiv Machine Learning
Aug 21

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

arXiv:2608. 19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage.

By Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi