arXiv Machine Learning By Zuyuan Zhang, Yongshan Chen, Mahdi Imani, Tian Lan

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

Read the original on arXiv Machine Learning →

arXiv:2607. 27132v1 Announce Type: new Abstract: An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Poincar\'e Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments

The paper presents a finite‑model synthesis that integrates operational state abstraction with optimal control within a stability‑evidence‑revision (SER) framework, termed "Poincaré meets Bellman." It distinguishes qualitative dynamics for reusable action‑response structures from dynamic programming that governs acquisition, retention, reuse, merging, and forgetting, and introduces a Bellman recursion over the joint law of hidden state and deployed memory. The authors derive explicit retention rules, demonstrate how factor sharing and informative observations improve identification, and verify coding and retention calculations through finite enumerations.

By Xin Li
arXiv Machine Learning
Aug 31

Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning

The paper challenges the common assumption that the successor measure in reinforcement learning is approximately low-rank, showing instead that a low-rank structure emerges in a shifted successor measure that ignores initial transitions. It provides finite-sample guarantees for estimating this low-rank approximation, introduces Type II Poincaré inequalities to bound spectral recoverability, and links the necessary shift to the decay of high-order singular values and local mixing properties. Experiments confirm that shifting the successor measure improves goal-conditioned RL performance.

By Bastien Dubail, Stefan Stojanovic, Alexandre Prouti\`ere
arXiv AI
2d ago

Nous: Learning and Certifying Memory Decisions Before Source Calibration

The paper introduces Nous, a framework that separates learning, calibration, and revision certification for belief‑based agent memory. It shows that learning and certifying useful memory decisions can require quadratically fewer records than source calibration, and provides finite‑sample certificates for policy improvement without needing to recover source reliability. Experiments on MiniGrid environments demonstrate that the new certificate reliably accepts improvements over incumbents, outperforming earlier methods.

By Pranav Singh