arXiv Machine Learning

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

arXiv:2607. 27132v1 Announce Type: new Abstract: An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown.

arXiv Machine Learning
1d ago

Poincar\'e Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments

The paper presents a finite‑model synthesis that integrates operational state abstraction with optimal control within a stability‑evidence‑revision (SER) framework, termed "Poincaré meets Bellman." It distinguishes qualitative dynamics for reusable action‑response structures from dynamic programming that governs acquisition, retention, reuse, merging, and forgetting, and introduces a Bellman recursion over the joint law of hidden state and deployed memory. The authors derive explicit retention rules, demonstrate how factor sharing and informative observations improve identification, and verify coding and retention calculations through finite enumerations.

By Xin Li
arXiv Machine Learning
Aug 31

Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning

The paper challenges the common assumption that the successor measure in reinforcement learning is approximately low-rank, showing instead that a low-rank structure emerges in a shifted successor measure that ignores initial transitions. It provides finite-sample guarantees for estimating this low-rank approximation, introduces Type II Poincaré inequalities to bound spectral recoverability, and links the necessary shift to the decay of high-order singular values and local mixing properties. Experiments confirm that shifting the successor measure improves goal-conditioned RL performance.

By Bastien Dubail, Stefan Stojanovic, Alexandre Prouti\`ere
arXiv AI
2d ago

Nous: Learning and Certifying Memory Decisions Before Source Calibration

The paper introduces Nous, a framework that separates learning, calibration, and revision certification for belief‑based agent memory. It shows that learning and certifying useful memory decisions can require quadratically fewer records than source calibration, and provides finite‑sample certificates for policy improvement without needing to recover source reliability. Experiments on MiniGrid environments demonstrate that the new certificate reliably accepts improvements over incumbents, outperforming earlier methods.

By Pranav Singh
arXiv AI
2d ago

Measuring the Stability Assumption Behind Action Chunking

The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.

By Aryan Goyal
arXiv AI
Jul 20

From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

arXiv:2607. 15459v1 Announce Type: new Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit.

By Eduardo C. Garrido-Merch\'an