arXiv Machine Learning

Poincar\'e Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments

The paper presents a finite‑model synthesis that integrates operational state abstraction with optimal control within a stability‑evidence‑revision (SER) framework, termed "Poincaré meets Bellman." It distinguishes qualitative dynamics for reusable action‑response structures from dynamic programming that governs acquisition, retention, reuse, merging, and forgetting, and introduces a Bellman recursion over the joint law of hidden state and deployed memory. The authors derive explicit retention rules, demonstrate how factor sharing and informative observations improve identification, and verify coding and retention calculations through finite enumerations.

arXiv AI
Aug 20

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

The paper introduces Harness Continual Learning (HCL), a paradigm where an agent’s state evolves through prompts, memories, tools, skills, and routing rules while keeping the underlying foundation model frozen. HCL defines harness-level forgetting and proposes a guarded evolution process involving a Continual Optimizer and Evaluator to ensure improvements without losing prior behavior. Experiments across textual reasoning, multimodal perception, and open‑world interaction show over 10% performance gains and demonstrate how the stability–plasticity trade‑off can be explicitly tuned.

By Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao
arXiv AI
Aug 26

Revelation Control

Revelation Control studies how to price interventions that reveal hidden state only when the revealed distinctions can alter a consequential decision, while separately accounting for any useful progress the intervention itself creates. The authors develop a framework for learning systems that defines decision‑sufficient revelation, revelation depth, and a cost‑adjusted factorization criterion, and they provide a target‑independent protocol for model‑specific instantiation. Experiments on Qwen2.5‑7B and Mistral‑7B‑v0.3 show that deeper future‑learning probes have positive decision value and that productive reuse yields strict equal‑compute utility advantages, supporting a structural transfer of the decision theory and evaluation protocol across architectures.

By Qinyou Wang
arXiv AI
Jun 10

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

arXiv:2606. 10616v1 Announce Type: new Abstract: Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts that exceed their finite context windows, making memory retention a fundamental resource-allocation problem.

By Qingcan Kang, Liu Mingyang, Shixiong Kai, Kaichao Liang, Tao Zhong, Mingxuan Yuan
arXiv AI
2d ago

Measuring the Stability Assumption Behind Action Chunking

The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.

By Aryan Goyal