arXiv Machine Learning By Xin Li

Poincar\'e Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments

Read the original on arXiv Machine Learning →

The paper presents a finite‑model synthesis that integrates operational state abstraction with optimal control within a stability‑evidence‑revision (SER) framework, termed "Poincaré meets Bellman." It distinguishes qualitative dynamics for reusable action‑response structures from dynamic programming that governs acquisition, retention, reuse, merging, and forgetting, and introduces a Bellman recursion over the joint law of hidden state and deployed memory. The authors derive explicit retention rules, demonstrate how factor sharing and informative observations improve identification, and verify coding and retention calculations through finite enumerations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 20

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

The paper introduces Harness Continual Learning (HCL), a paradigm where an agent’s state evolves through prompts, memories, tools, skills, and routing rules while keeping the underlying foundation model frozen. HCL defines harness-level forgetting and proposes a guarded evolution process involving a Continual Optimizer and Evaluator to ensure improvements without losing prior behavior. Experiments across textual reasoning, multimodal perception, and open‑world interaction show over 10% performance gains and demonstrate how the stability–plasticity trade‑off can be explicitly tuned.

By Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao
arXiv AI
Aug 26

Revelation Control

Revelation Control studies how to price interventions that reveal hidden state only when the revealed distinctions can alter a consequential decision, while separately accounting for any useful progress the intervention itself creates. The authors develop a framework for learning systems that defines decision‑sufficient revelation, revelation depth, and a cost‑adjusted factorization criterion, and they provide a target‑independent protocol for model‑specific instantiation. Experiments on Qwen2.5‑7B and Mistral‑7B‑v0.3 show that deeper future‑learning probes have positive decision value and that productive reuse yields strict equal‑compute utility advantages, supporting a structural transfer of the decision theory and evaluation protocol across architectures.

By Qinyou Wang