arXiv:2606. 12502v1 Announce Type: cross Abstract: We propose that value -- the quantity goal-directed agents create, destroy, and exchange -- is a lawful structural quantity in the same category as information.
By Cheng Qian
arXiv:2602. 09474v2 Announce Type: replace Abstract: We study reinforcement learning in MDPs whose transition function is stochastic at most steps but may behave adversarially at a fixed subset of $\Lambda$ steps per episode.
By Ofir Schlisselberg, Tal Lancewicki, Yishay Mansour
arXiv:2608. 12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards.
By Larissa Xu, King Bi, William Chang
SCAMP introduces a training‑free, damped Gauss‑Newton method that adjusts only the sparse anchor points in a frozen differentiable decoder, keeping the rest of the state unchanged. By operating solely in the space of the anchors’ Jacobian rows, it solves a system whose size matches the number of anchor constraints rather than the full state, enabling efficient control across diverse text‑to‑motion generators. Applied to seven existing generators, SCAMP achieves anchor errors that match or surpass all released control methods and can close anchors on hosts that originally lacked them.
By Pengcheng Fang, Tengjiao Sun, Xiaoyu Zhan, Yanwen Guo, Hansung Kim, Xiaohao Cai, Dongjie Fu
arXiv:2608. 07301v1 Announce Type: new Abstract: We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone.
By Neal Batra
arXiv:2607. 16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself.
By Venkatesh Saligrama