arXiv:2606. 12502v1 Announce Type: cross Abstract: We propose that value -- the quantity goal-directed agents create, destroy, and exchange -- is a lawful structural quantity in the same category as information.
By Cheng Qian
arXiv:2602. 09474v2 Announce Type: replace Abstract: We study reinforcement learning in MDPs whose transition function is stochastic at most steps but may behave adversarially at a fixed subset of $\Lambda$ steps per episode.
By Ofir Schlisselberg, Tal Lancewicki, Yishay Mansour
arXiv:2608. 12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards.
By Larissa Xu, King Bi, William Chang
SCAMP introduces a training‑free, damped Gauss‑Newton method that adjusts only the sparse anchor points in a frozen differentiable decoder, keeping the rest of the state unchanged. By operating solely in the space of the anchors’ Jacobian rows, it solves a system whose size matches the number of anchor constraints rather than the full state, enabling efficient control across diverse text‑to‑motion generators. Applied to seven existing generators, SCAMP achieves anchor errors that match or surpass all released control methods and can close anchors on hosts that originally lacked them.
By Pengcheng Fang, Tengjiao Sun, Xiaoyu Zhan, Yanwen Guo, Hansung Kim, Xiaohao Cai, Dongjie Fu
arXiv:2608. 07301v1 Announce Type: new Abstract: We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone.
By Neal Batra
arXiv:2607. 16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself.
By Venkatesh Saligrama
arXiv:2605. 25889v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models reach high success rates on clean inputs but collapse under small adversarial perturbations: a $16/255$ PGD attack drops OpenVLA-7B's LIBERO success from $95\%$ to under $5\%$.
By Jianwei Tai
arXiv:2607. 25408v1 Announce Type: new Abstract: A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al.
By Debjyoti Paul
arXiv:2606. 16515v1 Announce Type: cross Abstract: Hamilton-Jacobi-Bellman theory implies that the optimal goal-conditioned action depends on the goal only through the gradient of the goal-reaching distance at the current state, yet standard online GCRL still conditions the actor on the raw goal -- a signal that is geometrically uninformative when the goal is far from the data distribution.
By Swaminathan S K, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra
arXiv:2609. 21523v2 Announce Type: replace Abstract: A state may need to be compressed before its exact downstream task is known.
By Ronald Katende
arXiv:2608. 13651v1 Announce Type: cross Abstract: We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-\Delta, \Delta]$ is unknown in sign and magnitude and $\Delta$ is arbitrarily large.
By Dimitar Ho
arXiv:2607. 02196v1 Announce Type: new Abstract: We study online resource allocation when both rewards and consumption sizes may be continuously distributed.
By Jiawei Zhang