On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces Fed‑LSVI, a federated online reinforcement learning algorithm that uses linear function approximation in episodic Markov decision processes. It achieves a regret bound of ≥O(√{Md^3H^4T}) while only exchanging compressed sufficient statistics, thereby meeting privacy constraints. The method reduces communication cost to logarithmic in the number of episodes, a marked improvement over previous approaches that required linear communication.
arXiv:2607. 16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself.
arXiv:2608. 01151v1 Announce Type: cross Abstract: In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints.
arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.
The paper formulates and analyzes the linear exponential quadratic Gaussian (LEQG) covariance steering problem in continuous time over a finite horizon. It shows that the optimal controller, still a linear state feedback, cannot be expressed in closed form but is parameterized by a symmetric matrix solving an algebraic equation that captures the risk‑sensitivity parameter. The authors demonstrate that this controller generalizes the risk‑neutral case and prove existence‑uniqueness of solutions near the known risk‑neutral solution for matched noise and input channels, illustrated with a numerical example.
arXiv:2602. 12963v2 Announce Type: replace Abstract: An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world.