arXiv AI

Finite Reliability Representations: Noise-Calibrated Belief-Space Covers for Reliable Decision-Making

arXiv:2607. 04019v1 Announce Type: cross Abstract: Physical sensing and actuation noise floors should inform how much belief resolution a decision-making system can reliably use.

arXiv Machine Learning
Jul 21

Distributional Soft Bellman Operator under the Cram\'er Geometry

arXiv:2607. 17897v1 Announce Type: new Abstract: Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns.

By Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li
arXiv Machine Learning
Jul 22

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

arXiv:2607. 18554v1 Announce Type: cross Abstract: We develop the Continuous Distributed Coupled Policy Gradient (CDCPG) algorithm for cooperative reinforcement learning in networked Markov decision processes with continuous state and action spaces.

By Dongming Wang, Pengcheng Dai, Wenwu Yu, Wei Ren
arXiv Machine Learning
Jun 29

PAC-Bayesian Certificates for Quadratic Closed-Loop Control

arXiv:2606. 28281v1 Announce Type: cross Abstract: PAC-Bayesian bounds provide finite-sample guarantees for data-dependent randomized predictors, but applying them to learning-based control is difficult because the natural objective is a quadratic trajectory cost.

By Domagoj Herceg