arXiv AI

The Free Inference Dimension: Complexity Measure for Zero-Collision Navigation under Hypothesis Mixtures

The paper introduces the Free Inference dimension (dFI) as a new combinatorial measure of environmental complexity for value‑mixture agents in finite meta‑reinforcement learning. It shows that dFI is smaller than the VC‑dimension and relates to the Natarajan dimension, providing PAC‑style generalization bounds. The authors also define a complementary PMS identification dimension and demonstrate that a hybrid strategy—averaging until the first collision and then selecting—achieves optimal performance, supported by grid‑world experiments.

arXiv Machine Learning
Jun 16

Learning Policy from a Single Trajectory in Average-Reward Markov Decision Process

arXiv:2606. 16729v1 Announce Type: new Abstract: While there is an extensive body of work characterizing the sample complexity of discounted cumulative-reward MDPs, finite sample analyses for average-reward MDPs have been limited, and most existing works rely on restrictive assumptions such as ergodicity or access to a generative model.

By Jongmin Lee, Ernest K. Ryu, Vaneet Aggarwal
arXiv Machine Learning
Sep 4

Robust PAC Learning of Concurrent Stochastic Games

The paper presents the first PAC learning framework for general-sum concurrent stochastic games with uncertain transitions, addressing the challenge of Nash equilibrium existence. It introduces data‑driven L¹ confidence sets over transition kernels and a robust CSG solver that computes a social‑welfare optimal ε‑NE, or provides a certificate that no exact NE exists. The algorithm achieves polynomial sample complexity under a minimum reachability condition and is validated on benchmark CSGs with near‑optimal performance.

By Angel Y. He, David Parker
Hugging Face Trending Papers
Sep 3

Robust PAC Learning of Concurrent Stochastic Games

We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty, while addressing the challenge of Nash equilibrium (NE) existence. Our algorithm maintains data-driven $L^1$ confidence sets over transition kernels and solves a robust CSG to compute a social-welfare optimal $\varepsilon$-NE, using a robust MDP-based exploration mechanism to drive joint state-action coverage.

arXiv Machine Learning
Sep 24

Limiting-Kernel Q($\lambda$): Bridging Short and Long Horizons

Limiting‑Kernel Q(λ) (LKQL) is an off‑policy value estimator that blends n‑step truncation with a long‑horizon approximation based on the limiting kernel. It maintains the computational efficiency of n‑step methods while improving policy evaluation accuracy, especially for long‑horizon tasks. The authors prove faster convergence of LKQL’s operator under aperiodicity and near‑on‑policy conditions, and demonstrate empirical gains on MuJoCo continuous‑control benchmarks.

By Tolga Ok, Arman Sharifi Kolarijani, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi Kolarijani
arXiv Machine Learning
Sep 25

Active Client Selection in Federated Trajectory Prediction with Uncertainty-Awareness and Heterogeneous Complexity

The paper introduces active client selection strategies for federated learning in autonomous vehicle trajectory prediction, addressing challenges of high scene uncertainty and heterogeneous complexity across different driving environments. It proposes uncertainty-aware selectors that use per-client negative log-likelihood and aleatoric uncertainty, as well as a joint selector that balances scene complexity and uncertainty to prioritize informative clients. Experiments on the Argoverse dataset show that federated models outperform local training, with uncertainty-aware selection speeding convergence and improving key metrics, while the joint selector yields the best generalization under strong heterogeneity.

By Yiming Xie, Muzi Peng, Fei Miao, Ningfang Mi, Lili Su