arXiv:2603. 00374v2 Announce Type: replace Abstract: Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectories.
By Austin A. Nguyen, Michael P. Wellman
arXiv:2606. 11284v1 Announce Type: cross Abstract: Real-world multi-agent systems, from traffic coordination to resource allocation, are often modeled as general-sum games where individual incentives conflict with collective welfare.
By Wongyu Lee, Francesco Lelli, Omran Ayoub, Massimo Tornatore
arXiv:2608. 09389v1 Announce Type: cross Abstract: This note aims to serve as an entry point to the literature on learning in games, a topic with significant theoretical appeal and a wide range of applications -- from machine learning and data science to economics and beyond.
By Panayotis Mertikopoulos
arXiv:2605. 00155v3 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) is a central post-training tool for aligning large language models, but its training reward is only a learned proxy for true human utility.
By Yikai Wang, Shang Liu, Jose Blanchet
arXiv:2606. 06486v1 Announce Type: new Abstract: In this paper, we study regret minimization in repeated games with \emph{adaptive} opponents who can respond based on histories of play.
By Mingyang Liu, Asuman Ozdaglar, Tiancheng Yu, Kaiqing Zhang
The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework that adjusts an agent’s caution based on epistemic uncertainty by using a Bayesian posterior over dynamics and a Wasserstein ambiguity set whose radius depends on that posterior. As evidence accumulates, the radius shrinks, smoothly transitioning the agent’s behavior from worst-case robustness to risk-neutral reward maximization. The authors prove a Safety Sandwich theorem showing RATTL’s value lies between the uninformed robust value and the full-knowledge optimum, and demonstrate the method on a binary-hazard example where the criterion reduces to Conditional Value-at-Risk.
By Deep Kumar Ganguly, Jan Kretinsky
arXiv:2607. 10630v1 Announce Type: cross Abstract: Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data.
By Tong Nie, Yuewen Mei, Junlin He, Yihong Tang, Jian Sun, Wei Ma
arXiv:2510. 07424v3 Announce Type: replace Abstract: We study linear contextual bandits with paid observations, where at each round the learner observes a context, selects an action, and may pay a fixed cost to observe feedback from a subset of arms.
By Nathan Boyer, Dorian Baudry, Patrick Rebeschini
arXiv:2609.38938v1 Announce Type: new
Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies...
By Xinyi Ni, Lifeng Lai
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
By Christoph Dann, Yishay Mansour, Mehryar Mohri
Robust Markov Decision Processes (RMDPs) generalize classical MDPs by allowing uncertainty in transition probabilities and optimizing against their worst-case realization. We consider $(s,a)$-rectangu...
The paper introduces a new approach to safety in contextual bandits with continuous actions, focusing on high‑probability constraints on the realized cost rather than expected cost. It presents the High‑Probability Constrained UCB algorithm, which balances reward exploration with conservative safety estimation, and provides theoretical regret guarantees for linear models and extensions to general function classes. Experiments demonstrate that this realized‑cost safety framework significantly reduces safety violations compared to expected‑cost constrained methods.
By Spyros Dragazis, Aldo Pacchiano