arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.
By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
arXiv:2609.38093v1 Announce Type: new
Abstract: Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor saf...
By Arav Dhoot, Punya Syon Pandey, Jamie Johnson, Daniel Tan, Elliott Thornley, David Demitri Africa
arXiv:2606. 19883v1 Announce Type: new Abstract: We study a multi-agent multi-armed bandit problem in the competitive setup with two-sided matching markets under a human centric decision making model.
By Ananya Kunisetty, Avishek Ghosh
arXiv:2607. 10251v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities.
By Xuankun Rong, Wenke Huang, Bo Du, Dacheng Tao, Mang Ye
arXiv:2606. 03238v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) makes large-scale post-training possible by replacing an underspecified human objective with learned and scalable proxies.
By Zelalem Abahana
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan