arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
The paper introduces CURB, a reward‑shaping framework that penalizes the total variation distance between an agent’s action distributions under cooperation and defection histories, thereby preventing collusive equilibria in repeated games. By linking empirical Q‑learning collusion to Simple Penal Codes, the authors prove that any non‑trivial SPC can be neutralized, and demonstrate CURB’s effectiveness in both tabular and deep Q‑learning settings for Bertrand and Cournot competition.
By Karthik Sivachandran, Rohan Paleja
arXiv:2607. 00155v1 Announce Type: new Abstract: We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes.
By Yunjin Tong
arXiv:2609.06816v1 Announce Type: new
Abstract: Decision-time search in perfect and imperfect information games with enumerable belief states are effective methods for game AI. Collectible card games...
By Dustin Rubin
The paper introduces Preference-based Opponent Shaping (PBOS), a method that incorporates a preference parameter into an agent’s loss function to directly consider an opponent’s loss during strategy updates. By jointly learning strategy and preference parameters, PBOS aims to guide agents toward cooperative or competitive behaviors without relying on simple opponent predictions. Experiments on differentiable games demonstrate that PBOS enables agents to achieve better reward distributions across various environments.
By Xinyu Qiao, Yudong Hu, Congying Han, Weiyan Wu, Tiande Guo
The paper shows that fine‑tuning a large language model (LLM) with a debate framework—where a generator and a critic compete and a weaker LLM judge adjudicates—reduces reward hacking compared to standard reinforcement learning from AI feedback (RLAIF). In experiments on mathematics tasks, the debate approach keeps the judge’s performance stable, achieving a 45% higher peak validation accuracy than the RLAIF baseline and mitigating the rapid exploitation of judge errors. Additional findings indicate that weakening the judge speeds hacking unless countered by extra debate rounds, that debate can override misalignment prompts, and that word‑limit constraints on critiques help balance the game and prevent judge hacking.
whyItMatters":"The study demonstrates a practical method to curb reward hacking in RL‑based AI systems, addressing a key obstacle for safely scaling AI oversight."
By Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah
arXiv:2605. 28863v2 Announce Type: replace-cross Abstract: Imperfect-information multiplayer games test whether agents can act under hidden information, sparse rewards, and non-stationary opponents.
By Aalok Patwa
arXiv:2609.38881v1 Announce Type: new
Abstract: Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and att...
By Xinhe Tian, Xiaoyue Zhang, Ziyou Zhang, Jiacheng Li, Xiaoqiang Jin, Qianchuan Zhao, Gaochen Cui
Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.
arXiv:2606. 00017v1 Announce Type: new Abstract: Training language model agents for multi-agent strategic interaction presents a core difficulty: the quality of any action may depend on future events that never materialize, on moves that violate game rules, or on decisions made by other players.
By Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov
arXiv:1908.08773v3 Announce Type: replace
Abstract: In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit...
By Victor Gallego, Roi Naveiro, David Rios Insua, David Gomez-Ullate Oteiza
arXiv:2608. 02440v1 Announce Type: cross Abstract: In noisy social dilemmas, intended actions are stochastically corrupted before execution, so an observed defection may reflect hostile intent or action error.
By Kival Mahadew, Jonathan Shock