arXiv AI

CoupVisor: Strategy Optimization by Round and Challenge Decision Support

arXiv:2608. 15868v1 Announce Type: new Abstract: This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup.

arXiv Machine Learning
Aug 13

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
arXiv AI
Sep 18

Mitigating Retaliatory Algorithmic Collusion in Repeated Games

The paper introduces CURB, a reward‑shaping framework that penalizes the total variation distance between an agent’s action distributions under cooperation and defection histories, thereby preventing collusive equilibria in repeated games. By linking empirical Q‑learning collusion to Simple Penal Codes, the authors prove that any non‑trivial SPC can be neutralized, and demonstrate CURB’s effectiveness in both tabular and deep Q‑learning settings for Bertrand and Cournot competition.

By Karthik Sivachandran, Rohan Paleja
arXiv AI
6d ago

Preference-based opponent shaping in differentiable games

The paper introduces Preference-based Opponent Shaping (PBOS), a method that incorporates a preference parameter into an agent’s loss function to directly consider an opponent’s loss during strategy updates. By jointly learning strategy and preference parameters, PBOS aims to guide agents toward cooperative or competitive behaviors without relying on simple opponent predictions. Experiments on differentiable games demonstrate that PBOS enables agents to achieve better reward distributions across various environments.

By Xinyu Qiao, Yudong Hu, Congying Han, Weiyan Wu, Tiande Guo
arXiv Machine Learning
Aug 19

Debate Training Reduces Reward Hacking in RLAIF

The paper shows that fine‑tuning a large language model (LLM) with a debate framework—where a generator and a critic compete and a weaker LLM judge adjudicates—reduces reward hacking compared to standard reinforcement learning from AI feedback (RLAIF). In experiments on mathematics tasks, the debate approach keeps the judge’s performance stable, achieving a 45% higher peak validation accuracy than the RLAIF baseline and mitigating the rapid exploitation of judge errors. Additional findings indicate that weakening the judge speeds hacking unless countered by extra debate rounds, that debate can override misalignment prompts, and that word‑limit constraints on critiques help balance the game and prevent judge hacking. whyItMatters":"The study demonstrates a practical method to curb reward hacking in RL‑based AI systems, addressing a key obstacle for safely scaling AI oversight."

By Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah
Hugging Face Trending Papers
Aug 12

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.

arXiv AI
Jun 2

MindGames Arena Generalization Track: In2AI Solution with Delayed Per-Step Reward Attribution

arXiv:2606. 00017v1 Announce Type: new Abstract: Training language model agents for multi-agent strategic interaction presents a core difficulty: the quality of any action may depend on future events that never materialize, on moves that violate game rules, or on decisions made by other players.

By Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov