arXiv AI

Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment

arXiv:2606. 04750v1 Announce Type: new Abstract: Instilling virtuous behavior in artificial intelligence has seen increasing interest.

arXiv AI
6d ago

Preference-based opponent shaping in differentiable games

The paper introduces Preference-based Opponent Shaping (PBOS), a method that incorporates a preference parameter into an agent’s loss function to directly consider an opponent’s loss during strategy updates. By jointly learning strategy and preference parameters, PBOS aims to guide agents toward cooperative or competitive behaviors without relying on simple opponent predictions. Experiments on differentiable games demonstrate that PBOS enables agents to achieve better reward distributions across various environments.

By Xinyu Qiao, Yudong Hu, Congying Han, Weiyan Wu, Tiande Guo
arXiv AI
Sep 18

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.

By Wenjie Liao, Liangjie Zhao, Zehong Cao
arXiv Computation and Language
Sep 21

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

ArenaFlow is a hierarchical credit propagation framework designed to improve reinforcement learning for open-ended agent tasks. It uses tournament-based relative ranking to generate trajectory-level rewards and structured reflective evaluation to identify pivotal success steps, reusable strategy skills, and skill usage attribution. The framework propagates advantages to high-confidence steps and maintains a global skill memory, enabling more targeted optimization and reusable skill priors for future exploration.

By Qiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang, Yinfeng Huang, Yi Zheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha
arXiv Machine Learning
Sep 11

UBCL: A Reinforcement Learning Framework for Controllable and Diverse Player Behaviors

The paper presents UBCL, a reinforcement learning framework that generates controllable and diverse player behaviors without using human gameplay data. By defining behavior in an N‑dimensional continuous space and training a single PPO‑based multi‑agent policy with target behavior vectors, the method learns how actions affect behavioral statistics such as aggressiveness, mobility, and cooperativeness. Experiments in a custom Unity multiplayer game demonstrate that UBCL achieves greater behavioral diversity than a win‑only baseline and accurately matches specified behavior vectors across a range of targets.

By Atahan Cilan, Atay \"Ozg\"ovde