Policy Improvement with Style-Specific Demonstrations
arXiv:2506. 16995v4 Announce Type: replace Abstract: Proficient game agents with diverse play styles enrich the gaming experience and enhance the replay value of games.
The paper presents UBCL, a reinforcement learning framework that generates controllable and diverse player behaviors without using human gameplay data. By defining behavior in an N‑dimensional continuous space and training a single PPO‑based multi‑agent policy with target behavior vectors, the method learns how actions affect behavioral statistics such as aggressiveness, mobility, and cooperativeness. Experiments in a custom Unity multiplayer game demonstrate that UBCL achieves greater behavioral diversity than a win‑only baseline and accurately matches specified behavior vectors across a range of targets.
arXiv:2506. 16995v4 Announce Type: replace Abstract: Proficient game agents with diverse play styles enrich the gaming experience and enhance the replay value of games.
arXiv:2607. 27574v1 Announce Type: new Abstract: Activation steering has emerged in large language models as a lightweight alternative for dynamically changing a model's behavior at inference time.
arXiv:2512. 09706v2 Announce Type: replace Abstract: The paradigm of agentic AI is shifting from engineered complex workflows to post-training native models.
arXiv:2506. 01568v4 Announce Type: replace Abstract: Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima.
The paper introduces Trajectory-guided Joint Policy Optimization (TJPO), a reinforcement‑learning framework that explicitly encourages diversity in the trajectories of large‑language‑model agents. By defining task‑specific trajectory descriptors, TJPO measures and optimizes diversity as a set‑level function, avoiding population‑based training. Experiments on Sokoban and ALFWorld demonstrate that TJPO yields diverse, interpretable behaviors while preserving strong task performance.
arXiv:2607. 00642v1 Announce Type: new Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models.
arXiv:2601. 05675v2 Announce Type: replace Abstract: Hybrid action space, which combines discrete choices and continuous parameters, is prevalent in domains such as robot control and game AI.
arXiv:2509. 23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences.
The paper introduces Preference-based Opponent Shaping (PBOS), a method that incorporates a preference parameter into an agent’s loss function to directly consider an opponent’s loss during strategy updates. By jointly learning strategy and preference parameters, PBOS aims to guide agents toward cooperative or competitive behaviors without relying on simple opponent predictions. Experiments on differentiable games demonstrate that PBOS enables agents to achieve better reward distributions across various environments.
UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.
The paper investigates zero‑shot task generalisation in offline multi‑agent reinforcement learning by extending sequence‑modeling architectures to support multi‑task observation and action spaces and variable agent counts. It finds that increasing task diversity, rather than merely enlarging the dataset, is the key driver for robust zero‑shot transfer. Experiments on four challenging environments show a 3.2× mean improvement on held‑out tasks compared to single‑task models and outperform strong behaviour‑cloning baselines.
arXiv:2506. 13741v2 Announce Type: replace-cross Abstract: Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefined reward functions.