The paper explores a runtime strategy-selection framework where a large language model (LLM) guides a pre‑trained reinforcement learning (RL) policy for non‑player characters (NPCs) in a Unity combat game without altering the underlying policy. Five NPC agents sharing a PPO policy were compared in a baseline setup and an LLM‑augmented setup, where a locally hosted Mistral 7B model assigns one of four tactical tags every five seconds based on live game state. Across 600 episodes against three scripted opponents, the LLM‑augmented agents more than doubled their win rate against a Balanced opponent, improved performance against an Evasive opponent, but struggled against an Aggressive opponent due to over‑reliance on encirclement; analysis of 2,430 strategy selections revealed limited zero‑shot differentiation with the model favoring Surround in 83.8% of cases.
By Hrithika Deepu Nair, Kayvan Karim
The paper introduces a curriculum reinforcement learning approach to overcome the cold‑start problem in prompt‑injection red‑teaming of frontier large language models. By training an attacker LLM sequentially against increasingly robust target models and ensuring partial success at each stage, the method achieves high attack success rates (93.8% against GPT‑5.6‑Luna and 45.0% against GPT‑5.6‑Terra) where prior RL methods fail. The attacker LLM also transfers its effectiveness to other frontier models it was not explicitly trained on.
By Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia
The paper introduces Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that selectively simulates only payoff-matrix cells relevant to a Nash equilibrium and uses a surrogate model for the remaining entries. RWPS provides an instance-dependent error bound weighted by the opponent’s equilibrium mixture and a coverage result guaranteeing that, once the deviation-relevant set is simulated, surrogate error does not affect either player’s regret. Experiments on three 21×21 general-sum games, including an asymmetric Colonel Blotto, show that RWPS achieves four to six times tighter bounds than previous methods and outperforms other sampling strategies on the CyGym and ANSG cyber simulators at low budgets.
By Michael Lanier, David Farmer, Yevgeniy Vorobeychik
arXiv:2608. 04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied.
By Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi, Sanggeon Yun, Hyunwoo Oh, SungHeon Jeong, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani
The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.
By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments and interaction datasets.
The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.
By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv:2609.38889v1 Announce Type: new
Abstract: Constrained multi-agent control requires more than predicting rewarding actions: an action can cease to be executable as contact windows, shared capaci...
By Bo Yin, Dongbo Li, Hongkai Chen, Jie Liu, Guoliang Xing
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
The paper studies a stationary decentralized Markov game where a focal agent experiences drifting rewards and dynamics due to learning peers, framing this as an agent‑centric continual reinforcement‑learning problem. It introduces the concept of an invariant core—maximal abstract patterns common to many successful trajectories—and proves a worst‑case conditioning theorem linking trajectory‑law drift to success coverage. The authors provide theoretical guarantees for survival horizon, first‑exit law, and regret, and validate their predictions with solvable models and empirical studies in continual control, cue‑MNIST, and Level‑Based Foraging.
By Dane Malenfant
arXiv:2603. 13026v2 Announce Type: replace Abstract: Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents.
By Chenlong Yin, Runpeng Geng, Yanting Wang, Jinyuan Jia
arXiv:2609.34373v2 Announce Type: replace
Abstract: Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message s...
By Mihir Chauhan, Aniket Bera