arXiv AI By Phillip Jiang

UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms

Read the original on arXiv AI →

arXiv:2608. 15372v1 Announce Type: new Abstract: We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games

The paper explores a runtime strategy-selection framework where a large language model (LLM) guides a pre‑trained reinforcement learning (RL) policy for non‑player characters (NPCs) in a Unity combat game without altering the underlying policy. Five NPC agents sharing a PPO policy were compared in a baseline setup and an LLM‑augmented setup, where a locally hosted Mistral 7B model assigns one of four tactical tags every five seconds based on live game state. Across 600 episodes against three scripted opponents, the LLM‑augmented agents more than doubled their win rate against a Balanced opponent, improved performance against an Evasive opponent, but struggled against an Aggressive opponent due to over‑reliance on encirclement; analysis of 2,430 strategy selections revealed limited zero‑shot differentiation with the model favoring Surround in 83.8% of cases.

By Hrithika Deepu Nair, Kayvan Karim
arXiv Machine Learning
4d ago

Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

The paper introduces a curriculum reinforcement learning approach to overcome the cold‑start problem in prompt‑injection red‑teaming of frontier large language models. By training an attacker LLM sequentially against increasingly robust target models and ensuring partial success at each stage, the method achieves high attack success rates (93.8% against GPT‑5.6‑Luna and 45.0% against GPT‑5.6‑Terra) where prior RL methods fail. The attacker LLM also transfers its effectiveness to other frontier models it was not explicitly trained on.

By Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia
arXiv AI
Sep 18

Efficient Nash Equilibrium Computation for Cybersecurity Games

The paper introduces Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that selectively simulates only payoff-matrix cells relevant to a Nash equilibrium and uses a surrogate model for the remaining entries. RWPS provides an instance-dependent error bound weighted by the opponent’s equilibrium mixture and a coverage result guaranteeing that, once the deviation-relevant set is simulated, surrogate error does not affect either player’s regret. Experiments on three 21×21 general-sum games, including an asymmetric Colonel Blotto, show that RWPS achieves four to six times tighter bounds than previous methods and outperforms other sampling strategies on the CyGym and ANSG cyber simulators at low budgets.

By Michael Lanier, David Farmer, Yevgeniy Vorobeychik
arXiv AI
Aug 6

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

arXiv:2608. 04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied.

By Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi, Sanggeon Yun, Hyunwoo Oh, SungHeon Jeong, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani
arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
Hugging Face Trending Papers
Aug 5

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments and interaction datasets.