arXiv AI By Jose Luis Lima de Jesus Silva

Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security Response

Read the original on arXiv AI →

arXiv:2606. 13832v1 Announce Type: cross Abstract: Autonomous network-security response systems promise to reduce Security Operations Centre (SOC) reaction latency, but reward-only multi-agent reinforcement learning (MARL) can improve security reward while remaining non-deployable.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.

By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
arXiv Machine Learning
Sep 23

Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry

The paper introduces Network Feasibility Geometry Reinforcement Learning (NFG‑RL), a method that enforces multi‑layer network constraints—such as interference, power‑rate coupling, flow conservation, service chains, capacity, latency, and reliability—by transporting a proto‑policy through a differentiable feasibility map. By compiling heterogeneous constraints into typed residual blocks and using a variational transport operator, NFG‑RL ensures almost‑sure feasible execution and shapes exploration and gradients to respect active constraints. Experiments on two wireless‑edge surrogate environments show that NFG‑RL boosts feasible utility by 37.5–41.5 %, cuts raw‑action violations by 48.5–60.8 %, and reduces P99 delay by 57.0–75.5 % compared to leading baselines.

By Zuyuan Zhang, Zeyu Fang, Mahdi Imani, Nathaniel D. Bastian, Tian Lan
arXiv Machine Learning
Aug 3

Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

arXiv:2607. 28826v1 Announce Type: new Abstract: Autonomous Cyber Operations (ACO) are increasingly important for defending enterprise networks as cyber threats continue to evolve in sophistication.

By Konur Tholl, Fran\c{c}ois Rivest, Mariam El Mezouar, Adrian Taylor, Ranwa Al Mallah
arXiv AI
Aug 6

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

arXiv:2608. 04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied.

By Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi, Sanggeon Yun, Hyunwoo Oh, SungHeon Jeong, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani