arXiv AI By Duc Huy Le, Rolf Stadler

Learning Intrusion Response Strategies for OT Systems

Read the original on arXiv AI →

The paper presents a formal model for responding to cyber intrusions in Operational Technology (OT) systems using a Partially Observable Markov Decision Process (POMDP) framework. It incorporates realistic partial observability derived from traffic measurements and develops learning‑based solution methods based on Proximal Policy Optimization (PPO). The resulting response strategies are evaluated on an emulated OT system and shown to be effective against several MITRE attack types for the studied use case.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
5d ago

Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System

The paper introduces a new method for generating universal adversarial perturbations (UAPs) against deep reinforcement learning (DRL)-based intrusion detection systems (IDS). It leverages Probabilistic Robustness (PR) as a post‑hoc metric to guide UAP creation, integrating PR directly into the optimization objective. The authors further develop PX‑UAP, which incorporates explainable AI (XAI) to shape perturbations within realistic domain constraints, and provide a theoretical analysis of its design. Experiments show PX‑UAP outperforms existing UAP techniques in attack effectiveness.

By Hongsen Zhang, Lu Zhang, Mingjing Xu, Yi Zhang, Gregory Epiphaniou, Carsten Maple
arXiv AI
Sep 16

A Cyber Range Evaluation of Autonomous Network Incident Response Agents

The study evaluates autonomous agents that respond to network intrusions within a cyber range designed for human operator training. Using an emulated network with variable topology, red‑team attacks, and simulated users, the agents aim to block unauthorized access while minimizing defensive costs. Experiments compare heuristic policies with reinforcement‑learning‑derived policies, finding that the latter generally defend more efficiently, though performance varies with adversary strategy and user simulation.

By Jakob Nyberg, Teodor Sommestad, Andrei Buhaiu, Joakim Loxdal, Pontus Johnson, Mathias Ekstedt