arXiv AI

A Knowledge-Based Multi-Agent Framework for Security Control Recommendation

arXiv:2607. 09954v1 Announce Type: cross Abstract: Hardening IT on-premises environments can be a daunting task for teams without access to adequate cybersecurity expertise.

arXiv AI
Sep 11

Dont Just Teach, Explain! A Gamified 20Q Recommender for Cybersecurity Education

The paper presents a gamified 20Q-style recommender for cybersecurity education that uses reinforcement learning and explainable AI to guide learners through interactive questioning. By acting as a knowledgeable questioner, the system narrows down user-described security scenarios, identifies the underlying threat, and transparently explains its reasoning. The authors detail the system architecture, algorithmic foundations, and provide case studies covering attack vectors such as the Cyber Kill Chain, phishing, ransomware, and web application vulnerabilities.

By Mary Nusrat, Sarfuddin Bhuiyan, Gahangir Hossain
arXiv AI
Aug 24

Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making

The paper investigates whether large language models (LLMs) can perform structured security reasoning in cybersecurity decision-making. By testing LLMs on defense selection over attack graphs from real-world threat scenarios, the study finds that LLMs can produce coherent strategies when the attack-graph structure is explicitly provided, yet their performance is fragile, highly sensitive to prompt framing, and deteriorates with increasing graph complexity. Additionally, LLM-generated solvers recover the correct high-level formulation but scale poorly compared to specialized solvers.

By Pasquale Malacaria, Yunxiao Zhang
arXiv AI
Jun 12

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

arXiv:2606. 13079v1 Announce Type: cross Abstract: Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross.

By Jiaqi Luo, Jiarun Dai, Zhile Chen, Jia Xu, Weibing Wang, Yawen Duan, Brian Tse, Geng Hong, Xudong Pan, Yuan Zhang, Min Yang
arXiv AI
Sep 15

Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents

The paper evaluates a task-based permission scoping architecture for AI agents, comparing a fine‑tuned RoBERTa‑large encoder to few‑shot Claude Haiku 4.5 on a 600‑prompt dataset. It shows the new system achieves comparable macro‑F1 (0.881 vs. 0.886) and higher precision (0.897 vs. 0.842), while reducing severity‑weighted residual risk from 1.12 to 0.63. The study also introduces an attack‑surface elimination metric, demonstrating that task‑granular control can close 84.4% of the severity‑weighted surface, far surpassing role‑based ceilings alone.

By Halil Burak Noyan
arXiv Machine Learning
Aug 3

Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

arXiv:2607. 28826v1 Announce Type: new Abstract: Autonomous Cyber Operations (ACO) are increasingly important for defending enterprise networks as cyber threats continue to evolve in sophistication.

By Konur Tholl, Fran\c{c}ois Rivest, Mariam El Mezouar, Adrian Taylor, Ranwa Al Mallah