arXiv AI By Saeedeh Lohrasbi, Mohammad Mamun, Ahmed Yehia, Scott Buffett, Sherif Saad

Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents

Read the original on arXiv AI →

The paper presents a diagnostic study of a multi‑stage LLM‑based cyber agent system, examining its orchestrator, executor, and validator components in enterprise‑style lateral‑movement scenarios. Six advanced LLMs were tested across expert‑defined, self‑scaffolded, and fully autonomous modes, with metrics that include validator consistency, evidence grounding, token usage, retries, and runtime. Findings show that while validators are generally relevant, they are often nonspecific and overly optimistic, and the main bottlenecks lie in credential acquisition and lateral‑movement tasks, especially under full autonomy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 11

RangeFactory: Scalable Construction of Multi-Hop Cyber Ranges

arXiv:2608. 09526v1 Announce Type: cross Abstract: Real-world cyberattacks often require sustained progress across multiple hosts and network segments, making multi-hop cyber ranges essential infrastructure for studying and improving LLM agents' ability to sustain complete attack chains.

By Hanlin Jiang, Puyi Wang, Jiandong Jin, Shaofei Li, Zhan Shen, Pengli Wang, Ziming Wang, Yifeng Cai, Ning Jia, Yuxin Ren, Peng Jiang, Yao Guo, Ding Li
arXiv AI
Jun 12

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

arXiv:2606. 13079v1 Announce Type: cross Abstract: Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross.

By Jiaqi Luo, Jiarun Dai, Zhile Chen, Jia Xu, Weibing Wang, Yawen Duan, Brian Tse, Geng Hong, Xudong Pan, Yuan Zhang, Min Yang