arXiv Computation and Language

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild

arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche
arXiv AI
Jun 12

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

arXiv:2606. 13079v1 Announce Type: cross Abstract: Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross.

By Jiaqi Luo, Jiarun Dai, Zhile Chen, Jia Xu, Weibing Wang, Yawen Duan, Brian Tse, Geng Hong, Xudong Pan, Yuan Zhang, Min Yang
arXiv AI
Jul 3

Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens

arXiv:2507. 02964v2 Announce Type: replace-cross Abstract: The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scalable and sustainable.

By Salahuddin Salahuddin, Ahmed Hussain, Jussi L\"opp\"onen, Toni Jutila
arXiv AI
Jul 9

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

arXiv:2607. 06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing cybersecurity, enabling both automated defense and sophisticated attacks.

By Kiarash Ahi, Saeed Valizadeh
arXiv AI
Sep 10

Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models

Feyospace‑v1 describes a data‑centric framework that enables a seven‑person team to train open‑weight cyber agents, addressing bottlenecks such as environment cost, supervision, and teacher access. The system combines five complementary components—Choulea, SkyReal, Hongzwang, PSBreakup, and Kreator—to analyze reasoning, reduce sampling costs, bypass API limits, restore merged model capabilities, and convert expert interventions into trainable signals. Using a diverse data engine, the team produced 164,269 verified trajectories for supervised fine‑tuning, achieving an average 23.76% improvement on CyberGym and 10.49% on pooled CTF suites, with the final checkpoint ranking 10th on the CyberGym leaderboard and first among comparable‑scale models.

By Zongjie Li, Alan Z. W, John Nicolas J, Walter H. F, Scott Donald L, Gordon Y. P, Deke X Jr