arXiv AI

Polar: LLM-Powered Synthesis of Real-World Cyber Evidence for Prioritization and Mitigation

arXiv AI
1d ago

From Sandbox to Enforcement: Confidence-Qualified Threat Intelligence for Critical Infrastructure

arXiv:2610. 07310v1 Announce Type: cross Abstract: Security operations centres and national incident-response teams defending critical infrastructure collect abundant threat data yet struggle to turn it into actionable intelligence.

By Nikolaos Kekatos, Mihaela Curc\u{a}, Georgios Koutidis, Mihai Nena, Tom Nianios, Robert-\c{S}tefan \c{S}andru, Michael Ioannou, Charalambos Bratsas
arXiv AI
Aug 5

DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

arXiv:2608. 03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions.

By Xuyang Liu, Yibin Han, Zhenwei Zhang, Kai Chang, Zhiwei Xu, Tian Qiu, Weixian Deng, Jiabao Gao, Xiaolin Peng, Hai Wan, Xibin Zhao
arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche
arXiv AI
Aug 19

Future-Back Threat Modeling: A Foresight-Driven Security Framework

Future-Back Threat Modeling (FBTM) is a predictive security framework that starts with envisioned future threat states and works backward to uncover assumptions, gaps, blind spots, and vulnerabilities in current defense architectures. It aims to reveal both known unknowns and unknown unknowns, including emerging tactics, techniques, and procedures, thereby improving the predictability of adversary behavior under future uncertainty. By anticipating future threats such as AI, information warfare, and supply chain attacks, FBTM helps security leaders make informed decisions today to build more resilient security postures for the future.

By Vu Van Than
arXiv AI
Jul 31

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.

By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
arXiv AI
Aug 24

Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making

The paper investigates whether large language models (LLMs) can perform structured security reasoning in cybersecurity decision-making. By testing LLMs on defense selection over attack graphs from real-world threat scenarios, the study finds that LLMs can produce coherent strategies when the attack-graph structure is explicitly provided, yet their performance is fragile, highly sensitive to prompt framing, and deteriorates with increasing graph complexity. Additionally, LLM-generated solvers recover the correct high-level formulation but scale poorly compared to specialized solvers.

By Pasquale Malacaria, Yunxiao Zhang
arXiv AI
Jun 17

Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

arXiv:2604. 09998v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasingly marketing autonomous AI solutions for SOCs.

By Souradip Nath, Chih-Yi Huang, Aditi Ganapathi, Kashyap Thimmaraju, Jaron Mink, Gail-Joon Ahn