arXiv AI

CTIConnect: A Benchmark for Retrieval-Augmented LLMs over Heterogeneous Cyber Threat Intelligence

arXiv:2510. 11974v2 Announce Type: replace-cross Abstract: Cyber Threat Intelligence (CTI) is foundational to modern cybersecurity, enabling organizations to proactively defend against evolving threats.

arXiv AI
Sep 10

Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports

The paper introduces AHLERT, a system that automatically extracts environment-aware hunt leads from Cyber Threat Intelligence reports. It combines a hybrid retriever—dense vector search plus multi-hop knowledge‑graph traversal seeded with MITRE ATT&CK—with ontology‑grounded retrieval‑augmented generation to constrain leads to a defender’s assets. Evaluations on public CTI reports show that AHLERT doubles mean F1 scores and achieves an effectiveness score of ~86.95% compared to off‑the‑shelf LLM models.

By Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi
arXiv Machine Learning
Aug 20

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model

MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. The authors also introduce MITRE‑QA, a benchmark of 3,000 question‑answer pairs, and show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight configuration achieving top performance on most tasks.

By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
arXiv Machine Learning
Aug 19

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model

MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. Experiments show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight Qwen2.5‑based configuration excelling on most benchmark tasks.

By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
arXiv AI
Aug 20

CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence

CTIFoundry is an agent‑native corpus scaffold designed to improve cyber threat intelligence (CTI) investigations by LLM agents. It transforms traditional CTI data—such as CVE, CWE, CAPEC, and ATT&CK—into a deterministic ontology graph with typed, traversable edges, a span‑grounded report layer that resolves entity aliases and provenance, and hybrid dense‑plus‑lexical retrieval surfaces. When integrated with a standard open‑source agent harness, CTIFoundry boosts overall F1 scores by 0.19 to 0.28 on the CTIConnect benchmark, achieving higher accuracy with fewer tool calls compared to agents using flat, retrieval‑augmented corpora.

By Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao
arXiv Computation and Language
Aug 31

DisCTI: Who Needs to Know Timely? Automated Sector-Aware Cyber Threat Intelligence Dissemination

The paper introduces DisCTI, a system that automatically maps cyber threat intelligence (CTI) events to relevant industry sectors using a multilabel classification approach. By creating a dataset of 872 sector‑labelled CTI events and applying a BERT transformer model, the authors achieve a macro‑averaged F1‑score of 0.89, correctly assigning 94.5% of sector labels. This demonstrates that embedding expert knowledge into machine learning can enable timely, sector‑aware CTI dissemination, improving defensive response.

By Fajar Wijitrisnanto (National Cyber and Crypto Agency, Jakarta, Indonesia), Alsharif Abuadbba (CSIRO, Sydney, Australia), Yansong Gao (CSIRO, Sydney, Australia, The University of Western Australia, Perth, Australia), Nan Wu (CSIRO, Sydney, Australia)
arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche
arXiv Machine Learning
Jun 17

Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports

arXiv:2606. 18166v1 Announce Type: cross Abstract: Classifying Cyber Threat Intelligence (CTI) using MITRE Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) is essential for proactive defense, but historically required extensive human effort.

By Ahmed Ryan, Saad Sakib Noor, Md Erfan, Shaswata Mitra, Sudip Mittal, Md Rayhanur Rahman
arXiv AI
Jul 22

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA

arXiv:2607. 18725v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to adapt, before paying the cost of adaptation, remains difficult.

By Shaswata Mitra, Subash Neupane, Trisha Chakraborty, Himanshu Tripathi, Sudip Mittal, Aritran Piplai, Shahram Rahimi