arXiv:2608.28394v1 Announce Type: cross
Abstract: Cyber threat intelligence (CTI) is foundational to modern cyber defense, yet much of it resides in unstructured reports whose volume and heterogeneit...
By Changze Li, Yutong Cheng, Tsania Camila Finnisa, Qian Cui, Wei Ding, Peng Gao
The paper introduces AHLERT, a system that automatically extracts environment-aware hunt leads from Cyber Threat Intelligence reports. It combines a hybrid retriever—dense vector search plus multi-hop knowledge‑graph traversal seeded with MITRE ATT&CK—with ontology‑grounded retrieval‑augmented generation to constrain leads to a defender’s assets. Evaluations on public CTI reports show that AHLERT doubles mean F1 scores and achieves an effectiveness score of ~86.95% compared to off‑the‑shelf LLM models.
By Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi
MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. The authors also introduce MITRE‑QA, a benchmark of 3,000 question‑answer pairs, and show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight configuration achieving top performance on most tasks.
By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. Experiments show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight Qwen2.5‑based configuration excelling on most benchmark tasks.
By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
CTIFoundry is an agent‑native corpus scaffold designed to improve cyber threat intelligence (CTI) investigations by LLM agents. It transforms traditional CTI data—such as CVE, CWE, CAPEC, and ATT&CK—into a deterministic ontology graph with typed, traversable edges, a span‑grounded report layer that resolves entity aliases and provenance, and hybrid dense‑plus‑lexical retrieval surfaces. When integrated with a standard open‑source agent harness, CTIFoundry boosts overall F1 scores by 0.19 to 0.28 on the CTIConnect benchmark, achieving higher accuracy with fewer tool calls compared to agents using flat, retrieval‑augmented corpora.
By Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao
The paper introduces DisCTI, a system that automatically maps cyber threat intelligence (CTI) events to relevant industry sectors using a multilabel classification approach. By creating a dataset of 872 sector‑labelled CTI events and applying a BERT transformer model, the authors achieve a macro‑averaged F1‑score of 0.89, correctly assigning 94.5% of sector labels. This demonstrates that embedding expert knowledge into machine learning can enable timely, sector‑aware CTI dissemination, improving defensive response.
By Fajar Wijitrisnanto (National Cyber and Crypto Agency, Jakarta, Indonesia), Alsharif Abuadbba (CSIRO, Sydney, Australia), Yansong Gao (CSIRO, Sydney, Australia, The University of Western Australia, Perth, Australia), Nan Wu (CSIRO, Sydney, Australia)