arXiv AI By Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao

CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence

Read the original on arXiv AI →

CTIFoundry is an agent‑native corpus scaffold designed to improve cyber threat intelligence (CTI) investigations by LLM agents. It transforms traditional CTI data—such as CVE, CWE, CAPEC, and ATT&CK—into a deterministic ontology graph with typed, traversable edges, a span‑grounded report layer that resolves entity aliases and provenance, and hybrid dense‑plus‑lexical retrieval surfaces. When integrated with a standard open‑source agent harness, CTIFoundry boosts overall F1 scores by 0.19 to 0.28 on the CTIConnect benchmark, achieving higher accuracy with fewer tool calls compared to agents using flat, retrieval‑augmented corpora.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 7

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

arXiv:2607. 03233v1 Announce Type: cross Abstract: The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation.

By Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche
arXiv Machine Learning
Jun 17

Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports

arXiv:2606. 18166v1 Announce Type: cross Abstract: Classifying Cyber Threat Intelligence (CTI) using MITRE Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) is essential for proactive defense, but historically required extensive human effort.

By Ahmed Ryan, Saad Sakib Noor, Md Erfan, Shaswata Mitra, Sudip Mittal, Md Rayhanur Rahman
arXiv Machine Learning
2d ago

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model

MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. The authors also introduce MITRE‑QA, a benchmark of 3,000 question‑answer pairs, and show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight configuration achieving top performance on most tasks.

By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani