arXiv AI By Shaswata Mitra, Subash Neupane, Trisha Chakraborty, Himanshu Tripathi, Sudip Mittal, Aritran Piplai, Shahram Rahimi

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA

Read the original on arXiv AI →

arXiv:2607. 18725v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to adapt, before paying the cost of adaptation, remains difficult.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche
arXiv Machine Learning
Aug 20

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model

MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. The authors also introduce MITRE‑QA, a benchmark of 3,000 question‑answer pairs, and show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight configuration achieving top performance on most tasks.

By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
arXiv Machine Learning
Aug 19

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model

MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. Experiments show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight Qwen2.5‑based configuration excelling on most benchmark tasks.

By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani