arXiv Machine Learning

LLMs for Zero-Shot Threat Detection via Structured Risk Indicators

arXiv:2608. 16508v1 Announce Type: cross Abstract: We propose a two-stage large language model (LLM) framework for zero-shot detection of insider threats and advanced persistent threats (APTs) from heterogeneous security logs.

arXiv AI
Sep 17

MiST: Mid-Training LLMs for Cybersecurity

MiST (Mid-trained Security Transformer) is a suite of 8B and 32B language models tailored for cybersecurity, achieving strong performance on public benchmarks. The approach uses a mid-training stage that adapts general pre-trained models to the domain by curating a compact, expert-vetted seed corpus and generating high-quality synthetic training data, rather than continual pre-training on large raw text. MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over Qwen baselines for 8B and 32B models, respectively, and provide a stronger initialization for downstream task-specific fine-tuning and reinforcement learning.

By Oded Ovadia, Elad Ben Zaken, Elad Guttman, Orly Moreno Kadosh
arXiv Machine Learning
Aug 21

From Noise to Signal: Improving Security Log Anomaly Detection Using LLMs with Endpoint-Specific Logs

arXiv:2608. 19938v1 Announce Type: cross Abstract: Existing approaches to anomalous behaviour log detection, such as Wazuh rely primarily on predefined detection rules, while statistical anomaly detection approaches such as OpenSearch identify deviations from previously observed behavioural patterns.

By Christopher Henshaw, Gour Karmakar
arXiv AI
Sep 10

Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports

The paper introduces AHLERT, a system that automatically extracts environment-aware hunt leads from Cyber Threat Intelligence reports. It combines a hybrid retriever—dense vector search plus multi-hop knowledge‑graph traversal seeded with MITRE ATT&CK—with ontology‑grounded retrieval‑augmented generation to constrain leads to a defender’s assets. Evaluations on public CTI reports show that AHLERT doubles mean F1 scores and achieves an effectiveness score of ~86.95% compared to off‑the‑shelf LLM models.

By Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi
arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche
arXiv Computation and Language
Aug 27

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots

The paper introduces a three‑layer security framework designed to protect retrieval‑augmented generation (RAG) chatbots from both direct and indirect prompt injection attacks. Layer 1 filters user input with rule‑based patterns and a semantic anomaly classifier; Layer 2 enforces a provenance‑based instruction hierarchy during context assembly; Layer 3 audits model output with a policy rule engine and semantic drift detector. Evaluations on GPT‑4o, Llama 3, and Mistral 7B demonstrate a reduction in attack success rate from 71.4 % to 11.3 %, outperforming existing single‑layer defenses while keeping false positives low and latency acceptable.

By Gulshan Saleem, Nisar Ahmed, Muhammad Imran Zaman, Ali Hassan, Umar Mujahid