arXiv AI

SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification

arXiv:2607. 09936v1 Announce Type: cross Abstract: Cybersecurity systems must adapt rapidly to emerging threats.

arXiv AI
Sep 17

MiST: Mid-Training LLMs for Cybersecurity

MiST (Mid-trained Security Transformer) is a suite of 8B and 32B language models tailored for cybersecurity, achieving strong performance on public benchmarks. The approach uses a mid-training stage that adapts general pre-trained models to the domain by curating a compact, expert-vetted seed corpus and generating high-quality synthetic training data, rather than continual pre-training on large raw text. MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over Qwen baselines for 8B and 32B models, respectively, and provide a stronger initialization for downstream task-specific fine-tuning and reinforcement learning.

By Oded Ovadia, Elad Ben Zaken, Elad Guttman, Orly Moreno Kadosh
arXiv Machine Learning
Jun 17

Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports

arXiv:2606. 18166v1 Announce Type: cross Abstract: Classifying Cyber Threat Intelligence (CTI) using MITRE Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) is essential for proactive defense, but historically required extensive human effort.

By Ahmed Ryan, Saad Sakib Noor, Md Erfan, Shaswata Mitra, Sudip Mittal, Md Rayhanur Rahman
arXiv AI
Jun 16

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.

By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche
arXiv Computation and Language
Aug 31

DisCTI: Who Needs to Know Timely? Automated Sector-Aware Cyber Threat Intelligence Dissemination

The paper introduces DisCTI, a system that automatically maps cyber threat intelligence (CTI) events to relevant industry sectors using a multilabel classification approach. By creating a dataset of 872 sector‑labelled CTI events and applying a BERT transformer model, the authors achieve a macro‑averaged F1‑score of 0.89, correctly assigning 94.5% of sector labels. This demonstrates that embedding expert knowledge into machine learning can enable timely, sector‑aware CTI dissemination, improving defensive response.

By Fajar Wijitrisnanto (National Cyber and Crypto Agency, Jakarta, Indonesia), Alsharif Abuadbba (CSIRO, Sydney, Australia), Yansong Gao (CSIRO, Sydney, Australia, The University of Western Australia, Perth, Australia), Nan Wu (CSIRO, Sydney, Australia)
arXiv AI
Aug 3

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

arXiv:2607. 28862v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage.

By Chengshuai Zhao, Pingchuan Ma, Dawei Li, Bohan Jiang, Zhiyuan Yu, Zhen Tan, Huan Liu