arXiv Machine Learning

Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition

The paper introduces a Noise Contrastive Estimation (NCE)-based matching framework for recognizing low‑resource security attack patterns, specifically Tactics, Techniques, and Procedures (TTPs). It reframes TTP mapping as a semantic similarity matching problem rather than a traditional multi‑class classification, thereby mitigating issues of large label sets, skewed distributions, and hierarchical complexity. The proposed neural architecture employs a sampling‑based learn‑to‑compare mechanism with two objectives: an α‑balanced NCE to manage the overall mass of sampled negatives and an asymmetric focusing objective to handle incomplete annotations during individual comparisons.

arXiv AI
Sep 17

MiST: Mid-Training LLMs for Cybersecurity

MiST (Mid-trained Security Transformer) is a suite of 8B and 32B language models tailored for cybersecurity, achieving strong performance on public benchmarks. The approach uses a mid-training stage that adapts general pre-trained models to the domain by curating a compact, expert-vetted seed corpus and generating high-quality synthetic training data, rather than continual pre-training on large raw text. MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over Qwen baselines for 8B and 32B models, respectively, and provide a stronger initialization for downstream task-specific fine-tuning and reinforcement learning.

By Oded Ovadia, Elad Ben Zaken, Elad Guttman, Orly Moreno Kadosh
arXiv Machine Learning
1d ago

SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

SAGE is a defense against clean‑label data poisoning that relies on a very small set of verified examples—both clean and poisoned—rather than a large clean set. It trains a generic feature extractor on a separate dataset and then uses a non‑parametric, similarity‑weighted prediction to flag poisoned training examples. Experiments on standard benchmarks show that even a handful of verified poisoned examples give a substantial advantage, and that the distribution of verified clean examples across classes is more important than their sheer number.

By Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka
arXiv Computation and Language
Sep 1

OASIS: Optimizing Attacker Sequences for Hard-Label Black-Box Text Attacks

OASIS is a method for optimizing attacker sequences in hard‑label black‑box text attacks. It first performs a one‑time bi‑objective search over candidate sequences to balance attack success rate and perturbation, then reuses the selected fixed global chain during execution. Experiments on multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains.

By Qian Chen, Shiliang Xiao, Yuzhi Liang
arXiv AI
Jul 3

Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens

arXiv:2507. 02964v2 Announce Type: replace-cross Abstract: The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scalable and sustainable.

By Salahuddin Salahuddin, Ahmed Hussain, Jussi L\"opp\"onen, Toni Jutila