arXiv Machine Learning By Tu Nguyen, Nedim \v{S}rndi\'c, Alexander Neth

Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition

Read the original on arXiv Machine Learning →

The paper introduces a Noise Contrastive Estimation (NCE)-based matching framework for recognizing low‑resource security attack patterns, specifically Tactics, Techniques, and Procedures (TTPs). It reframes TTP mapping as a semantic similarity matching problem rather than a traditional multi‑class classification, thereby mitigating issues of large label sets, skewed distributions, and hierarchical complexity. The proposed neural architecture employs a sampling‑based learn‑to‑compare mechanism with two objectives: an α‑balanced NCE to manage the overall mass of sampled negatives and an asymmetric focusing objective to handle incomplete annotations during individual comparisons.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

MiST: Mid-Training LLMs for Cybersecurity

MiST (Mid-trained Security Transformer) is a suite of 8B and 32B language models tailored for cybersecurity, achieving strong performance on public benchmarks. The approach uses a mid-training stage that adapts general pre-trained models to the domain by curating a compact, expert-vetted seed corpus and generating high-quality synthetic training data, rather than continual pre-training on large raw text. MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over Qwen baselines for 8B and 32B models, respectively, and provide a stronger initialization for downstream task-specific fine-tuning and reinforcement learning.

By Oded Ovadia, Elad Ben Zaken, Elad Guttman, Orly Moreno Kadosh
arXiv Machine Learning
1d ago

SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

SAGE is a defense against clean‑label data poisoning that relies on a very small set of verified examples—both clean and poisoned—rather than a large clean set. It trains a generic feature extractor on a separate dataset and then uses a non‑parametric, similarity‑weighted prediction to flag poisoned training examples. Experiments on standard benchmarks show that even a handful of verified poisoned examples give a substantial advantage, and that the distribution of verified clean examples across classes is more important than their sheer number.

By Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka
arXiv Computation and Language
Sep 1

OASIS: Optimizing Attacker Sequences for Hard-Label Black-Box Text Attacks

OASIS is a method for optimizing attacker sequences in hard‑label black‑box text attacks. It first performs a one‑time bi‑objective search over candidate sequences to balance attack success rate and perturbation, then reuses the selected fixed global chain during execution. Experiments on multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains.

By Qian Chen, Shiliang Xiao, Yuzhi Liang