SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification
arXiv:2607. 09936v1 Announce Type: cross Abstract: Cybersecurity systems must adapt rapidly to emerging threats.
The paper introduces a Noise Contrastive Estimation (NCE)-based matching framework for recognizing low‑resource security attack patterns, specifically Tactics, Techniques, and Procedures (TTPs). It reframes TTP mapping as a semantic similarity matching problem rather than a traditional multi‑class classification, thereby mitigating issues of large label sets, skewed distributions, and hierarchical complexity. The proposed neural architecture employs a sampling‑based learn‑to‑compare mechanism with two objectives: an α‑balanced NCE to manage the overall mass of sampled negatives and an asymmetric focusing objective to handle incomplete annotations during individual comparisons.
arXiv:2607. 09936v1 Announce Type: cross Abstract: Cybersecurity systems must adapt rapidly to emerging threats.
arXiv:2607. 11994v1 Announce Type: cross Abstract: Classifying cybersecurity vulnerabilities using the Common Weakness Enumeration (CWE) taxonomy is challenging due to extreme class imbalance and strong hierarchical dependencies among weakness categories.
arXiv:2409. 08521v2 Announce Type: replace-cross Abstract: In cybersecurity practice, new forms of cyberattacks continuously emerge, deliberately designed to evade defense systems that rely on previously observed behaviors.
MiST (Mid-trained Security Transformer) is a suite of 8B and 32B language models tailored for cybersecurity, achieving strong performance on public benchmarks. The approach uses a mid-training stage that adapts general pre-trained models to the domain by curating a compact, expert-vetted seed corpus and generating high-quality synthetic training data, rather than continual pre-training on large raw text. MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over Qwen baselines for 8B and 32B models, respectively, and provide a stronger initialization for downstream task-specific fine-tuning and reinforcement learning.
SAGE is a defense against clean‑label data poisoning that relies on a very small set of verified examples—both clean and poisoned—rather than a large clean set. It trains a generic feature extractor on a separate dataset and then uses a non‑parametric, similarity‑weighted prediction to flag poisoned training examples. Experiments on standard benchmarks show that even a handful of verified poisoned examples give a substantial advantage, and that the distribution of verified clean examples across classes is more important than their sheer number.
OASIS is a method for optimizing attacker sequences in hard‑label black‑box text attacks. It first performs a one‑time bi‑objective search over candidate sequences to balance attack success rate and perturbation, then reuses the selected fixed global chain during execution. Experiments on multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains.
arXiv:2601. 14300v4 Announce Type: replace Abstract: Hard-label black-box attacks, relying solely on top-1 predictions, represent one of the most challenging yet practically threat models.
arXiv:2606. 28953v1 Announce Type: cross Abstract: Poisoning attacks entail attackers intentionally tampering with training data.
arXiv:2608.28394v1 Announce Type: cross Abstract: Cyber threat intelligence (CTI) is foundational to modern cyber defense, yet much of it resides in unstructured reports whose volume and heterogeneit...
arXiv:2507. 02964v2 Announce Type: replace-cross Abstract: The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scalable and sustainable.
arXiv:2602. 18934v2 Announce Type: replace Abstract: Membership inference attacks (MIAs) threaten the privacy of machine learning models by revealing whether a specific data point was used during training.
arXiv:2512. 05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their "open" nature introduces more avenues for malicious actors to misuse them for harmful purposes.