arXiv AI By Vadym Hadetskyi, Dario Pasquini, Artem Sorokin

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

Read the original on arXiv AI →

arXiv:2607. 02714v1 Announce Type: cross Abstract: There is no doubt that safety alignment is an essential step in LLM training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 8

Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis

arXiv:2607. 05842v1 Announce Type: cross Abstract: Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for legitimate code review, triage, and repair can closely resemble terminology associated with misuse.

By Mingchen Li, Meikang Qiu, Zifan Peng, Heng Fan, Song Fu, Junhua Ding, Yunhe Feng
arXiv AI
Sep 3

FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

The paper introduces FUSE, a modular framework that evaluates large language models (LLMs) for dangerous capabilities across three orthogonal pipelines: Knowledge (K), Defense (D), and Harm (H). Using a chemical‑biological module, the authors assess 12 commercial LLMs, revealing divergent profiles among models and families, and showing that newer models increase knowledge while only partially improving defense. The framework’s reliability is supported by high cross‑judge consistency and low inter‑pipeline correlations.

By Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang
arXiv Machine Learning
Jul 21

LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats

arXiv:2607. 16227v1 Announce Type: new Abstract: LLMs are increasingly deployed in security-critical systems across healthcare, finance, education, and decision support, yet their inability to forget creates serious cybersecurity, privacy, and safety risks.

By Ruppikha Sree Shankar, Abhishek Bhardwaj, Arnav Doshi, Anusri Nagarajan, Troy Paulus Asia, Saptarshi Sengupta