arXiv AI By Zacharie Bugaud

Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs

Read the original on arXiv AI →

arXiv:2606. 04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5 models (12B--70B) in 4,200 interactions with dual-judge validation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

The paper introduces FUSE, a modular framework that evaluates large language models (LLMs) for dangerous capabilities across three orthogonal pipelines: Knowledge (K), Defense (D), and Harm (H). Using a chemical‑biological module, the authors assess 12 commercial LLMs, revealing divergent profiles among models and families, and showing that newer models increase knowledge while only partially improving defense. The framework’s reliability is supported by high cross‑judge consistency and low inter‑pipeline correlations.

By Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang