arXiv AI By Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

Read the original on arXiv AI →

arXiv:2607. 12792v1 Announce Type: cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings

The paper investigates whether response safety can be measured by the cosine similarity between a response embedding and the mean embedding of known‑safe responses. Using four frozen encoders and prompt‑controlled datasets, the authors find that a simple prototype (mean safe embedding) performs poorly (ROC‑AUC 0.457‑0.545) while an explicit safe‑minus‑unsafe reference achieves higher scores (0.588‑0.738). The study shows that a class mean is merely a location, not a safety direction, and that a reference with sufficient unsafe mass is needed to orient safety judgments.

By Sahil Kadadekar
arXiv Machine Learning
Jun 25

RAS: Measuring LLM Safety Through Refusal Alignment

arXiv:2606. 25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy.

By Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu, Wei-Bin Lee