arXiv AI By Dominik Schwarz

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators

Read the original on arXiv AI →

arXiv:2607. 13075v1 Announce Type: cross Abstract: Context can change whether a request is harmful without changing its topic or surface form.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
23h ago

A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings

The paper investigates whether response safety can be measured by the cosine similarity between a response embedding and the mean embedding of known‑safe responses. Using four frozen encoders and prompt‑controlled datasets, the authors find that a simple prototype (mean safe embedding) performs poorly (ROC‑AUC 0.457‑0.545) while an explicit safe‑minus‑unsafe reference achieves higher scores (0.588‑0.738). The study shows that a class mean is merely a location, not a safety direction, and that a reference with sufficient unsafe mass is needed to orient safety judgments.

By Sahil Kadadekar