arXiv AI

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators

arXiv:2607. 13075v1 Announce Type: cross Abstract: Context can change whether a request is harmful without changing its topic or surface form.

arXiv Machine Learning
Jul 31

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

arXiv:2607. 09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation).

By Kwan Soo Shin