Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Read the original on arXiv Machine Learning →arXiv:2607. 09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation).
Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.