arXiv Machine Learning By Kwan Soo Shin

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

Read the original on arXiv Machine Learning →

arXiv:2607. 09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

A Probe Direction Is a Property of Its Prompt

arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.

By Valentin No\"el