arXiv AI By Jingkun Luo, Da-Tian Peng

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

Read the original on arXiv AI →

arXiv:2607. 24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 31

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

arXiv:2607. 09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation).

By Kwan Soo Shin