arXiv AI By Kwan Soo Shin, In Seok Kang, Yunkyung Min, Munho Lee

A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss

Read the original on arXiv AI →

arXiv:2607. 09306v2 Announce Type: replace-cross Abstract: Whether a language model behaves as it claims is a judgement on which independent human raters cannot agree (Fleiss kappa = 0.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 1

Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

arXiv:2606. 32008v1 Announce Type: new Abstract: Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens.

By Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo, Chia-Tse Shao, Yingxiao Ye, Aobo Yang, Vivek Miglani, Nehal Bandi