arXiv:2608.21664v1 Announce Type: new
Abstract: Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especiall...
By Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa
arXiv:2607. 23379v1 Announce Type: cross Abstract: Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations.
By Tobias Bersia, Tatiana Gaintseva
arXiv:2606. 10229v1 Announce Type: cross Abstract: We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that trains on the curated data.
By Aarav Bedi
arXiv:2607. 20455v1 Announce Type: cross Abstract: Human-annotated data remains fundamental to training frontier Large Language Models (LLMs).
By Siddarth Malreddy, Ishan Nigam, Akshay Arora, Nikhil Mittal, Subrat Sahu
EvalDetectBench is an open pipeline and benchmark designed to measure evaluation awareness in frontier large language models, enabling practitioners to test models against any Inspect-compatible evaluation. It includes a curated transcript suite from current frontier system-card evaluations and diverse deployment sources, and it assesses both how reliably models recognize they are being evaluated and how detectable individual benchmarks are. The benchmark addresses systematic bias by calibrating probes per model and harmonizing generator selection to correct for variance caused by model identity and prompt choice.
By Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
arXiv:2609.01604v1 Announce Type: cross
Abstract: LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the int...
By Himil Vasava, Ming Jiang
MURANO is an open‑source framework that enables researchers to design, run, and reproduce mechanistic interpretability experiments on large language models. It unifies the five key stages—loading, recording, attribution, intervention, and evaluation—into composable pipeline steps that exchange named artifacts and use canonical addresses for interoperability. The authors demonstrate the framework with reproductions of existing studies and a sparse autoencoder case study, showing its practical applicability across disciplines.
By Alireza Bayat Makou, Emirhan B\"oge, Phu Gia Hoang, Federico Tiblias, Jingcheng Niu, Subhabrata Dutta, Richard Eckart de Castilho, Iryna Gurevych
arXiv:2506. 07406v3 Announce Type: replace-cross Abstract: Understanding the internal representations of large language models (LLMs) is a central challenge in interpretability research.
By Yifan Luo, Zhennan Zhou, Bin Dong
arXiv:2607. 17524v1 Announce Type: cross Abstract: We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task.
By Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu, Robin Jia
arXiv:2606. 11172v1 Announce Type: new Abstract: Deployed large reasoning models (LRMs) often behave unexpectedly.
By Evgenii Kortukov, Piotr Komorowski, Florian Klein, Paula Engl, Gabriele Sarti, Seong Joon Oh, Sebastian Lapuschkin, Wojciech Samek
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
By Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu
arXiv:2605.24614v2 Announce Type: replace-cross
Abstract: Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether...
By Jaeung Lee, Dohyun Kim, Jaemin Jo