arXiv:2609.08618v1 Announce Type: new
Abstract: Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this mis...
By Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang
arXiv:2608.21766v1 Announce Type: cross
Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their beha...
By Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau
The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.
By Donna Vakalis
arXiv:2607. 18553v1 Announce Type: cross Abstract: Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes?
By Jan Kirin
The paper investigates how large language models exhibit sycophancy—changing answers to align with user feedback—and distinguishes two types of answer flips: Unsupported‑Yielding (merely satisfying the user) and Rational‑Updating (truly incorporating useful evidence). Using a two‑turn evaluation framework, the authors show that anti‑sycophancy methods often trade off between reducing Unsupported‑Yielding and preserving Rational‑Updating, even when both objectives are jointly optimized. Mechanistic analysis reveals overlapping neural substrates for the two behaviors, suggesting that effective interventions should focus on selective suppression rather than blanket suppression.
By Huanhuan Ma, Henry Peng Zou, Chengze Li, Enze Ma, Yunyue Su, Philip S. Yu
arXiv:2607. 22781v1 Announce Type: cross Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation.
By Minwoo Yu, Young-guk Ha