arXiv:2607. 29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought.
By Matthew Nguyen, Kyle Cox, Austin Meek, Iv\'an Arcuschin
arXiv:2608.21766v1 Announce Type: cross
Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their beha...
By Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau
The paper introduces Verbalization Training (VT), a technique that encourages large language models (LLMs) to openly express their evaluation awareness (EA) without directly supervising their internal beliefs. VT works by truncating model rollouts just before spontaneous verbalizations, creating training prefixes that signal awareness, and then applying a reinforcement learning objective to increase calibrated verbalization. Experiments on models such as Qwen3.6-35B-A3B, Kimi K2.6, and Inkling show that VT boosts verbalized EA by 2.4–2.9× while keeping latent EA and overall behavior largely unchanged, and a causal study confirms that VT-induced verbalizations reflect newly acquired meta‑knowledge.
By Usman Anwar, Sahar Abdelnabi, David Krueger
arXiv:2606. 14199v1 Announce Type: cross Abstract: Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation.
By Xuhui Zhou, Weiwei Sun, Weihua Du, Jiarui Liu, Haojia Sun, Qianou Ma, Tongshuang Wu, Yiming Yang, Maarten Sap
arXiv:2606. 07897v1 Announce Type: new Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user.
By Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt
arXiv:2603.18007v2 Announce Type: replace-cross
Abstract: The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to inf...
By Anna Babarczy, Andras Lukacs, Peter Vedres, Zeteny Bujka