arXiv:2606. 18327v1 Announce Type: cross Abstract: Language models (LMs) that faithfully describe their own behavior can more easily be audited, understood, and trusted by users.
By Itamar Pres, Laura Ruis, Melat Ghebreselassie, Belinda Z. Li, Jacob Andreas
arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.
By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.
The paper investigates how agentic systems decide between acting and abstaining, focusing on the fidelity of their reasoning explanations. Using Qwen3‑8B in a multi‑party conversation setting, the authors compare direct decision policies, reasoning policies, supervised fine‑tuning, and reinforcement learning, finding a trade‑off: strong direct policies yield higher performance but no traceable reasoning, while reasoning policies provide an audit trail at the cost of lower recall. The study also uncovers that exposing reasoning can alter the agent’s policy and that common faithfulness metrics may overstate the alignment between reasoning and decisions.
By Shreya Mendi, Brinnae Bent
arXiv:2609.23939v1 Announce Type: new
Abstract: Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where...
By Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord, Been Kim
arXiv:2609.26037v1 Announce Type: new
Abstract: Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for lear...
By Nicholas Tagliapietra, Florian Peter Busch, Moritz Willig, Matej Ze\v{c}evi\'c, Lavdim Halilaj, Juergen Luettin, Kristian Kersting