arXiv AI

The Impossibility of Eliciting Latent Knowledge

arXiv:2606. 12268v1 Announce Type: new Abstract: Advanced AI systems have extensive knowledge of their environments; in fact, their knowledge may (far) exceed that of their developers or users.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv AI
Aug 24

Why2Speak: Faithful Reasoning for Abstaining Action Policies

The paper investigates how agentic systems decide between acting and abstaining, focusing on the fidelity of their reasoning explanations. Using Qwen3‑8B in a multi‑party conversation setting, the authors compare direct decision policies, reasoning policies, supervised fine‑tuning, and reinforcement learning, finding a trade‑off: strong direct policies yield higher performance but no traceable reasoning, while reasoning policies provide an audit trail at the cost of lower recall. The study also uncovers that exposing reasoning can alter the agent’s policy and that common faithfulness metrics may overstate the alignment between reasoning and decisions.

By Shreya Mendi, Brinnae Bent
arXiv Computation and Language
Sep 22

XYEval: Agents say yes to bad advice

arXiv:2609.23939v1 Announce Type: new Abstract: Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where...

By Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord, Been Kim
arXiv Machine Learning
Sep 23

xWhyL: Causal Interactive Learning

arXiv:2609.26037v1 Announce Type: new Abstract: Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for lear...

By Nicholas Tagliapietra, Florian Peter Busch, Moritz Willig, Matej Ze\v{c}evi\'c, Lavdim Halilaj, Juergen Luettin, Kristian Kersting
arXiv AI
3d ago

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning

The paper introduces FAME, a training‑free framework that evaluates false memory in autonomous agents by tracking how their internal beliefs shift under counterfactual scenarios. False memory, defined as biases arising from spurious correlations, environment shifts, or knowledge conflicts, is hard to detect with standard methods. FAME measures concept drift in hidden states, achieving AUROCs between 76.2% and 96.7% and outperforming baselines on benchmarks such as GSM‑Symbolic, GitChameleon, and BigBench‑Hard.

By Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu
arXiv Computation and Language
Sep 7

From Plausible to Actionable: A Position on LLM Self-Explanations

The paper discusses how Large Language Models can produce natural language self‑explanations that appear plausible but may not accurately reflect the model’s reasoning. It critiques current evaluation methods for such explanations and offers practical guidelines to assess their plausibility and faithfulness. Additionally, it argues that evaluation should also consider the actionability of these explanations, showing how they can aid decision‑making for various stakeholders.

By Elize Herrewijnen, Benedetta Muscato, Gizem Gezici, Fosca Giannotti
arXiv AI
Sep 16

Verifiable Social Reasoning for LLM Assistants

arXiv:2609.17496v1 Announce Type: new Abstract: LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (...

By Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish, Ariel Goldstein, Marian Croak, Avinatan Hassidim, Yossi Matias, Amir Feder
Hugging Face Trending Papers
Aug 12

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior.