arXiv AI By Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

Read the original on arXiv AI →

arXiv:2606. 28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.