arXiv:2606. 28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs.
By Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath
The paper investigates counterfactual self‑explanations in large language models, where a model edits an input minimally to change its own prediction. Experiments on sentiment analysis and natural language inference with ten instruction‑tuned models from the LLaMA‑3 and Qwen‑2.5 families show that larger models produce more faithful, minimal, and human‑aligned counterfactuals. While rationale‑guided prompts improve minimality and alignment, they do not consistently enhance faithfulness, indicating that explanation quality depends heavily on model capacity and requires empirical validation.
By Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis, Vassilis Lyberatos, Giorgos Stamou
arXiv:2503. 13445v3 Announce Type: replace-cross Abstract: When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans.
By Noah Y. Siegel, Nicolas Heess, Maria Perez-Ortiz, Oana-Maria Camburu
arXiv:2606. 30653v1 Announce Type: cross Abstract: Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification.
By Marina Mancoridis, Zo\"e Hitzig
arXiv:2608. 05188v1 Announce Type: cross Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses.
By Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas
The paper discusses how Large Language Models can produce natural language self‑explanations that appear plausible but may not accurately reflect the model’s reasoning. It critiques current evaluation methods for such explanations and offers practical guidelines to assess their plausibility and faithfulness. Additionally, it argues that evaluation should also consider the actionability of these explanations, showing how they can aid decision‑making for various stakeholders.
By Elize Herrewijnen, Benedetta Muscato, Gizem Gezici, Fosca Giannotti
arXiv:2606. 14867v1 Announce Type: cross Abstract: Proof autoformalization aims to translate a mathematical informal proof written in natural language into a formal proof in a formal language such as Lean~4.
By Zhengtao Gui, Sheng Yang, Zhouxing Shi
arXiv:2606. 18327v1 Announce Type: cross Abstract: Language models (LMs) that faithfully describe their own behavior can more easily be audited, understood, and trusted by users.
By Itamar Pres, Laura Ruis, Melat Ghebreselassie, Belinda Z. Li, Jacob Andreas
The paper investigates how quantization affects large language models’ self‑explanations, examining natural language explanations and counterfactual examples across three quantization techniques and bit widths. Results show moderate declines in explanation quality (up to 4.4%) and faithfulness (up to 3.9%), with user studies indicating up to an 8.5% drop in coherence and trustworthiness. Larger models are less resilient in quality but remain more faithful, and no single quantization method consistently outperforms others across accuracy, quality, and faithfulness.
By Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann, Sebastian M\"oller, Vera Schmitt
arXiv:2603. 14894v3 Announce Type: replace-cross Abstract: Trust and ethical concerns due to the widespread deployment of opaque machine learning (ML) models motivating the need for reliable model explanations.
By Sumedha Chugh, Ranjitha Prasad, Nazreen Shah
arXiv:2606. 08129v1 Announce Type: new Abstract: Large language models (LLMs) differ in architecture, training data, and optimization procedures, yet they may still develop similar internal inference patterns.
By Siyu Lou, Yao Yan, Yuntian Chen, Quanshi Zhang
arXiv:2605. 07527v2 Announce Type: replace-cross Abstract: Recent work has observed that explanations produced by Self-Interpretable Graph Neural Networks (SI-GNNs) can be self-inconsistent: when the model is reapplied to its own explanatory graph subset, it may produce a different explanation.
By Wenxin Tai, Yaqian Liu, Ting Zhong, Fan Zhou