arXiv AI By Noah Y. Siegel, Nicolas Heess, Maria Perez-Ortiz, Oana-Maria Camburu

Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

Read the original on arXiv AI →

arXiv:2503. 13445v3 Announce Type: replace-cross Abstract: When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 16

An Empirical Study of Counterfactual Self-Explanations in LLMs

The paper investigates counterfactual self‑explanations in large language models, where a model edits an input minimally to change its own prediction. Experiments on sentiment analysis and natural language inference with ten instruction‑tuned models from the LLaMA‑3 and Qwen‑2.5 families show that larger models produce more faithful, minimal, and human‑aligned counterfactuals. While rationale‑guided prompts improve minimality and alignment, they do not consistently enhance faithfulness, indicating that explanation quality depends heavily on model capacity and requires empirical validation.

By Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis, Vassilis Lyberatos, Giorgos Stamou
arXiv AI
Aug 26

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

The paper investigates how quantization affects large language models’ self‑explanations, examining natural language explanations and counterfactual examples across three quantization techniques and bit widths. Results show moderate declines in explanation quality (up to 4.4%) and faithfulness (up to 3.9%), with user studies indicating up to an 8.5% drop in coherence and trustworthiness. Larger models are less resilient in quality but remain more faithful, and no single quantization method consistently outperforms others across accuracy, quality, and faithfulness.

By Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann, Sebastian M\"oller, Vera Schmitt