arXiv Computation and Language By Phuong Q. Le, Kemal Kurniawan, Jey Han Lau

Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength

Read the original on arXiv Computation and Language →

The paper introduces an LLM-as-a-judge method to quantify perturbation strength for assessing self-consistency in LLM-generated explanations. It evaluates explanations from various LLMs under controlled perturbation strengths, comparing input and chain-of-thought (CoT) perturbations. Results show the proposed measure outperforms embedding- and probability-based approaches, and that input perturbations impact LLMs more than CoT perturbations, implying fair self-consistency judgments require consistent perturbation types.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 16

An Empirical Study of Counterfactual Self-Explanations in LLMs

The paper investigates counterfactual self‑explanations in large language models, where a model edits an input minimally to change its own prediction. Experiments on sentiment analysis and natural language inference with ten instruction‑tuned models from the LLaMA‑3 and Qwen‑2.5 families show that larger models produce more faithful, minimal, and human‑aligned counterfactuals. While rationale‑guided prompts improve minimality and alignment, they do not consistently enhance faithfulness, indicating that explanation quality depends heavily on model capacity and requires empirical validation.

By Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis, Vassilis Lyberatos, Giorgos Stamou
arXiv AI
Aug 7

Position: It's Time to Optimize LLMs for Self-Consistency

arXiv:2608. 05188v1 Announce Type: cross Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses.

By Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas
arXiv Computation and Language
Sep 7

From Plausible to Actionable: A Position on LLM Self-Explanations

The paper discusses how Large Language Models can produce natural language self‑explanations that appear plausible but may not accurately reflect the model’s reasoning. It critiques current evaluation methods for such explanations and offers practical guidelines to assess their plausibility and faithfulness. Additionally, it argues that evaluation should also consider the actionability of these explanations, showing how they can aid decision‑making for various stakeholders.

By Elize Herrewijnen, Benedetta Muscato, Gizem Gezici, Fosca Giannotti