arXiv AI

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

arXiv:2608. 16747v1 Announce Type: cross Abstract: Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors.

arXiv AI
3d ago

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning

The paper introduces FAME, a training‑free framework that evaluates false memory in autonomous agents by tracking how their internal beliefs shift under counterfactual scenarios. False memory, defined as biases arising from spurious correlations, environment shifts, or knowledge conflicts, is hard to detect with standard methods. FAME measures concept drift in hidden states, achieving AUROCs between 76.2% and 96.7% and outperforming baselines on benchmarks such as GSM‑Symbolic, GitChameleon, and BigBench‑Hard.

By Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu
arXiv Computation and Language
Sep 16

An Empirical Study of Counterfactual Self-Explanations in LLMs

The paper investigates counterfactual self‑explanations in large language models, where a model edits an input minimally to change its own prediction. Experiments on sentiment analysis and natural language inference with ten instruction‑tuned models from the LLaMA‑3 and Qwen‑2.5 families show that larger models produce more faithful, minimal, and human‑aligned counterfactuals. While rationale‑guided prompts improve minimality and alignment, they do not consistently enhance faithfulness, indicating that explanation quality depends heavily on model capacity and requires empirical validation.

By Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis, Vassilis Lyberatos, Giorgos Stamou
arXiv Machine Learning
Aug 27

iFlip: Iterative Feedback-driven Counterfactual Example Refinement

iFlip is an iterative refinement method for generating counterfactual examples using large language models. It incorporates three feedback types—model confidence, feature attribution, and natural language—to guide successive edits. Experiments show iFlip outperforms five state‑of‑the‑art baselines, achieving a 57.8% higher validity rate and improving model performance through counterfactual data augmentation.

By Yilong Wang, Qianli Wang, Nils Feldhus
arXiv Machine Learning
Jul 9

Counterfactual Modeling with Fine-Tuned LLMs for Health Intervention Design and Sensor Data Augmentation

arXiv:2601. 14590v3 Announce Type: replace Abstract: Counterfactual explanations (CFEs) provide human-centric interpretability by identifying the minimal, actionable changes required to alter a machine learning model's prediction.

By Shovito Barua Soumma, Asiful Arefeen, Stephanie M. Carpenter, Melanie Hingle, Hassan Ghasemzadeh
arXiv AI
Jun 2

From Features to Actions: Explainability in Traditional and Agentic AI Systems

arXiv:2602. 06841v4 Announce Type: replace Abstract: Over the last decade, Explainable AI has primarily focused on interpreting individual model predictions, producing post-hoc explanations that relate inputs to outputs under a fixed decision structure.

By Sindhuja Chaduvula, Jessee Ho, Kina Kim, Aravind Narayanan, Ahmed Y. Radwan, Mahshid Alinoori, Muskan Garg, Dhanesh Ramachandram, Shaina Raza
arXiv AI
Jul 3

From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents

arXiv:2604. 19775v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments.

By Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, Adam D. Cobb, Daniel Elenius, Manoj Acharya, Colin Samplawski, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha, Ugur Kursuncu, Anirban Roy