arXiv AI By Abhinaba Basu, Pavan Chakraborty

ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs

Read the original on arXiv AI →

The paper introduces ICE (Intervention-Consistent Explanation), a framework that evaluates the faithfulness of large language model explanations by comparing them to random baselines of equal size across multiple intervention operators. It demonstrates that faithfulness varies with the chosen operator, with significant differences observed when switching between deletion and retrieval infill operators. The study evaluates seven LLMs on four tasks, revealing that operator changes can cross the positive-evidence threshold in 18% of configurations and that random baselines uncover anti-faithfulness in nearly a third of English deletion setups, findings that also hold across six non‑English languages and two attribution methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen