arXiv AI By Avni Mittal, Rauno Arike

C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning

Read the original on arXiv AI →

arXiv:2603. 05167v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as judges of chain-of-thought (CoT) reasoning, yet it remains unclear whether they can reliably assess process faithfulness rather than merely answer plausibility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

The paper investigates whether the text of chain‑of‑thought reasoning steps actually reflects their true importance for a model’s final answer. By defining step importance as the advantage in expected reward when a step is included, the authors use Monte Carlo rollouts to estimate ground truth and then test whether large language model judges can identify high‑advantage steps. They find that capable LLMs can beat a prevalence baseline but still fall far short of a noise ceiling, and that fine‑tuning a step‑level critic improves detection for incorrect responses but remains distant from the ceiling for correct ones, indicating that step importance is only partially recoverable from the reasoning trace text.

By Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli
arXiv AI
Sep 3

ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs

The paper introduces ICE (Intervention-Consistent Explanation), a framework that evaluates the faithfulness of large language model explanations by comparing them to random baselines of equal size across multiple intervention operators. It demonstrates that faithfulness varies with the chosen operator, with significant differences observed when switching between deletion and retrieval infill operators. The study evaluates seven LLMs on four tasks, revealing that operator changes can cross the positive-evidence threshold in 18% of configurations and that random baselines uncover anti-faithfulness in nearly a third of English deletion setups, findings that also hold across six non‑English languages and two attribution methods.

By Abhinaba Basu, Pavan Chakraborty