arXiv AI

C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning

arXiv:2603. 05167v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as judges of chain-of-thought (CoT) reasoning, yet it remains unclear whether they can reliably assess process faithfulness rather than merely answer plausibility.

arXiv Machine Learning
Sep 4

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

The paper investigates whether the text of chain‑of‑thought reasoning steps actually reflects their true importance for a model’s final answer. By defining step importance as the advantage in expected reward when a step is included, the authors use Monte Carlo rollouts to estimate ground truth and then test whether large language model judges can identify high‑advantage steps. They find that capable LLMs can beat a prevalence baseline but still fall far short of a noise ceiling, and that fine‑tuning a step‑level critic improves detection for incorrect responses but remains distant from the ceiling for correct ones, indicating that step importance is only partially recoverable from the reasoning trace text.

By Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli
arXiv AI
Sep 3

ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs

The paper introduces ICE (Intervention-Consistent Explanation), a framework that evaluates the faithfulness of large language model explanations by comparing them to random baselines of equal size across multiple intervention operators. It demonstrates that faithfulness varies with the chosen operator, with significant differences observed when switching between deletion and retrieval infill operators. The study evaluates seven LLMs on four tasks, revealing that operator changes can cross the positive-evidence threshold in 18% of configurations and that random baselines uncover anti-faithfulness in nearly a third of English deletion setups, findings that also hold across six non‑English languages and two attribution methods.

By Abhinaba Basu, Pavan Chakraborty
arXiv Computation and Language
Aug 25

GRACE: Step-Level Benchmark for Faithful Reasoning over Context

GRACE is a step‑level benchmark for evaluating the faithfulness of chain‑of‑thought reasoning over context. It provides human annotations for each step in CoT traces from 10 models across 4 datasets, labeling faithfulness, error category, and natural‑language explanations. The benchmark introduces a data‑driven taxonomy that splits errors into GRACE‑Inference (deductive) and GRACE‑Grounding (factual) tracks, each with four categories, and demonstrates that incorporating step‑level faithfulness signals can improve downstream accuracy and reasoning reliability.

By Hoang Pham, Dong Le, Anh Tuan Luu
arXiv AI
3d ago

LSR-Ben: A Logical and Scientific Reasoning Benchmark for Evaluating Process Reward Models

The paper introduces LSR‑Ben, a benchmark designed to evaluate process reward models (PRMs) on scientific and logical reasoning tasks, addressing a gap left by existing math‑focused benchmarks. Experiments on 22 models reveal that PRMs and LLMs perform poorly in non‑mathematical domains, with LLMs tending to over‑identify errors while PRMs tend to overlook them. LSR‑Ben aims to spur research that broadens PRM applicability and improves LLM reasoning.

By Zhouhao Sun, Xuan Zhang, Xiao Ding, Bibo Cai, Li Du, Kai Xiong, Xinran Dai, Fei Zhang, weidi tang, Zhiyuan Kan, Yang Zhao, Bing Qin, Ting Liu
arXiv AI
Aug 14

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

arXiv:2608. 12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation.

By Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar
arXiv Computation and Language
Aug 27

Adaptive Triggering for Bias Correction in LLM Reasoning

The paper introduces an adaptive triggering mechanism for bias correction in large language model (LLM) reasoning. By framing bias intervention as an online change‑point detection problem, the authors update a CUSUM statistic at each step using either a white‑box next‑token probability signal or a black‑box LLM judge signal, and inject corrective prompts only when the accumulated evidence exceeds a calibrated threshold. Experiments on gpt‑4o‑mini and six open‑weight models show that adaptive black‑box triggering restores most of the accuracy lost by fixed‑interval interventions while reducing the number of corrections, whereas the white‑box signal improves ambiguous‑item accuracy but can hurt disambiguated‑item accuracy due to difficulty distinguishing stereotype reliance from correct evidence.

By Nayoung Kim, Mickey Mancenido, Huan Liu
arXiv Computation and Language
Aug 27

ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability

ReFIne is a training framework that augments large reasoning models with three trustworthiness properties: interpretability, faithfulness, and reliability. It combines supervised fine‑tuning with GRPO to produce structured, tag‑based reasoning traces, explicitly disclose decisive information, and provide self‑assessments of soundness and confidence. Applied to Qwen3 models, ReFIne improves interpretability by 44.0 %, faithfulness by 18.8 %, and reliability by 42.4 % on mathematical benchmarks.

By Chung-En Sun, Ge Yan, Akshay Kulkarni, Tsui-Wei Weng