arXiv AI By Jiawen Tao, Xiaokun Yuan, Yaoming Li, Chenxu Liu, Mengzhou Wu, Tong Yang, Maxm Pan

REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

Read the original on arXiv AI →

REALHOP introduces a behavioral auditing framework to assess multi‑hop reasoning by measuring the Behavioral Necessity Rate (BNR), which quantifies how often removing targeted evidence prevents correct answers. Across five benchmarks, the framework reveals a wide gap between annotated reasoning chains and actual evidence dependence, with panel‑mean BNR ranging from 16.6% to 48.9%. By re‑binding entities, factorizing relations, adding competing paths, and placing evidence at traceable locations, REALHOP raises BNR dramatically—from 27.4% to 94.4% on MuSiQue questions—while maintaining high overall accuracy and improving performance on long‑context tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

Are Stated Reasoning Steps Causally Load-Bearing?

The study investigates whether the reasoning steps a language model writes are causally responsible for its answers. Using a causal intervention method on the activation stream, the authors find that for Qwen3-4B, about 77% of stated steps are causally load‑bearing, while behavioral tests overestimate this by roughly 11 percentage points. The faithfulness of reasoning decreases with model size and depth of reasoning, especially for the smaller Qwen3-1.7B.

By Abhiram Bhupatiraju, Rayan Nyaupane