arXiv AI

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

arXiv:2608. 07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction.

arXiv Machine Learning
Sep 11

Legible Failures: Detecting and Repairing In-Context Binding Errors

The paper investigates in-context binding errors in language models, showing that a linear probe can recover correct entity bindings from frozen hidden states even when the model outputs incorrect bindings. Across 16 checkpoints, the probe’s accuracy on failure cases surpasses a baseline by about 0.196, and a probe‑based score improves failure detection over the model’s confidence by 0.079 AUROC. Steering the residual stream toward the probe‑decoded binding further boosts accuracy by an average of 0.168 across eight models.

By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari
arXiv AI
Aug 28

Invocation-Level Reliability of Tool-Using Agents

The paper investigates the reliability of tool‑using agents, focusing on two failure modes: selecting the wrong tool and constructing incorrect arguments. It introduces a correct‑invocation rate metric to distinguish these errors and evaluates five open‑weight models on multi‑step tasks up to depth 8, finding that by depth 6 about 70% of a model’s clean‑context capability is lost due to earlier mistakes. The study reveals that exact‑match scoring against a fixed gold trajectory forces severity and recovery parameters to extreme values, and proposes a conditional‑on‑state scoring remedy that yields more realistic severity estimates.

By Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta
arXiv AI
Aug 12

UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

arXiv:2607. 17188v2 Announce Type: replace Abstract: While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and underthinking, which we formulate as reasoning state--action mismatch.

By Cheng Yan, Zhijun Fan, Guangyang Ye, Fan Xu, Xiang Xia, Yawei Wang, Wuyang Zhang
arXiv AI
Sep 2

Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

The study investigates whether hints that convert failing code generation attempts into passing ones provide new information or simply guide models toward solutions they could already generate. Using Qwen2.5-3B-Instruct and Phi-3.5-mini on HumanEval+ and MBPP+, the authors find that relevant hints rescue a significant portion of failures, yet many of those solutions are also recoverable through ordinary sampling. Mechanistic tests reveal a shared activation direction between relevant and unrelated hints, but adding this direction does not improve overall accuracy, indicating limited task-general transfer.

By Will Badr
arXiv Machine Learning
Sep 21

Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

arXiv:2609.20973v1 Announce Type: cross Abstract: Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-ste...

By Jiazhang Cai, Tao Wang, Ruidong Zhang, Siyuan Li, Terry Ma, Luyang Fang, Haoran Lu, Huimin Cheng, Yingchuan Zhang, Shushan Wu, Rui Xie, Lin Tang, Chao Huang, Rongjie Liu, Ziyu Liu, Meizhi Yu, Yongkai Chen, Yifan Zhou, Zeliang Sun, Chang Liu, Zhen Xiang, Wei Xiao, Zixin Rao, Xinyi Liu, Yutong Hu, Mengrui Zhang, Jing Zhang, Weidi Luo, Jincheng Yu, Zhengliang Liu, Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen, Xinliang Li, Tianming Liu, Wenxuan Zhong, Ping Ma