arXiv AI By Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

Read the original on arXiv AI →

arXiv:2608. 07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 11

Legible Failures: Detecting and Repairing In-Context Binding Errors

The paper investigates in-context binding errors in language models, showing that a linear probe can recover correct entity bindings from frozen hidden states even when the model outputs incorrect bindings. Across 16 checkpoints, the probe’s accuracy on failure cases surpasses a baseline by about 0.196, and a probe‑based score improves failure detection over the model’s confidence by 0.079 AUROC. Steering the residual stream toward the probe‑decoded binding further boosts accuracy by an average of 0.168 across eight models.

By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari
arXiv AI
Aug 28

Invocation-Level Reliability of Tool-Using Agents

The paper investigates the reliability of tool‑using agents, focusing on two failure modes: selecting the wrong tool and constructing incorrect arguments. It introduces a correct‑invocation rate metric to distinguish these errors and evaluates five open‑weight models on multi‑step tasks up to depth 8, finding that by depth 6 about 70% of a model’s clean‑context capability is lost due to earlier mistakes. The study reveals that exact‑match scoring against a fixed gold trajectory forces severity and recovery parameters to extreme values, and proposes a conditional‑on‑state scoring remedy that yields more realistic severity estimates.

By Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta