arXiv Machine Learning

Grounded verification of chemical and materials reasoning: detection is the bottleneck

arXiv:2607. 17417v1 Announce Type: new Abstract: Large language models confabulate chemical objects (molecular formulas, space groups, formation energies) in fluent reasoning traces, concentrated on long-tail entities where confidence is least trustworthy.

Hugging Face Trending Papers
Jul 19

Grounded verification of chemical and materials reasoning: detection is the bottleneck

Large language models confabulate chemical objects (molecular formulas, space groups, formation energies) in fluent reasoning traces, concentrated on long-tail entities where confidence is least trustworthy. Deterministic, database-grounded verification can catch and repair such errors without the coverage cost of blanket retrieval; the binding constraint, we find, is detection, not repair.

arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut
arXiv Computation and Language
3d ago

CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search

CORE is a search controller that uses a verifier to obtain a certified conflict core, backjumps to the latest decision in that core, and caches the conflict to prevent repetition. In experiments on 2,000 graph‑coloring instances, CORE cuts median verifier calls by up to 39.8% compared to chronological repair, and improves success rates on five reasoning tasks, achieving 75.9% with Qwen2.5‑7B‑Instruct and 84.2% with Qwen3‑8B versus 72.5% and 81.8% for Tree of Thoughts. The approach also reduces verifier calls and generated tokens on both language‑model backbones.

By Siyu Song, Rui Xu, Jia Lin, Kai Liu, Weifang Wang
arXiv Machine Learning
Sep 11

Legible Failures: Detecting and Repairing In-Context Binding Errors

The paper investigates in-context binding errors in language models, showing that a linear probe can recover correct entity bindings from frozen hidden states even when the model outputs incorrect bindings. Across 16 checkpoints, the probe’s accuracy on failure cases surpasses a baseline by about 0.196, and a probe‑based score improves failure detection over the model’s confidence by 0.079 AUROC. Steering the residual stream toward the probe‑decoded binding further boosts accuracy by an average of 0.168 across eight models.

By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari