Hugging Face Trending Papers

Grounded verification of chemical and materials reasoning: detection is the bottleneck

Large language models confabulate chemical objects (molecular formulas, space groups, formation energies) in fluent reasoning traces, concentrated on long-tail entities where confidence is least trustworthy. Deterministic, database-grounded verification can catch and repair such errors without the coverage cost of blanket retrieval; the binding constraint, we find, is detection, not repair.

arXiv Computation and Language
3d ago

CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search

CORE is a search controller that uses a verifier to obtain a certified conflict core, backjumps to the latest decision in that core, and caches the conflict to prevent repetition. In experiments on 2,000 graph‑coloring instances, CORE cuts median verifier calls by up to 39.8% compared to chronological repair, and improves success rates on five reasoning tasks, achieving 75.9% with Qwen2.5‑7B‑Instruct and 84.2% with Qwen3‑8B versus 72.5% and 81.8% for Tree of Thoughts. The approach also reduces verifier calls and generated tokens on both language‑model backbones.

By Siyu Song, Rui Xu, Jia Lin, Kai Liu, Weifang Wang
arXiv AI
Jul 7

Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

arXiv:2607. 05199v1 Announce Type: new Abstract: Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows.

By Raj Jaiswal, Dhruv Jain, Rishabh Dhawan, Sree Krishna Uppalapati, Shin'ichi Satoh, Tanuja Ganu, Rajiv Ratn Shah
arXiv Computation and Language
Aug 28

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.

By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher