Hugging Face Trending Papers

Knowledge Knows, Verbalization Tells: Disentangling Latent Directions for Mathematical Solvability in LLMs

Although LLMs have made significant progress in mathematical reasoning, determining whether a mathematical problem is solvable remains a fundamental yet challenging capability. While recent studies have probed internal representations of model solvability beliefs, verbalization has primarily been studied behaviorally rather than as an internal representation, limiting its analysis and manipulation.

arXiv AI
Aug 20

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

The paper investigates how LLaMA 3.1‑8B models numerical sequence patterns, focusing on time‑series prediction. By designing a task that requires detecting structural cues—specifically first differences in a sequence—the authors show that the model performs well and internally computes and stores these differences. Probing and activation‑patching experiments reveal that LLaMA retrieves and applies the first‑difference via an induction‑like circuit, marking one of the first demonstrations of concept induction in large language models.

By Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang
arXiv AI
Jul 22

Fluid Reasoning Representations

arXiv:2602. 04843v2 Announce Type: replace Abstract: Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking.

By Dmitrii Kharlapenko, Terry Jingchen Zhang, Arth Singh, Alessandro Stolfo, Arthur Conmy, Mrinmaya Sachan, Zhijing Jin
arXiv AI
Jul 10

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

arXiv:2607. 08046v1 Announce Type: cross Abstract: Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast.

By Rapha\"el Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho
arXiv AI
Aug 24

Can LLMs Introspect? A Reality Check

The paper questions whether large language models (LLMs) truly introspect by critiquing recent studies that claim they can detect and report their internal states. It proposes two necessary conditions for genuine introspection: privileged access to internal representations and second‑order computation that distinguishes from first‑order task performance. Re‑examining two existing paradigms, the authors find that apparent introspective abilities can be explained by input‑based classifiers or generic anomaly detection, concluding that current evidence does not support metacognitive monitoring in LLMs.

By Shashwat Singh, Tal Linzen, Shauli Ravfogel
arXiv AI
Sep 21

When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap

The paper investigates the effectiveness of activation steering in latent chain-of-thought (CoT) reasoning compared to explicit CoT. It finds that steering continuous latent thoughts yields weaker impacts on language generation, even when hidden representations are shifted similarly. The authors propose a latent-to-language transition gap, supported by evidence of abrupt output distribution changes at the transition boundary and weaker bidirectional control in latent CoT.

By Gaoxiang Huang, Lei Qi