arXiv AI

When Chain-of-Thought Fails, the Solution Hides in the Hidden States

arXiv Computation and Language
Aug 25

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

The paper introduces a sequential activation patching framework to study how Chain-of-Thought (CoT) prompting influences large language models over multiple generated tokens. By tracking CoT-conditioned attention-head activations across token positions and aggregating them with Part-of-Speech guidance, the authors identify distributed head sets that jointly contribute to answer generation. Targeted zero-ablation experiments confirm that these heads are functionally important, affecting mechanisms such as reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation.

By Murat Dura, Serkan \"Ozt\"urk, Selma Tekir
arXiv Computation and Language
3d ago

Detecting Hidden Chain-of-Thought in Large Language Models with Linguistic, Behavioral, and Mechanistic Indicators

arXiv:2608.29956v1 Announce Type: new Abstract: Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete p...

By Armaan Singh, Ryan Trinh Le, Jasmine Kaur, Abdullah Sultan, Edward Lue Chee Lip, Kiran Nijjer, Adnan Ahmed, Vasu Sharma