arXiv AI By Houman Mehrafarin, Amit Parekh, Ioannis Konstas

When Chain-of-Thought Fails, the Solution Hides in the Hidden States

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Aug 25

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

The paper introduces a sequential activation patching framework to study how Chain-of-Thought (CoT) prompting influences large language models over multiple generated tokens. By tracking CoT-conditioned attention-head activations across token positions and aggregating them with Part-of-Speech guidance, the authors identify distributed head sets that jointly contribute to answer generation. Targeted zero-ablation experiments confirm that these heads are functionally important, affecting mechanisms such as reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation.

By Murat Dura, Serkan \"Ozt\"urk, Selma Tekir