arXiv Machine Learning By Zehao Liu, Vasant G. Honavar

Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models

Read the original on arXiv Machine Learning →

The paper investigates ‘thinking leakage’ in post‑training hybrid reasoning models that operate in NoThink mode. Using a causal mediation framework and bidirectional interventions, the authors show that the performance gains attributed to NoThink training largely stem from the model drifting toward its Think mode. Across three models and three training methods on competition math benchmarks, leakage ratios between 42% and 79% were observed, indicating that much of the accuracy improvement is due to re‑invoking existing Think behavior rather than genuine NoThink capability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models

The paper introduces a causal, layerwise audit method called the CoT Mediation Index (CMI) to evaluate whether chain-of-thought (CoT) prompting truly influences a language model’s internal computation. By comparing performance degradation from patching CoT-token hidden states against matched control patches, the authors find that CoT influence is often confined to narrow reasoning windows and can be nearly absent even when the model produces fluent rationales. The study shows that models explicitly tuned for reasoning exhibit stronger mediation, while Mixture-of-Experts models display more distributed mediation, indicating that CoT faithfulness varies across models and tasks.

By Anish Sathyanarayanan, Aditya Nagarsekar, Aarush Rathore