arXiv AI By Agatha Duzan, Asa Cooper Stickland

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Read the original on arXiv AI →

arXiv:2608. 04735v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models

The paper introduces a causal, layerwise audit method called the CoT Mediation Index (CMI) to evaluate whether chain-of-thought (CoT) prompting truly influences a language model’s internal computation. By comparing performance degradation from patching CoT-token hidden states against matched control patches, the authors find that CoT influence is often confined to narrow reasoning windows and can be nearly absent even when the model produces fluent rationales. The study shows that models explicitly tuned for reasoning exhibit stronger mediation, while Mixture-of-Experts models display more distributed mediation, indicating that CoT faithfulness varies across models and tasks.

By Anish Sathyanarayanan, Aditya Nagarsekar, Aarush Rathore
arXiv Machine Learning
Aug 14

A Probe Direction Is a Property of Its Prompt

arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.

By Valentin No\"el