arXiv Machine Learning By Kia-J\"ung Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

Read the original on arXiv Machine Learning →

arXiv:2605. 26772v1 Announce Type: cross Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.