arXiv:2608. 03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring.
By Dominik Meier, Luca Joshua Francis, Marco Bernhard Kaiser, Terry Ruas, Jan Philip Wahle, Bela Gipp
arXiv:2608. 15673v1 Announce Type: cross Abstract: Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a prompt-response pair and what those facts imply under a given policy.
By Satchit Chatterji, Shihan Wang, Giovanni Sileno, Erman Acar
The paper investigates why large reasoning models (LRMs) lose safety alignment when faced with harmful queries. By analyzing token-level refusal dynamics, the authors identify a vulnerability called Onset Refusal Collapse (ORC), where the refusal signal drops sharply at the first generated token, leading to unsafe responses. They introduce SafeToken, a lightweight inference-time intervention that injects a learned safety anchor at reasoning onset, which mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility.
By Yizheng Yang, Haining Yu, Yuechen Wang, Yikai Hou, Xing Fu, Jinbo Yang, Tianqing Zhu
EOPSA (Efficient On-Policy Self-Distilled Safety Alignment) addresses inefficiencies in On-Policy Self-Distillation (OPSD) for safety alignment by focusing training on safety-critical tokens. It introduces Adaptive Rollout Scheduling, which limits generation length based on a Teacher Rescue Rate metric, and Selective Distillation, which filters out safety-neutral tokens to concentrate gradient updates on safety-pivotal transitions. Experiments on models up to 32B parameters show that EOPSA reduces rollout computation by about 50% and backpropagates through only roughly 2% of tokens, outperforming full-token distillation baselines in safety compliance and reasoning retention.
arXiv:2610.00601v1 Announce Type: cross
Abstract: Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface fo...
By Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal
arXiv:2606. 30128v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning, but the source is contested: do the intermediate steps help because they carry useful semantic content, or because conditioning on more tokens buys extra computation before the model commits to an answer?
By Wenlong Wang, Fergal Reid