arXiv Machine Learning By Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, Eugene Bagdasarian

OverThink: Slowdown Attacks on Reasoning LLMs

Read the original on arXiv Machine Learning →

The paper introduces OverThink, a slowdown attack that forces reasoning language models (RLMs) to produce many more reasoning tokens while still giving correct answers. By injecting decoy reasoning problems—such as Markov decision processes, language translation, or graphic comprehension—into the model’s context, attackers can dramatically increase token generation (up to 46× on SQuAD and 17× on coding agents). The study evaluates the attack on both proprietary and open-source RLMs across multiple datasets, explores multimodal and coding‑agent variants, and tests several defenses, concluding that defending against OverThink is challenging and that newer RLMs are even more vulnerable due to higher per‑token costs and increased reasoning token usage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure

Large Reasoning Models (LRMs) use explicit chain‑of‑thought reasoning and large context windows to perform complex tasks, but these features create new attack surfaces. The paper introduces SRCF, an attack that steers LRMs by prepending counter‑aligned few‑shot conversations with explicit CoT traces, causing unsafe outputs on harmful queries and unwarranted refusals on benign ones, without needing model internals. To counter this, the authors propose ARCF, a post‑training defense that exposes models to counter‑aligned conversational contexts while enforcing aligned targets, improving safety and helpfulness without harming utility.

By Xiangyu Zhou, Saleh Zare Zade, Dongxiao Zhu
arXiv AI
5d ago

Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs

The paper investigates a new attack method called output‑prefix attacks on reasoning LLMs, where an attacker prepends a malicious text to the model’s output, thereby conditioning all subsequent tokens on that prefix. The study systematically isolates the scratchpad reasoning channel as a vulnerable vector and compares three attack types—reasoning‑only, output‑prefix‑only, and combined reasoning‑plus‑output‑prefix—across both exposed and hidden reasoning models. Experiments on three 2026‑era frontier models (Gemini 3 Flash Preview, DeepSeek V4 Flash, and Claude Haiku 4.5) show that reasoning alone is largely ineffective, but adding a trivial output prefix can raise attack success rates to as high as 99% for some models, with contextual prefixes outperforming static ones and susceptibility varying by model.

By Luk\'a\v{s} Br\r{u}na, Robert Bridges, Adam Ek
arXiv AI
Jun 3

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

arXiv:2606. 02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assumption that longer reasoning is consistently beneficial remains under-examined.

By Simone Caldarella, Davide Talon, Rahaf Aljundi, Elisa Ricci, Massimiliano Mancini