arXiv:2607.22629v3 Announce Type: replace
Abstract: Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate t...
By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv:2508. 09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving.
By Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai, Dengfeng Liu, Qiaobo Hao, Huang Liu, Zhiguo Yang, Ji Xie, Ninglun Gu, Jin Yang, Kailai Zhang, Yelun Bao, Jun Wang
arXiv:2607. 22629v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time.
By Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati
Negative Self-Distillation (NSD) is a new framework for improving large language models by encouraging them to diverge from their own flawed reasoning rather than imitate privileged solutions. Unlike On-Policy Self-Distillation, which can suppress uncertainty and exploratory behavior, NSD generates a question‑specific negative condition (e.g., a careless reasoner) and uses a dynamic gating mechanism to target only reasoning‑critical tokens for penalization. This approach preserves foundational language capabilities while consistently outperforming OPSD and other label‑free self‑bootstrapping reinforcement learning baselines.
By Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng
arXiv:2609.39346v1 Announce Type: new
Abstract: Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (S...
By Bohan Zhang (Southeast University), Linan Yue (Southeast University), Weibo Gao (Hong Kong Polytechnic University), Pengyu Chen (Southeast University), Hong Guo (Southeast University), Yanqi Hao (ZTE Corporation)
Negative Self-Distillation (NSD) is a new framework for improving large language models by encouraging them to diverge from their own flawed reasoning rather than imitate privileged solutions. Unlike On-Policy Self-Distillation (OPSD), which can suppress uncertainty and exploratory behavior, NSD generates a question‑specific negative condition (e.g., a careless reasoner) and pushes the student’s distribution away from it. A dynamic gating mechanism isolates reasoning‑critical tokens so that only behavioral flaws are penalized, preserving linguistic capabilities, and empirical results show NSD consistently outperforms OPSD and other label‑free self‑bootstrapping RL baselines.