arXiv:2604.02007v3 Announce Type: replace
Abstract: Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely adopte...
By Rafael Pardinas, Ehsan Kamalloo, David Vazquez, Alexandre Drouin
arXiv:2609.38409v1 Announce Type: new
Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable r...
By \.Ibrahim Ethem Deveci, Funda Tan \c{C}al{\i}k, Bar{\i}\c{s} Deniz Sa\u{g}lam, Duygu Ataman
arXiv:2606. 31048v1 Announce Type: cross Abstract: This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.
By Gaurab Baral, Aaditya Khanal, Yangyang Tao, Junxiu Zhou
arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.
By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
The paper presents a unified evaluation of seven open reasoning language models across four benchmarks (ARC-Challenge, GSM8K, MATH levels 1–3, and TruthfulQA MC1) using a consistent 238-example subset and three prompting strategies (zero-shot, chain-of-thought, few-shot CoT). It reports not only accuracy but also Wilson confidence intervals, latency, VRAM usage, weighted aggregate performance, Pareto-efficient points, prompt-sensitivity, and compatibility diagnostics, revealing that Gemma-4-26B-A4B tops the weighted score while Gemma-4-E4B offers a strong practical trade-off. The study emphasizes that model rankings shift with prompting strategy and that deployment trade-offs remain crucial, advocating for a deployment-aware, multi-objective evaluation framework rather than a single-score leaderboard.
By Md Motaleb Hossen Manik, Ge Wang
arXiv:2607. 20448v1 Announce Type: cross Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license.
By Simone Angarano, Francesco Bertolotti, Federico D'Ambrosio, Michele Resta, Alessandro Rognoni, Nicol\`o Ruggeri, Dario Salvati, Andrea Valenti, Alberto Veneri, Martin Cimmino