arXiv:2509.00961v3 Announce Type: replace-cross
Abstract: Active learning is a general learning mechanism shared by artificial and human learners. Whether AI can teach humans such a strategy that tra...
By Lun Ai, Johannes Langer, Ute Schmid, Stephen Muggleton
arXiv:2604. 22565v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can reason well, yet often miss decisive evidence when it is buried in long, noisy contexts.
By Shaoang Li, Yanhang Shi, Yufei Li, Mingfu Liang, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Frank Shyu, Luke Simon, Sandeep Pandey, Xi Liu, Jian Li
arXiv:2607. 28657v1 Announce Type: new Abstract: Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users.
By Oliver Savolainen, Emanuele Bastianelli, Hosein Azarbonyad
arXiv:2605. 27642v2 Announce Type: replace-cross Abstract: Soft prompting, also known as continuous prompting, is a parameter-efficient method for tuning LLMs to specific tasks.
By Pitipat Kongsomjit, Suryansh Goyal, Jacob Whitehill
The paper proposes Layer-Informed Fine-Tuning (LIFT), a method that identifies and updates only the most functionally critical layers of large language models (LLMs) using a bottleneck identification mechanism based on sensitivity analysis. By focusing on layers that handle conceptualization, reasoning, and textualization, LIFT aims to accelerate training and enhance performance on reasoning tasks. Experiments demonstrate that this selective fine-tuning approach both speeds up the training process and yields significant performance gains.
By Junning Shao, Siwei Wang, Zhixuan Fang
StateTree is a reinforcement learning approach that improves long‑term dialogue reasoning by building a tree‑structured auxiliary task from limited dialogue data. The method embeds key‑value records across multiple sessions into a binary tree, requiring the model to traverse from root to leaf, retrieve records, compare timestamps, and identify a target question among distractors. Curriculum RL training increases tree depth, and a compositional variant trains the model to combine partial reasoning fragments, enabling cross‑session retrieval, temporal reasoning, knowledge updates, and multi‑hop reasoning while generalizing from 10K‑token to 128K‑token contexts.
By Naen Xu, Wanqing Cui, Yibo Hu, Shixin Hong, Hengyu An, Meiguang Jin, Junfeng Ma, Tianyu Du
SPIRAL is a reinforcement‑learning framework that trains language models to employ three inference primitives—sequential reasoning within a trace, parallel sampling of independent traces, and aggregation of those traces—within a single compute pipeline. The model first generates multiple independent chain‑of‑thought traces in parallel, then produces a final aggregation trace conditioned on them, with all components optimized end‑to‑end for the reward of the aggregated response. Experiments on reasoning tasks demonstrate that SPIRAL scales efficiently with inference compute, achieving up to 11× better scaling efficiency and 15% higher performance compared to the GRPO baseline when all three primitives are scaled.
By Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li, Omar Shaikh, Yoonho Lee, Dorsa Sadigh, Chelsea Finn, Noah Goodman
arXiv:2607. 16205v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts.
By Dayu Wang, Jiaye Yang, Weikang Li, Jiahui Liang, Liwei Qian, Xin Pei, Jizhou Huang
arXiv:2605.06165v2 Announce Type: replace
Abstract: As the widespread adoption of Large Language Models (LLMs) accelerates, token consumption from intermediate reasoning traces increasingly contribut...
By Richmond Sin Jing Xuan, Rishabh Bhardwaj, Soujanya Poria
InternBootcamp is an open‑source framework that offers over 1,000 domain‑diverse task environments for large language model (LLM) reasoning research. It introduces Bootcamp‑Eval, an automatically generated benchmark for comprehensive performance assessment. Experiments show that training on InternBootcamp significantly improves reasoning performance, with a 32B model achieving state‑of‑the‑art results on Bootcamp‑Eval and other established benchmarks, demonstrating that scaling the number of training tasks yields consistent gains.
By Peiji Li, Jiasheng Ye, Yongkang Chen, Linyang Li, Yichuan Ma, Zijie Yu, Ganqu Cui, Haozhan Li, Jiacheng Chen, Chengqi Lyu, Wenwei Zhang, Qipeng Guo, Dahua Lin, Bowen Zhou, Kai Chen
The paper reports that in on‑policy distillation for large language models, reasoning performance can be improved by supervising only a tiny fraction of generated tokens—sometimes just one or two tokens per reasoning trajectory, about 0.05% of all tokens. This sparse supervision consistently matches or exceeds full‑token training across nine teacher‑student setups on mathematical reasoning, and is also validated on coding reasoning, Llama models, and PPO‑based reinforcement learning with verifiable reward. The findings suggest that effective post‑training does not require token‑intensive supervision and may align more closely with natural learning processes that focus on critical reasoning steps.
By Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane
The paper introduces Echo-GRPO, a method that rewrites privileged reasoning traces into a model’s own idiolect to align off‑policy supervision with the student policy’s vocabulary. By preserving semantics through Dual‑Reference Decoding, Echo‑GRPO mitigates gradient clipping on critical reasoning tokens and improves reasoning distillation. The approach is instantiated as VideoEcho‑R1 for video reasoning, yielding consistent gains across multiple multimodal LLM backbones and benchmarks, and it can be applied as a plug‑in to both RL and supervised fine‑tuning frameworks.
By Ji Soo Lee, Jinyoung Park, Seohyun Lee, Jongha Kim, Joonmyung Choi, Jinsung Yoon, Hyunwoo J. Kim